AI engineering with a security company's reflexes
We design, deploy, and optimise AI systems inside your perimeter: local and sovereign LLM deployments, model selection backed by measurements rather than vendor claims, retrieval over your own material with its access rules preserved, and the guardrails that make the result stand up to a security review.
The AI consultancy market sells enthusiasm. We sell measurements: what a model actually does on your workload, what the hardware actually costs, and what your auditor will actually accept. Everything we deploy is built to run where your data already lives.
Where we are typically brought in
Organisations that need AI over sensitive material without sending it to a cloud provider
Teams with a GPU purchase to justify and no measured basis to size it
AI pilots that stalled on performance, cost, or a security review
What an engagement covers
Local and sovereign LLM deployment: model serving, GPU sizing, and quantisation chosen from measured throughput, not datasheets
Model selection bake-offs: candidates benchmarked on your workload and in your language before any hardware is bought
Retrieval and knowledge-base design over the manuals, tickets, and runbooks you already have, under your existing access rules
Inference performance tuning: context caching, batching, and prompt structure that cut latency and hardware cost
Guardrails and egress control between your AI and the internet, enforced by architecture rather than by policy
Integration with the stack you run, and the evidence trail your auditors will ask for
Measured before bought
Model and hardware choices are made from benchmarks run on your workload, in your language, before anything is purchased: tokens per second under real prompt sizes, time to first token at your context lengths, and quality on your material rather than on a public leaderboard.
When a smaller or local model is good enough, the measurements say so and you keep the difference. When it is not, you learn that before the hardware arrives, not after.
Defensible by architecture
Guardrails that live in a policy document fail the first time somebody is in a hurry. We build the constraint into the deployment instead: what may reach the internet, what must stay local, and what gets logged are properties of the architecture, verifiable from your own network records.
That is the same discipline our products are built on, applied to systems we build for you.
Products that carry this work
When the engagement calls for a platform rather than a build, these are ours.
Marmot
Sovereign AI with a Gate
On-premises AI with a per-request gate that decides what may reach a cloud model, and what never leaves the building.
Learn More about MarmotIbex
Air-gapped AI Security & Attack Surface Platform
See everything, connect nothing: air-gapped AI security and attack-surface management for the classified, isolated, and operational networks the cloud can never reach.
Learn More about Ibex
