Working with businesses in the US, UK, UAE, Canada and Australia info@vrisic.com

AI architecture consulting and engineering for systems that must hold up in production

For CTOs and heads of engineering with an LLM or agent prototype that works in a demo but not yet at scale. We review it, measure it, harden it and leave your team with the tests and dashboards to keep it that way.

Updated Reviewed by Nitesh Dan Charan

What is AI architecture consulting and engineering?

AI architecture consulting and engineering is the work of turning an AI prototype into a system you can run in production: choosing and routing models, designing retrieval and agent orchestration, building evaluation and observability, controlling cost and latency, and closing security and governance gaps. Vrisic reviews your existing LLM or agentic AI system, measures how it performs today, then hardens it with working code in your repository rather than a slide deck.

Key facts
Who it suitsCTOs and engineering leads with a prototype that is not production ready
Architecture review2 to 4 weeks, $8,000 to $20,000
Evaluation and hardening4 to 10 weeks, $20,000 to $60,000
Ongoing engineeringRetainer from $4,000 per month
Frameworks we map toOWASP Top 10 for LLM Applications, NIST AI RMF, ISO/IEC 42001, EU AI Act
What you keepTarget architecture, golden sets, dashboards, runbooks and code
Watch with sound, 70 seconds

AI architecture and engineering for systems that must hold up, explained

  1. 1Review the architecture and the risks
  2. 2Build an evaluation suite on real cases
  3. 3Add guardrails and fallbacks
  4. 4Monitor quality, cost and latency
Read the video transcript

AI architecture and engineering for systems that must hold up. Evaluation, guardrails and observability for AI in production. Here is the problem. The prototype works in a demo, not at scale. Nobody can say how accurate it really is. And costs and latency drift every month. Here is how it works. First, review the architecture and the risks. Second, build an evaluation suite on real cases. Third, add guardrails and fallbacks. And finally, monitor quality, cost and latency. What do you get? A written report and a fix plan. Dashboards in your own account. Your team trained to run it. Architecture review $8,000 to $20,000, in 2 to 4 weeks. The easiest way to start is a $1,500 AI Pilot on your own data. You see a working version in ten business days, and the fee is credited to the full build. Book a free call, or message us on WhatsApp, at vrisic.com.

Why AI prototypes stall before production

A prototype proves that a model can do a task once. Production asks whether it does the task correctly thousands of times a day, within budget, without leaking data, and in a way you can explain to a customer or regulator. Most teams find the gap after launch, when a model update changes answers overnight.

Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls (Gartner, June 2025). All three causes are architecture problems. Cost is controllable with routing, caching and budgets. Value becomes provable when you measure task success. Risk becomes manageable when you model threats and log every action.

The usual symptoms are familiar: nobody can say whether last week’s prompt change helped, answers cannot be traced back to sources, the agent has broader API access than any employee, and there is no plan for the day your model provider retires the version you tested on. A review names each one and ranks the fixes by risk.

What an AI architecture review delivers

  • A current state diagram of every model call, data flow and tool
  • Baseline scores for quality, latency at the 95th percentile and cost per task
  • A starter golden set of real cases your team can extend
  • A threat model mapped to the OWASP Top 10 for LLM Applications
  • A governance gap check against NIST AI RMF and the EU AI Act
  • A target architecture with a prioritised, costed roadmap
  • A readout session for engineering and leadership

A reference AI architecture for LLM and agent systems

We judge every system against the same request path. Each stage has one job, and the model is deliberately the easiest part to swap.

Whether you run a copilot, an enterprise RAG assistant or are building agentic AI systems with many tools, a production request passes through the same seven stages.

  1. Entry and identity. Every request carries the user’s identity, tenant and permissions from your identity provider. Nothing downstream should have to guess who is asking.
  2. Orchestration. Deterministic code decides what can be decided without a model. Where judgement is needed, an agent graph in LangGraph or the OpenAI Agents SDK takes over, with explicit states, step limits and approval nodes.
  3. Context assembly. Retrieval, memory and tool results are gathered, filtered by the user’s permissions, ranked and trimmed to a token budget. Oversized context is the most common hidden cost we find.
  4. Model gateway and router. One internal service owns every model call. It routes by task, applies timeouts, retries and fallbacks to a second provider, caches repeatable results, redacts personal data and records tokens and cost.
  5. Tools behind a policy layer. Agents reach systems through typed tools, often Model Context Protocol servers, and each tool checks permissions and business limits before it acts.
  6. Output validation. Responses are checked against schemas, citation rules and content policies before a user or system sees them.
  7. Telemetry and evaluation. Every stage writes to one trace, and the release pipeline runs the golden set before anything ships.

Model selection and routing sit at the heart of this. We score candidate models, typically from the GPT 5 family, Claude, Gemini and open weight options such as Llama or Mistral, on your own golden set for accuracy, latency, cost per thousand requests, context length, data terms and regional availability. The result is rarely one model. A small, fast model classifies and routes, a stronger one handles hard reasoning, and a fallback provider covers outages. When data cannot leave your network, we assess private and self hosted LLMs on the same scorecard.

We use the same structure for new builds through custom AI software development.

LLM evaluation: golden sets, LLM as a judge and human review

Without evaluation, every prompt change is a guess. With it, releasing a new model becomes a routine decision backed by numbers.

Evaluation is the largest gap we find in systems built under deadline pressure. We build it in layers, each catching what the others miss.

  • Golden sets. A versioned collection of 150 to 500 real inputs with expected outputs or grading notes, drawn from production logs and expert input. Cases are tagged by type and difficulty, and every production incident becomes a new case.
  • Automated checks. Cheap, deterministic tests run first: valid JSON, required fields, correct tool chosen, citation present, retrieval recall at 5 and 10 results, refusal where the policy requires it.
  • LLM as a judge. For qualities code cannot check, such as faithfulness to sources, completeness or tone, a separate model grades each output against a written rubric. We calibrate the judge on human labelled examples, measure its agreement with people, randomise answer order to reduce position bias, and prefer a judge from a different model family to the one under test.
  • Human review. Domain experts review a weekly sample, every case where the judge and the automated checks disagree, and all high stakes outputs. Their labels feed back into the golden set and the judge calibration.

The golden set runs in CI on every pull request and on a nightly schedule against production models, because providers do update model behaviour. Release gates block a change if any tracked metric falls below its threshold. Online evaluation then samples live traffic to catch drift. For retrieval heavy systems we also track the metrics described in our RAG vs fine tuning guide, since the right fix for a bad answer depends on whether retrieval or generation failed.

AI engineering workstreams we take on

After the review, most clients ask us to fix a handful of these. Each one ships as working code and documentation in your repository.

  1. 01

    AI agent architecture and orchestration

    We redesign agent systems so behaviour is explicit and testable: clear states, bounded loops, typed tools, approval nodes and durable execution for long tasks. Where a multi agent AI design is justified, we define the supervisor, the handoffs and the shared state, and we collapse agents that only add cost. Teams building new agents from scratch usually start with our AI agent development services instead.

    Best for: Teams whose agents loop, stall or behave differently on every run

    • Agent graphs in LangGraph or the OpenAI Agents SDK
    • Step, time and spend limits per run
    • Tools exposed through Model Context Protocol
    • Temporal for long running work
  2. 02

    AI observability

    We instrument every model call, retrieval and tool call into a single trace with prompt version, tokens, cost and latency attached, then build dashboards and alerts the on call engineer will actually use. Traces link to user feedback, so a reported bad answer becomes a golden set case the next day.

    Best for: Systems where nobody can explain why a specific answer was wrong

    • Langfuse, LangSmith or Arize Phoenix
    • OpenTelemetry GenAI conventions
    • Alerts on quality, cost and latency
    • Personal data masked in traces
  3. 03

    Cost control

    We break spend down per feature, per tenant and per successful task, then cut it with model routing, prompt and context trimming, semantic and exact caching, batch processing for non urgent work, and hard budgets on agent loops. Every change is checked against the golden set so quality holds while the bill falls.

    Best for: Products where model spend is growing faster than revenue

    • Cost per successful task as the headline metric
    • Per tenant budgets and rate limits
    • Routing to smaller models where scores allow
    • Monthly cost forecast
  4. 04

    Latency engineering

    We measure time to first token and total response time at the 95th percentile, then reduce them with streaming, parallel tool calls, smaller routing models, prompt caching offered by providers, fewer sequential agent steps and region choices closer to your users.

    Best for: Chat, voice and in product assistants where speed shapes adoption

    • Latency budget per stage of the request
    • Streaming and parallel retrieval
    • Provider and region benchmarking
    • Load tests before launch
  5. 05

    Reliability patterns

    Model APIs fail, slow down and change. We add timeouts, retries with backoff, circuit breakers, fallback models from a second provider, idempotent tools so a retried action never runs twice, schema validation with automatic repair, and graceful degradation that hands work to people rather than failing silently.

    Best for: Systems that customers or staff depend on during business hours

    • Multi provider fallback
    • Idempotency keys on every write tool
    • Dead letter queues for failed jobs
    • Runbooks for common incidents
  6. 06

    Enterprise RAG architecture

    We review ingestion, chunking, embeddings, hybrid keyword and vector search, reranking and permission filtering, then fix whichever stage is losing answers. Larger rebuilds move to our enterprise RAG and AI search team.

    Best for: Knowledge assistants that miss obvious answers or show documents users should not see

    • Retrieval recall measured per document type
    • Permission aware retrieval
    • Hybrid search and reranking
    • Freshness and reindexing jobs

How an AI architecture review works

A fixed scope, fixed price engagement of 2 to 4 weeks. Hardening, if you want it, is quoted separately once you have seen the findings.

  1. 01

    Kickoff and access

    Days 1 to 3

    We sign an NDA, agree goals and constraints, and get read access to the repository, infrastructure, logs and any existing test data. We interview the product owner, the lead engineer and whoever handles security or compliance.

    You get: Agreed goals and scope, Access checklist

  2. 02

    System walkthrough and code read

    Week 1

    We trace real requests end to end, read the orchestration, prompt and tool code, and draw the current architecture including every model call, data store and external service.

    You get: Current state diagram, Dependency list

  3. 03

    Baseline measurement

    Weeks 1 to 2

    With your team we assemble a starter golden set of real cases and measure quality, latency at the 95th percentile and cost per task.

    You get: Starter golden set, Baseline scorecard

  4. 04

    Threat model and governance check

    Week 2

    We map data flows and trust boundaries, test for prompt injection and excessive agency, review tool permissions and data retention, and compare current controls with NIST AI RMF and the EU AI Act where relevant.

    You get: Threat model, Security findings, Governance gap list

  5. 05

    Target architecture and roadmap

    Weeks 2 to 3

    We design the target architecture and rank every fix by risk reduction, effort and cost, separating quick wins from structural changes. Each item has an estimate your team can plan around.

    You get: Target architecture, Prioritised roadmap

  6. 06

    Readout and handover

    Weeks 3 to 4

    A working session with engineering and a shorter briefing for leadership. You keep every artefact, and your team can implement the roadmap alone, with us, or with another partner.

    You get: Written report, Readout sessions, Recorded walkthrough

Vrisic vs big consultancies vs freelance AI consultants

Each option fits a different job. Here is an honest comparison, including where we are not the right choice.

Large consultancyFreelance AI consultantVrisic AI architecture and engineering
Who does the workPartners sell, mixed teams deliverOne individualSenior engineers who also write the code
Board level strategy and change management Yes No Technical strategy only
Large programmes across many countries Yes No Focused teams, not large programmes
Hands on fixes in your codebase Sometimes, often subcontracted Yes Yes
Evaluation and observability built, not just recommended Varies Depends on the person Yes
Continuity if someone leaves Yes No Small team, shared knowledge
Time to startWeeks to monthsDays1 to 2 weeks
Typical cost referenceDay rates up to about £2,855 in UK public sector dataHourly, widely variableReview from $8,000, fixed price
Best whenYou need an organisation wide AI programmeYou need a narrow task done quicklyYou need a prototype made production ready

Rate reference: UK government rate cards compiled by Fulkerson Advisors (September 2026) show Deloitte at £450 to £2,450 per day and KPMG at £400 to £2,855 per day.

AI architecture consulting pricing

Three ways to work with us, from a one off review to a standing engineering team. You get one fixed price in writing after a free scoping call.

Architecture review

Teams with a working prototype who need an honest view before scaling

$8,000 to $20,000£6,400 to £16,000AED 29,000 to AED 73,000

2 to 4 weeks

  • Code, infrastructure and data flow review
  • Baseline quality, latency and cost scorecard
  • Starter golden set
  • Threat model and governance gap check
  • Target architecture and costed roadmap

Ongoing engineering retainer

Teams that want senior AI engineering capacity every month

From $4,000 per monthFrom £3,200 per monthFrom AED 14,700 per month

Monthly, rolling

  • Agreed engineering days each month
  • Model upgrade testing and rollout
  • Monthly quality and cost report
  • Architecture advice on new features
  • Priority incident support
Not ready to commit to the full build? Start with a 10 day AI Pilot on your own data for $1,500 (£1,200, AED 5,500). The full fee is credited if you continue.See the pilot

Typical ranges; you get one fixed price in writing after a free scoping call. Model and hosting usage is billed at cost. For context, US federal rate data puts AI consulting at $130.82 to $194.82 per hour (Fulkerson Advisors, September 2026), and published consultancy pricing runs from about $25,000 for a focused assessment to $500,000 or more for enterprise strategy (Master of Code, September 2026). Build costs are covered in our AI development cost guide.

What changes the cost of an AI consultant engagement?

The price of AI consulting services depends mostly on how much there is to examine and how much we are asked to change.

Cost driverEffect on budgetWhy it matters
Advisory only or hands onHighA review that ends in a report costs far less than one where we also build the evaluation suite, gateway and dashboards. Most clients choose review first, then hardening.
Number of AI features and agentsHighEach distinct feature or agent needs its own golden set, threat review and metrics. Three agents sharing tools cost less to review than three unrelated systems.
Compliance scopeHighHealthcare, financial services or EU AI Act high risk use cases add documentation, evidence gathering and stricter testing.
Existing test dataMediumIf you already have labelled examples or user feedback, the golden set comes together quickly. If not, building one with your experts takes extra time.
Code quality and documentationMediumPrototypes spread across notebooks and scripts take longer to map than clean, documented code.
Data residency and hostingMediumMulti region, private network or self hosted model setups add infrastructure review and benchmarking work.
Speed of accessLowDelays in granting repository, cloud or log access extend the calendar more than the effort. We send the access checklist before kickoff.

Have a prototype that works in the demo but worries you in production?

Walk us through it in a 30 minute scoping call. You leave with the three risks we would look at first and a fixed price for the review, whether or not you hire us.

Book a free scoping call

Threat modelling, security review and AI governance

We map every system to the frameworks your security team, auditors and regulators already use. We help you meet them; we do not issue certificates.

OWASP Top 10 for LLM Applications

We test against the OWASP Top 10 for LLM Applications, including prompt injection, sensitive information disclosure, improper output handling, excessive agency and unbounded consumption, and report each finding with a fix.

Threat modelling per data flow

We draw trust boundaries around users, models, retrieval stores and tools, then walk attack paths using techniques catalogued in MITRE ATLAS. Indirect prompt injection through documents and emails gets particular attention.

Agent permissions and tool scopes

We review each agent’s service identity, tool list and write access against least privilege, check that limits live in code, and confirm audit logs capture every action and approval.

NIST AI RMF alignment

We map controls to the Govern, Map, Measure and Manage functions of the NIST AI Risk Management Framework and its Generative AI Profile, a common reference in US enterprise procurement.

ISO/IEC 42001 readiness

ISO/IEC 42001 sets requirements for an AI management system. We produce technical evidence, such as risk assessments, evaluation records and monitoring, that fits into your management system. Certification is done by accredited auditors.

EU AI Act classification

We help classify each system under the EU AI Act. Transparency duties apply from August 2026, and the AI Omnibus moved most high risk obligations to December 2027 (White & Case, July 2026).

Tooling for evaluation, observability and orchestration

We work with the tools you already run where possible and add open, well supported ones where there are gaps. More detail in our comparison of AI agent frameworks.

Evaluation

promptfooRagasDeepEvalBraintrustOpenAI Evals

Golden sets, judge rubrics and release gates that run in CI.

Observability

LangfuseLangSmithArize PhoenixOpenTelemetryDatadog LLM Observability

One trace per request across models, retrieval and tools.

Gateways and model access

LiteLLMPortkeyAWS BedrockAzure OpenAIGoogle Vertex AI

Routing, fallbacks, caching and spend limits in one place.

Orchestration

LangGraphOpenAI Agents SDKModel Context ProtocolTemporal

Explicit agent state and durable execution for long tasks.

Retrieval

pgvectorElasticsearchOpenSearchQdrantCohere Rerank

Hybrid search with permission filters applied at query time.

Guardrails and red teaming

NeMo GuardrailsLlama GuardMicrosoft PresidioPyRITgarak

Input and output checks, PII detection and automated attack testing.

AI architecture and data residency in the US, UK and UAE

Where models run and where logs are stored is an architecture decision. It differs by market, and it is hard to change later.

US

United States

US reviews usually centre on HIPAA for health data, CCPA/CPRA and California’s rules on automated decisionmaking technology, and enterprise customers asking for NIST AI RMF alignment in security questionnaires. We keep models, vector stores and traces in US regions and check that every provider’s data retention settings match your contracts.

  • US regions for models, stores and traces
  • Provider retention settings documented
  • Pricing in USD
  • Overlap with Eastern and Pacific time
UK

United Kingdom

UK systems fall under UK GDPR and the Data Protection Act 2018 as amended in 2025, and the ICO expects a data protection impact assessment for most AI processing of personal data. If you serve EU users, the EU AI Act may also apply. Our UK buyer’s guide covers local questions.

  • UK or EU hosting for models and logs
  • DPIA inputs from the architecture review
  • Pricing in GBP
  • Full overlap with UK hours
UAE

United Arab Emirates

UAE clients often need models and data kept in country under the UAE PDPL (Federal Decree Law No. 45 of 2021), DIFC Data Protection Law No. 5 of 2020 or ADGM rules. We benchmark model availability in Azure UAE North and the AWS Middle East (UAE) region, and add Arabic cases to every golden set. See our UAE AI development guide.

  • In country hosting options compared
  • Arabic and English evaluation sets
  • Pricing in AED
  • Overlap with Gulf Standard Time

How to choose an AI architecture consultant

Seven questions to ask any AI consulting firm or independent consultant before you share your codebase.

  1. 1

    Will you write code, or only recommendations?

    Many AI consultants stop at a report. If your team is stretched, you need people who can build the evaluation suite and gateway themselves, not just describe them.

  2. 2

    How will you measure our system before changing it?

    Every serious review starts with a baseline on real cases. Without one, nobody can prove the fixes worked, and the engagement becomes a matter of opinion.

  3. 3

    How do you validate an LLM as a judge?

    Listen for calibration against human labels and measured agreement. A consultant who trusts judge scores without checking them will hand you confident but meaningless numbers.

  4. 4

    Which security frameworks do you test against?

    Expect the OWASP Top 10 for LLM Applications at minimum, with agent specific checks for tool permissions and indirect prompt injection.

  5. 5

    Are you tied to a model vendor or platform?

    Resale margins and partner targets can bias model and cloud advice. Ask directly how they are paid and whether they earn anything from the tools they recommend.

  6. 6

    What do we keep when the engagement ends?

    You should keep the golden sets, rubrics, dashboards, diagrams and code in your own repository and accounts, with no licence needed to use them.

  7. 7

    Who exactly will do the work?

    Ask to meet the engineers, not only the sales lead, and ask them how they would approach your hardest open problem. See how we work for our answer.

Questions we hear every week

Still unsure about something? Ask us on a call. We will give you a straight answer, even if it means we are not the right partner.

How much does an AI consultant cost?

It depends on who you hire. Published US federal rate data puts AI consulting at about $131 to $195 per hour, and UK government rate cards show large consultancies charging £400 to £2,855 per day. Vrisic prices by outcome instead: an AI architecture review costs $8,000 to $20,000, evaluation and hardening $20,000 to $60,000, and an ongoing engineering retainer starts at $4,000 per month, fixed in writing after a free call.

What does an AI consultant do?

An AI consultant helps a company decide where AI fits and how to build it safely. Strategy consultants focus on use cases, business cases and operating models. Engineering consultants, which is what Vrisic is, review system architecture, choose models, design evaluation and monitoring, find security gaps and often write the code that fixes them. Ask any consultant which of the two they actually deliver.

What is AI architecture?

AI architecture is the design of how an AI system is put together: how requests flow, which models handle which tasks, how context is retrieved, how agents call tools, how outputs are checked, and how the whole thing is measured, secured and paid for. Good AI architecture keeps the model replaceable and makes quality, cost and risk visible to the people who own the system.

What is included in an AI architecture review?

A Vrisic architecture review includes a code and infrastructure walkthrough, a baseline measurement of quality, latency and cost on a starter golden set, a threat model mapped to the OWASP Top 10 for LLM Applications, a governance gap check, and a written target architecture with a prioritised roadmap. It takes 2 to 4 weeks and ends with a readout for your engineering and leadership teams.

What is LLM evaluation?

LLM evaluation is the practice of scoring an AI system’s outputs against defined expectations so you know whether a change made it better or worse. It combines automated checks such as schema validity, retrieval recall and exact answers, model graded scoring with an LLM as a judge, and human review of samples. Evaluation runs before every release and continuously on production traffic.

Can we trust an LLM as a judge to grade our AI system?

Only after you calibrate it. A judge model should score against a clear written rubric, be tested against a few hundred human labelled examples, and reach agreement with people that you have measured, not assumed. Judges have known biases, such as favouring longer answers or the first option shown, so we randomise order, use a different model family where possible and keep humans on disagreements.

What is AI observability?

AI observability means recording what every AI request did: the prompt version, retrieved documents, model and parameters, tool calls, tokens, cost, latency and final output, linked into one trace. It lets engineers explain any single answer, spot rising cost or error rates, and feed real failures back into the test set. Tools include Langfuse, LangSmith, Arize Phoenix and OpenTelemetry.

How do you reduce LLM costs in production?

The biggest savings usually come from routing easy requests to smaller models, caching repeated prompts and retrieval results, trimming oversized context, capping agent loops, and batching work that does not need an instant answer. We measure cost per successful task before and after each change against the golden set, so savings never come from quietly lowering quality.

When is a multi agent AI system worth the extra complexity?

Multi agent AI makes sense when a process has genuinely separate stages, different permission boundaries or specialist knowledge that would overload a single prompt. It is not worth it when one agent with good tools can finish the job. Multiple agents add coordination bugs, higher token spend and harder testing, so we usually start with one agent and split only when traces show a clear reason.

Does the EU AI Act apply to companies outside the EU?

It can. The EU AI Act applies to providers and deployers outside the EU when their AI system is placed on the EU market or its output is used in the EU. Transparency duties for certain AI systems apply from August 2026, while most high risk obligations were moved to December 2027 by the 2026 AI Omnibus amendment. Take legal advice on your specific case.

How is Vrisic different from a big consultancy for AI work?

Large consultancies are strong on enterprise strategy, change management and programmes across many countries. Vrisic is a small senior engineering team: the people who review your architecture also write the evaluation suites, fix the code and set up the dashboards. That usually means a faster start and lower cost for technical work, and less capacity for organisation wide transformation programmes.

Do we need ISO/IEC 42001 certification for our AI system?

Usually not, unless customers or regulators ask for it. ISO/IEC 42001 is a management system standard for governing AI across an organisation, and certification is done by accredited auditors, not by builders like us. Many teams use it as a checklist to organise policies, risk assessments and monitoring. We help produce the technical evidence an auditor will ask for.

Find the process where AI will pay off first

Message us on WhatsApp for the fastest reply, or book a free thirty minute call. You will leave with the processes most worth automating, whether to build new or upgrade what you have, a realistic timeline and a clear idea of cost, whether or not you work with us.

No obligation. NDA available on request. WhatsApp replies are usually fast; email within one business day. Or start with a 10 day AI Pilot for $1,500, credited in full
Chat with us on WhatsApp Chat on WhatsApp