Our AI development process, from first call to steady improvement
Every Vrisic project follows the same AI software development lifecycle: scope it honestly, price it once, build it in two week sprints, prove it with evaluation, launch it carefully and keep improving it.
What is the AI development process at Vrisic?
Our AI development process runs in seven stages: discovery and scoping, a fixed price proposal, design and prototype, data preparation, build in two week sprints, evaluation with red teaming and security review, then launch with hypercare. After launch, a monthly optimisation cycle keeps accuracy, cost and speed on target. A focused project reaches production in 2 to 6 weeks; a larger platform takes 3 to 9 months.
| Stages | 7, followed by monthly optimisation |
|---|---|
| Sprint length | 2 weeks, each ending in a live demo |
| Pricing | Fixed price per scope or per phase |
| Quality gates | Evaluation, red teaming, security review |
| Hypercare | 2 to 4 weeks after launch |
| Your time | About 1 to 2 hours a week |
How we build your AI software, step by step, explained
- 1Discovery and a fixed price
- 2Design with your team
- 3Two week sprints with weekly demos
- 4Launch, measure and improve
Read the video transcript
How we build your AI software, step by step. From the first call to steady improvement. Here is the problem. Scope and price that keep moving. No idea what is happening until the end. And code you cannot take elsewhere. Here is how it works. First, discovery and a fixed price. Second, design with your team. Third, two week sprints with weekly demos. And finally, launch, measure and improve. What do you get? One fixed price in writing. Working software every week. Full handover at the end. Start with a free 30 minute call. The easiest way to start is a $1,500 AI Pilot on your own data. You see a working version in ten business days, and the fee is credited to the full build. Book a free call, or message us on WhatsApp, at vrisic.com.
The AI software development lifecycle, stage by stage
Seven stages with clear outputs. Nothing moves forward until the stage before it has produced something you can see.
- 01
Discovery and scoping
1 to 10 daysA free 30 to 60 minute call about the problem, the people affected and the systems involved. For anything beyond a single workflow we add one or two working sessions, look at sample data under NDA and agree success metrics in numbers, such as calls answered, minutes saved per case or answer accuracy.
You get: Problem statement, Success metrics, Risk and data notes
- 02
Fixed price proposal
Usually 2 business daysA written proposal with scope, deliverables, acceptance criteria, team, timeline and one fixed price. Anything uncertain is named as an assumption rather than buried. When too much is unknown to price honestly, we propose a short paid scoping phase instead of padding the estimate.
You get: Scope and assumptions, Acceptance criteria, Fixed price and payment milestones
- 03
Design and prototype
1 to 2 weeksWe map the workflow, design the screens or conversation flows and build a thin prototype on real examples, such as sample calls from a voice agent or real drafts from an AI agent or copilot. Disagreements surface now, while changing direction is cheap.
You get: Workflow map, Prototype, Architecture note
- 04
Data preparation
1 to 3 weeks, often in parallelWe connect data sources, clean and structure what the AI will read, set permissions and build the evaluation set, typically 50 to 300 real examples with answers your team agrees on. If the data cannot support the accuracy you need, you hear it here, before the main build.
You get: Data connectors, Evaluation set, Data quality report
- 05
Build in two week sprints
2 to 20 weeksEach sprint has an agreed goal, ends with a live demo in a staging environment on your cloud account and reports evaluation scores against the test set. You see working software from the first sprint and can reorder priorities at every sprint boundary within the agreed scope.
You get: Sprint demos, Evaluation scores, Release notes
- 06
Evaluation, red teaming and security review
1 to 2 weeksThe system must pass the acceptance thresholds on the full evaluation set, survive a red team exercise aimed at misuse, and clear a security review of access, secrets, logging and data flows. Failures go back onto the sprint board, not into a footnote.
You get: Evaluation report, Red team findings, Security review notes
- 07
Launch and hypercare
2 to 4 weeksWe launch in stages, often to one team, one location or a share of traffic first, with rollback ready. Hypercare means daily review of real conversations and outputs, fast fixes and a short weekly report, ending with a handover session.
You get: Staged rollout plan, Hypercare reports, Handover session
Four quality gates every AI project passes before launch
AI outputs vary, so “it worked in the demo” is not a release criterion. These gates are.
- 01
Evaluation against an agreed test set
Evaluation means scoring the system on real examples with known correct answers, before launch and after every change. We measure what matters for the task: correctness and citation accuracy for search, task completion for agents, field accuracy for documents, and latency and cost per request for everything. Scores are tracked in tools such as Langfuse or LangSmith so trends are visible. Our AI evaluation and engineering service applies the same method to systems others built.
Best for: Proving the system meets the acceptance criteria
- Thresholds written into the proposal
- Rerun on every model or prompt change
- Full results shared, not summarised
- 02
Red teaming for misuse and failure
We deliberately try to make the system misbehave: prompt injection hidden in documents or emails, attempts to extract other users’ data, pressure to promise refunds or give advice it must not give, and inputs in other languages. We combine scripted attacks, open source tools such as promptfoo and garak, and manual testing by an engineer who did not build the feature.
Best for: Anything customer facing or connected to live systems
- Mapped to the OWASP Top 10 for LLM Applications
- Attacks added to the regression suite
- Findings fixed before launch
- 03
Security review
A written review of least privilege access, secret storage, network exposure, logging, retention, personal data redaction and the data terms of every model provider involved. If you have a security team, they review it with us; if you have a questionnaire, we complete it against the actual build rather than a template.
Best for: Every project, scaled to its risk
- Permissions checked per tool
- Data flow diagram for your records
- Provider data terms confirmed
- 04
Launch readiness
The last gate is operational. Monitoring, alerts and cost limits are live, a runbook covers provider outages and drifting answers, your team is trained, and the human fallback, such as a call transfer or review queue, is tested.
Best for: Making sure day one is uneventful
- Runbook and alerting in place
- Fallback to a person tested
- Rollback rehearsed
How AI project delivery feels week to week
Good AI project delivery is predictable: you always know what was done, what comes next and what is blocked. We run a fixed rhythm so nobody has to chase anyone.
A sprint starts with a 30 minute planning call where we agree the goal for the next two weeks. Midway, a short written update lands in the shared channel. The sprint ends with a live demo on real data, the latest evaluation scores and a list of decisions we need from you. Between those points, questions go into the channel and are answered the same working day.
Every decision that changes scope, cost or architecture goes into a decision log in your repository. The people running this rhythm are introduced on our about page.
Communication cadence
- Shared Slack or Microsoft Teams channel from day one
- Sprint planning every two weeks, 30 minutes
- Written progress update every week
- Live demo and evaluation scores every sprint
- Monthly steering call for phased programmes
- One named project lead as your contact
Engagement models compared side by side
Four ways to buy the same team. The right one depends on how settled your scope is and how long the work will run.
| Fixed scope project | Dedicated team | Retainer | Phased programme | |
|---|---|---|---|---|
| How it is priced | One fixed price | Monthly fee for named people | Monthly, from $1,500 for support or $4,000 for engineering | Fixed price per phase |
| Scope | Written and fixed | Your backlog, reprioritised each sprint | Agreed monthly allowance of work | Fixed per phase, open beyond it |
| Commitment | One project | Rolling, with notice | Rolling, with notice | One phase at a time |
| Who carries estimation risk | Vrisic | You | Shared | Vrisic, phase by phase |
| Changes midway | Written change request | Absorbed into the backlog | Swapped within the allowance | Folded into the next phase |
| Best for | One clear deliverable | Product companies with a roadmap | Systems already live | Platforms that combine several services |
Typical build prices for each service are listed on our AI development services overview. Retainer prices in other currencies: from £1,200 or AED 5,500 per month for support, and from £3,200 or AED 14,700 per month for engineering.
What we need from you to deliver on time
AI projects that slip usually slip because of access and decisions, not code. These keep yours on schedule.
- 1
A product owner who can decide
One person who can answer questions, accept work and make trade offs within a day or two. Without that, sprints stall waiting for a meeting.
- 2
System access in the first week
Sandbox or test credentials for the CRM, EHR, ERP, phone system or APIs involved. Late access is one of the most common reasons AI projects run over.
- 3
Real examples for the evaluation set
Historic emails, calls, documents or tickets, with your view of the right answer. Fifty real cases teach us more than a hundred invented ones. Our guide on building an AI knowledge base shows how to gather source material.
- 4
Your rules in writing
What the AI may say and do, what it must never say or do, when it hands off to a person, and your tone of voice. Existing policies and scripts are a fine starting point.
- 5
Time with the people who do the work today
An hour or two with the receptionist, paralegal or coordinator whose work is changing. They know the edge cases no process document mentions.
- 6
A security and legal contact
Someone who can approve data flows, sign the NDA and data processing terms and review security notes, so approvals run in parallel with the build.
- 7
Honest feedback at every demo
If a demo looks wrong, say so then. Changing course in week two costs hours; changing course in week ten costs weeks.
How change requests work on a fixed price project
Change is normal. Unpriced change is what breaks budgets and relationships.
A change request is any addition or change to the agreed scope, such as a new integration, a new channel, another language or a stricter accuracy target. Small swaps of similar size inside a sprint, for example replacing one report field with another, we absorb without paperwork.
For anything larger the process is short. You describe the change, we reply in writing within two business days with its effect on price, timeline and risk, and nothing starts until you approve it. If a change removes work, the price goes down too. Approved changes are logged in the project record, so the final invoice matches a paper trail you have already seen.
Sometimes evaluation results suggest a change, for instance when poor scans hold down accuracy on one document type. We raise it early with options rather than quietly missing the target. Our guide to AI development costs explains why scope moves price so directly.
Want to see this process applied to your project?
In a free scoping call we walk through your use case, name the likely risks and suggest the engagement model that fits, whether or not you hire us.
Handover and documentation you keep
Everything needed to run, change or move the system without us. It lives in your repository and cloud account from day one, not in a zip file at the end.
Code and infrastructure
Deployed to your cloud account throughout the build.
AI assets
The parts many vendors keep to themselves.
Operations
Tracing in Langfuse, LangSmith or OpenTelemetry based tooling.
Design and decisions
Written for an engineer who joins after we leave.
Knowledge transfer
Scheduled before hypercare ends.
Monthly optimisation after launch
An AI system is never finished. Models change, prices fall, your data drifts and users find new questions.
On a managed plan we run a monthly optimisation cycle. We review a sample of real conversations or outputs, rerun the evaluation set, add new failure cases to it and fix what we find. We check whether a newer or cheaper model now passes your thresholds, since switching can lower running costs without hurting quality. Then you get a one page report covering accuracy, volume, cost per task, incidents and what changed.
Model and hosting usage is billed at cost with no markup. Annual running costs typically land at 15% to 30% of the original build, a range consistent with NomadX (2026), which is why cost tuning belongs in the monthly routine. A cycle might tighten retrieval for a document type users ask about more than expected, move a simple classification step to a smaller model, or add a language the call data shows customers need. For sensitive workloads we also review whether a private or self hosted model has become a better fit.
Support options after hypercare
Pick the level of help you need now; you can change it later. Response times are agreed in writing per project.
Full handover
Your team takes over with the documentation, recorded sessions and a final walkthrough. No ongoing fee. We stay available for paid ad hoc work if you need help later.
Managed support
From $1,500 per month (£1,200, AED 5,500). Monitoring, incident response in agreed hours, the monthly optimisation cycle, model and dependency updates and small improvements. Usage is billed at cost on top.
Engineering retainer
From $4,000 per month (£3,200, AED 14,700). Everything in managed support plus an allowance of engineering time for new features, channels and integrations, prioritised with you each month.
Questions we hear every week
Still unsure about something? Ask us on a call. We will give you a straight answer, even if it means we are not the right partner.
What are the stages of the AI software development lifecycle?
The AI software development lifecycle covers discovery, scoping, design, data preparation, build, evaluation, deployment and ongoing monitoring. It differs from ordinary software in two places: data preparation, because the AI is only as good as what it reads, and evaluation, because outputs vary and must be scored on real examples before launch and after every change, not tested once.
How long does it take to get a fixed price from Vrisic?
For a focused project such as one agent or one workflow, you usually get a written fixed price within 2 business days of the scoping call. Larger or less certain programmes may need one or two follow up sessions, or a short paid scoping phase, so the price rests on real data and system access rather than guesswork.
What is red teaming for an AI system?
Red teaming for an AI system is a structured attempt to make it fail or misbehave before real users do. Testers try prompt injection, requests for other users’ data, pressure to break business rules and unusual inputs such as mixed languages. Each successful attack becomes a permanent test case, and every fix is verified before launch.
How often will we see progress during the build?
Every two weeks you get a live demo of working software on real data, with the latest evaluation scores. In between, a written update arrives weekly and questions are answered in a shared Slack or Microsoft Teams channel. You never have to wait until the end of a project to find out whether it works.
What happens if the AI does not reach the agreed accuracy?
Accuracy targets are written into the proposal as acceptance criteria, and we keep working within the agreed scope until the system meets them. If data preparation shows a target is unrealistic with your current data, we tell you before the main build and offer options: improve the data, narrow the task, add human review or stop.
What is hypercare after an AI launch?
Hypercare is a period of close monitoring right after go live, usually 2 to 4 weeks. We review real conversations and outputs daily, fix issues quickly, watch cost and latency, and send a short weekly report. It catches the problems no test set predicted and ends with a handover session and your choice of support option.
What documentation do we receive at handover?
You receive source code and infrastructure templates in your repository, versioned prompts, evaluation sets and scores, a runbook, monitoring dashboards and alert rules, an architecture overview, a data flow diagram, a decision log and recorded walkthrough sessions. It is enough for your own engineers, or another partner, to run and change the system without us.
How much of our team’s time does an AI project need?
Plan on 1 to 2 hours a week from a product owner who can make decisions, plus a few hours early on from the people who do the work today and from whoever approves security and legal terms. The busiest weeks for your team are the first two, when we gather access, examples and rules.
Can we switch from a fixed scope project to a retainer later?
Yes. Most systems start as a fixed scope project and move to managed support or an engineering retainer after hypercare, once the next priorities are clear. The switch needs no rework, because the code, documentation and monitoring already sit in your accounts. You can also step down to a full handover at any time.
Find the process where AI will pay off first
Message us on WhatsApp for the fastest reply, or book a free thirty minute call. You will leave with the processes most worth automating, whether to build new or upgrade what you have, a realistic timeline and a clear idea of cost, whether or not you work with us.