We scope one painful workflow, build against your real data, and prove it before it touches production.
No pilot on synthetic data. If we can't get to a passing bar on your records, you find that out in week three, not after signing a year.
Stages
- Stage 1
Scope
1 weekWe sit with the people doing the work and map the current path, including the exceptions nobody documented. We agree what success would look like as numbers.
we needTwo hours with the team doing the work, and a sample of real documents or questions.
you getA written scope with the per-field accuracy bars, the exceptions list, and a fixed price for the build.
- Stage 2
Evaluation set
1 weekWe build the test set from your records before building the agent, with the correct answers supplied by your team. This is what every later claim is measured against.
we need[N] real documents or questions, and someone to confirm the correct answers.
you getAn evaluation set you keep, and a baseline score for the current manual process.
- Stage 3
Build
2–4 weeksThe agent is built against your real data in a non-production environment, scored on the evaluation set after every change. You see the score as it moves.
we needRead access to source systems, a sandbox for the target system.
you getA working agent, its scores per field, and the list of cases it currently refuses.
- Stage 4
Review-gated launch
1–2 weeksIt goes live with every output passing through the review queue. Reviewers correct what's wrong, and those corrections feed the evaluation set.
we needOne or two reviewers for part of their day, and a decision on the go-live scope.
you getProduction output with a person on every record, and a report on what the queue caught.
- Stage 5
Widen and hand over
OngoingWhere the score holds above the bar, the automatic path widens field by field and vendor by vendor. What stays under the bar stays with a reviewer.
we needA monthly hour to review the scores and agree what to widen.
you getA documented system, the runbook, and the choice to run it yourselves or with us.
How we measure whether the agent is good enough
We build an evaluation set from your own records before we build the agent: a few hundred documents or questions with the correct answer attached, chosen by the people who do the work today. Every change is scored against that set, and the score is visible to you.
The passing bar is agreed with you in stage one, per field, not as one headline accuracy number. A total that is right 99% of the time and a freight charge that is right 80% of the time are different problems.
When the agent is unsure, it does not guess. Output below the field's threshold goes to the review queue with the source document, the extracted value, and the reason it stopped.
| Field | Bar | If below |
|---|---|---|
| invoice_no | ≥ 0.99 | Queue for review |
| total_amount | ≥ 0.99 | Queue for review |
| vendor_match | ≥ 0.98 | Queue with candidate vendors |
| line[].qty | ≥ 0.96 | Queue the line, keep the header |
| po_match | exact | Hold posting, notify buyer |
Bars shown are the defaults we open with; yours are set in stage one.
Engagement models
Pilot
One workflow, scoped and measured, ending in a decision. You leave with the evaluation set and the numbers whether or not you continue.
Suits you if you need evidence before a budget conversation.
Wrong fit if you already know the workflow and want it in production this quarter.
Full build
Scope through review-gated launch, then handover. We build in your environment and document as we go.
Suits you if the workflow is agreed and someone owns the outcome internally.
Wrong fit if the process itself is still changing week to week.
Managed
We run and maintain the agent after launch: monitoring, score reports, model and connector changes, and the review queue if you want us in it.
Suits you if you have no engineer who wants to own this.
Wrong fit if you have a platform team that would rather hold it themselves.
What we don't do
- We don't run pilots on synthetic data. If we can't see your records, there's nothing to prove.
- We don't ship an agent that guesses when it's unsure. Uncertain output goes to a person.
- We don't take on a workflow nobody internally owns. Those builds stall in week four.
- We don't sell seats. If the agent isn't doing work worth more than it costs, cancel it.
- We don't do staff augmentation or bodyshopping.
- We don't claim a certification we don't hold, or an accuracy number we haven't measured on your data.
Trust and security
Written for the person who has to sign off on this.
Facts only. Where a fact depends on your deployment or is not yet confirmed, it is bracketed rather than guessed.
Data handling
where it lives In the deployment you choose. In our managed option, data is stored in [REGION] on [CLOUD PROVIDER]; in a VPC or on-premise deployment it never leaves your account.
retention Documents and answers are retained for [RETENTION PERIOD] by default, configurable per workflow, and deleted on request within [DELETION SLA].
training Your data is not used to train models. We do not send it to model providers for training, and we contract for that with our providers.
subprocessors The current list is [SUBPROCESSOR LIST], available in full on request. We notify you before adding one that touches customer data.
Permission model
source of truth Your identity provider and your source systems. We store a reference to the identity, not a copy of your permission tree.
enforcement point At query time, on every request. The retrieval set is filtered before generation, so out-of-scope content never reaches the model.
revocation Removing access in the source system takes effect on the user's next request. There is no cached grant to expire.
service accounts Each connector runs on its own least-privilege account, scoped to what that workflow reads and writes. Credentials are held in [SECRETS MANAGER].
Deployment options
| Question | SaaS | Your VPC | On-premise |
|---|---|---|---|
| Where does data rest? | our account | your account | your hardware |
| Who holds credentials? | us, scoped | you | you |
| Model calls | our endpoint | your endpoint | self-hosted |
| Network egress | standard | model only | none required |
| Upgrades | by us, continuous | scheduled with you | you apply |
| Lead time | days | [N] weeks | [N] weeks |
What leaves your network
In a VPC or on-premise deployment, the model endpoint is the only external call, and it can be pointed at a model running inside your own network — [CONFIRM WHICH MODELS YOU SUPPORT SELF-HOSTED].
Audit and observability
agent actions Every read, extraction, answer and write is logged with the identity that triggered it and the sources used.
human actions Reviewer decisions are logged with the before and after value, so a posted record can be traced to the person who approved it.
retention Logs are retained for [LOG RETENTION] and exportable to your SIEM via [EXPORT METHOD].
access to logs Readable by the administrators you name. We access them only for support you've asked for, and that access is itself logged.
Compliance posture
We hold no certifications today. We will not claim one we do not have, and we won't put a badge on this page before an auditor's report exists.
Who you'll actually be working with
BizOp Solutions is a small studio. The people who scope your workflow are the people who build it, and they stay on it after it goes live. We take on a limited number of builds at a time so that stays true.
We work mostly with mid-market finance and operations teams, and mostly on workflows that involve documents, permissions, or both.
[TEAM — names, roles and photographs to supply]