Services · 03
AI & automationThe routine work, done before anyone sits down to it
Automation for what repeats the same way every time, and AI for the reading and sorting a person does all day. A person still approves what matters.
A clerk who never tires, with a person approving.
Somewhere in your business a person reads all day: orders arriving as WhatsApp messages, supplier invoices as PDFs, the same customer questions a hundred times a week, claims to be sorted before anyone can act on them. Somewhere else, a person repeats the same six steps between the same two screens, every time, exactly the same way. The first is where AI earns its keep. The second needs no model at all.
This work is both halves: plain automation for what repeats the same way every time, and a model for the reading, sorting, judging and drafting that arrives in messy human form. Before anything is built we agree in writing what it must get right, how you would know it had stopped, and what it costs to run per month at your volumes. When the honest answer is that you need better software rather than AI, we say that on the first call.
- Team
- The named owners, for the whole engagement
- Runs in
- Your own accounts, from day one
- Pricing
- Fixed price, billed by milestone; a fixed itemised quote follows the first call
- A queue of orders, messages or claims that someone reads all day
- The same steps repeated by hand between the same two systems
- A ChatGPT experiment that should become real work
- The test it must pass, agreed in writing before the build
- A working system in your own accounts, on real traffic
- A running-cost sheet: what it costs per month at your volumes
- A handover file your team can run it from
Four gates. You accept each one before the next starts.
The same four as every engagement, as they look for this work.
Diagnose
We find the queue: the one your people read all day, and the steps they repeat by hand. We count both, watch how they are handled, and collect real examples, labelled by the people who do the job now. The findings document says what can be automated outright, what needs a model, what must stay with a person, and what it would cost to run.
Prove
The test it must pass is agreed as a number, on your own examples. The smallest version goes live on real traffic with a person approving every output. Accuracy is measured against the test, not the demo, and the running cost is checked against the real bill.
Build
The rest of the job, fortnightly: more of the queue, fewer approvals where the test is met, a reason shown beside every suggestion. Everything runs in your own accounts, on paid business plans where what you send is never used for training.
Hand over
Your team runs it for two weeks while we are still on call. The examples and the test stay in your repository so anyone can check it later, an alert fires when the live work drifts from them, and the review says where it still needs a person.
Quote to dispatch on one workflow, confirmed from the floor.
A furniture maker with two workshops and dealers in nine cities. Orders moved on paper, and nobody knew where one stood without walking the floor.
Job cards and dealer updates retyped by hand
documents typed a second time from a record that already held the same facts

- Quote to dispatchdays from an accepted quote to the piece leaving the workshop, averaged over the month
- Office time spent chasing order statustime at the desk answering where an order has reached, counted against the month before the build
What owners ask before they call.
Is this an AI problem or a software problem?
Often the second. If the work is a fixed rule (an order over a limit needs approval, a part below its reorder level gets ordered), plain software does it exactly, and AI would only add uncertainty. AI earns its place when the work is reading, sorting, judging or drafting things that arrive in messy human form: a WhatsApp order in three languages, a scanned invoice, a complaint. The first call sorts your problem into one pile or the other, and we say which even when the answer is not what you hoped.
What daily work is AI good at in a small business?
Reading what arrives and turning it into something a system can act on: orders from messages, line items from supplier invoices, the intent behind a customer question. Sorting: which claim needs a person today, which can wait. Suggesting: the reorder list, the reply to a routine query, the likely reason a delivery was returned. Drafting: the quotation, the follow-up, the summary of a long thread. In every case a person approves, and the system shows why it suggested what it did.
When has a business outgrown Zapier, Make or n8n?
When a flow has to know something only your business knows: whether this customer is on credit hold, which godown has the part, what the reorder level is this month. Those tools move data between apps well, and for a handful of simple flows we will tell you to use them and go home. They become hard to trust when a flow has to make decisions, when it fails silently at three in the morning, or when the twentieth flow depends on the nineteenth. That is the point at which the connections need to be your own system, watched, tested and documented, rather than a subscription nobody dares to touch. When do you outgrow Zapier, Make or n8n? works through the three tests in full.
What must an AI system get right before it goes live?
The test it must pass, agreed in writing and expressed as a number on your own examples: a few hundred real orders or claims, labelled by the people who handle them today, with the ones they disagreed about marked as such. That set defines the job, defines done, and later tells you when the live work has drifted away from it. A vendor who will not show you the test is selling you the demo. Why the test is published before the model is chosen sets out the rest.
Where does our data go?
Into your own accounts and nowhere else. We build on paid business plans where what you send is not used to train anyone’s model, or on models running in your own cloud account, and every key and credential sits with you. An NDA is signed before you share anything sensitive. If a document must never leave the building, we say whether the job can still be done, and how.
What happens to the ones it is not sure about?
They go to a person, with the reason shown. A system that is unsure and says so is worth more than one that guesses confidently, so the threshold for asking is agreed with you and tuned on real traffic. Every correction a person makes becomes a new example in the test, so the system improves on your work rather than on someone else’s. Accuracy is reported monthly against the test, in plain numbers.
How do we know the automation worked?
Because the counts were agreed before the build and measured after it, on your data. For most businesses they are plain: orders that reached dispatch without a human retyping them, items sold that were already out of stock, the time between the last bill of the day and the books being closed. If a number has not moved, the review says so in writing, and why. Nothing else is promised, which is why it can be a contract term. The furniture maker is the shape of it: quote to dispatch, every stage confirmed on a phone.
Can we start from the ChatGPT experiment we already have?
Yes, and it is usually the best starting point, because it tells us what your people actually want the machine to do. What it lacks is what makes it safe to rely on: a test it must pass, a person in the loop where it matters, a record of every decision, and a running cost you have seen in writing. We keep the part that works and build the rest around it, in your own accounts, so it stops being one person’s browser tab and becomes part of how the business runs.
For your technical adviser
Retrieval, agents, fine-tuning, guardrails, observability, and the plain workflow engine for everything that does not need a model. The evals are published alongside the model, before the build starts.
- Eval set and baseline metrics, published before build
- Production system in your cloud account
- Cost model per request, reviewed with the code
- Runbooks, drift alerts, exit review
Stack Claude · GPT · Gemini · Llama / Mistral self-hosted · pgvector · Qdrant · Python · TypeScript
Why the test is published before the model is chosen
The essay behind the second gate: the test is the contract, and you keep it whatever happens.
Write a paragraph. A founder replies within one working day.
A 30-minute call with both founders follows. No deck and no proposal template, just a first read of whether we are the right shape for the problem.
Or by phone or WhatsApp: +91 94330 54299 · WhatsApp (opens in new tab)