AI & Machine Learning

Nine of our forty-one documented projects put a model or a rules engine into production. This page is about those nine — which models they use, what they replaced, and what we learned shipping them.

Language models inside a product

Generation and transcription, with the plumbing that makes them usable

A model call is the easy part. The work is everything around it: queueing the slow requests so a user is not watching a spinner, keeping a record of what the model produced so a human can correct it, and making sure a failed call degrades into something other than a blank screen.

Jynak — exam questions from study material

Instructors upload course material; OpenAI drafts candidate questions they then edit and approve. The platform randomises question order across exam versions so no two papers are identical. Built on Node, React and a serverless PostgreSQL database, with WebSockets for collaborative editing of a question bank.

Undu AI — voice search across a supplier marketplace

Six interconnected applications — supplier dashboard, supplier and traveller mobile apps, admin panel, a NestJS backend-for-frontend, and a Python service using OpenAI Whisper to turn spoken queries into structured searches. Multi-language throughout, because the users are not all working in English.

Document and language understanding

Extracting structure from text people wrote for other people

Resumes, inbound messages, scanned forms. The input is unstructured by nature and the volume makes manual handling impossible, but the cost of a wrong extraction is a real decision made on bad data — so these systems are built to show their working and route the uncertain cases to a human.

4dot5 — resume parsing and candidate matching

Daxtra NLP parses incoming resumes into structured candidate records; an NLP resolver matches those records semantically against open roles rather than on keyword overlap. AWS Lambda handles the asynchronous load when a client uploads a large candidate pool at once, with Quartz scheduling the background workflows and MongoDB holding documents whose shape varies by source.

Textellent — sentiment classification on inbound SMS

Stanford NLP classifies incoming messages so that a reply which needs a human gets one. It runs inside a multi-tenant Java platform serving 17 product variants, each with its own campaign types and permission model — the classification feeds routing, not an autoresponder.

Decision engines

When the rules change faster than you can deploy

Not every intelligent system needs a model. Five of our nine are rules engines, and they exist because the business logic changed weekly and the people who understood it were analysts, not engineers. A Drools BRMS lets those analysts edit rules in a spreadsheet and see the result without a release — which is both the point and the thing to be careful about.

Where we have built them

The honest caveat: rule sets grow, and a spreadsheet with four hundred rows is a codebase without tests. On every one of these we built the rule-set versioning and the fact-table evaluation model up front, because retrofitting them once the rules are live is considerably harder.

How we work

What an AI engagement looks like

We start by finding the decision you want to change. If there is no decision — no point where the output of the system causes something different to happen — there is no project worth doing, and we would rather establish that in a week than in a quarter.

From there: a narrow proof of concept against your real data, not a demo dataset; a measurement of how often it is wrong and how expensive each kind of wrong is; then the production build, which is mostly the unglamorous work of queues, retries, audit trails, cost controls and a path for a human to override the model. Undu AI's Whisper integration began as a standalone Python proof of concept for exactly this reason.

Things we will push back on

Replacing a working rules engine with a language model when the rules are known and auditable. Putting a model in front of a decision that carries regulatory consequence without a human in the loop. Any build where nobody has agreed in advance what accuracy would make it a success. And training a custom model when a hosted one with good prompting would get you there for a fraction of the cost — which, on current evidence, is most of the time.

Read the detail

Every project above has a written case study covering the stack, the deliverables and what shipped.

Tell us what you are building

Whether it is a system to build, one to replace, or one that needs rescuing — describe it and we will come back within one working day with who would work on it and how we would start.

Start a conversation See what we have built