Skip to content
Menu
How we work 6 min read ·

How our engineers use coding agents, and what stays human

Bizmap engineering team
The stance
Agents change how fast we reach a first working version, not who is accountable for the code; review stays human and client data stays out of prompts.
A question about this?Ask the team that wrote this. An engineer, not a ticket system, replies.Talk to an engineer →

Our engineers use coding agents: AI models that read a codebase, propose changes, write code and tests, and run commands under an engineer's direction. Clients ask about it, sometimes hoping for a discount and sometimes worried about their data. Both are fair questions. This is how we use agents in practice, what we do not let them decide, and what we think you should ask any partner who uses them.

We are not going to quote a productivity multiplier. We have not measured one we would stand behind, and the numbers that circulate in the industry are rarely measured on work like ours.

What changes for a client

Three things change, and they are worth having.

  • A first working version sooner. In discovery and early build, an engineer with an agent can turn a requirement into a working screen, a document type or a report on a staging site quickly. You get a first working demo in days rather than weeks, and a discussion about something you can click is better than a discussion about a document.
  • Tests and documentation that used to be cut. On a tight project, tests and documentation are what slip first. Agents make both cheaper to write, so they stay in scope. That matters most at upgrade time, when tests are what tell you something broke.
  • Cheaper experiments. Trying two approaches to a design problem and keeping the better one used to be a luxury. It is now often affordable before you commit.

What does not change is who is responsible.

What stays human

Architecture, data-model decisions, security review and every merged change are owned by a named senior engineer, who is accountable for them. This is not a formality, and the reasons are specific to the kind of systems we build.

  • The data model outlives the code. In ERPNext and Frappe, a document type designed wrongly is expensive to correct once real transactions sit on it. Migration, reports, permissions and integrations all depend on it. An agent will happily produce a model that works for the example in front of it. A person who has seen how month-end, audits and upgrades behave decides what the model should be.
  • Fit before build. Whether a requirement should be met by standard ERPNext, by configuration or by custom code is a judgement about your business, your budget and your next upgrade. An agent asked to build something builds it. An engineer asks whether it should be built. We mark every requirement standard, configuration or custom before anything is built, and that marking is done by people.
  • Security and permissions. Who can see which field, which roles can approve, what a portal user can reach. These are reviewed by a person who understands the client's organisation, not inferred from the code.
  • Conversations with you. Workshops, trade-offs, bad news. A model does not sit in your steering meeting.

How a change moves from requirement to merge

The loop is the same whether an agent is involved or not. Only the drafting step is faster.

  1. The requirement comes first. An engineer starts from a signed line in the requirements document, marked standard, configuration or custom, not from a loose instruction typed into a chat window.
  2. The engineer frames the task. What the change must do, which document types and permissions it touches, what it must not change, and how it will be tested.
  3. The agent drafts on a branch. Code, tests and notes, in the custom app, against a development or staging site with test data.
  4. The engineer reads and runs it. Anything the engineer cannot explain line by line is rewritten or thrown away.
  5. A second engineer reviews it. Then it is tested on staging and released with the rest of the sprint, and you see it at the next fortnightly demo.

The agent shortens step three. Steps one, two, four and five are where the quality comes from, and they take the time they take.

Review and tests

Every change is reviewed by a second engineer before it is merged, whoever or whatever wrote it. Code written with an agent gets the same review as code written by hand. In some ways it gets a harder one, because agent output tends to look finished whether it is right or not.

Tests get particular attention. An agent that writes both the code and the test for it can produce a test that checks what the code does rather than what it should do. So a person reads what each test asserts, not only whether it passes. A staging site mirrors production, and nothing is tried on live data by surprise.

Custom work goes into a separate Frappe app in your repository, never into ERPNext's core code, whether a person or an agent typed it. The rules that keep upgrades possible do not loosen because the code arrived faster.

Client data stays out of prompts

Your production data does not go into prompts. Agents work on code, on configuration and on staging or test data. When a problem can only be seen in real records, an engineer investigates it directly, and the agent sees the code path, not the customer list, the payroll or the patient record.

Your code never trains a third-party model. Model calls go through our gateway under no-training terms, and where a client requires it the work can run inside the client's own tenancy.

The gateway

Agents are only as safe as the way they are connected. Every coding agent our engineers use goes through a gateway Bizmap built for its own use. It is provider-agnostic, so we can change models without changing how we work, and it keeps a budget and attribution per client and project.

That answers questions you are entitled to ask. Which tools touched your project? On whose credentials? Within what spending limit? Without a gateway, the honest answer is often "an engineer's personal account", which is shadow IT with your code in it. We would not accept that from a supplier, so we do not do it ourselves. The same gateway carries the model calls in the AI features we build for clients; our rules for delivering and governing AI apply to our own use as well.

What we do not claim

  • No productivity multiplier. We do not know a single number that holds across discovery, configuration, custom code, migration and support, and we will not invent one.
  • No "AI discount" on judgement. Agents shorten some tasks. They do not shorten workshops, user testing, data cleansing, training or the first month-end, and those are where most implementation time goes.
  • No unreviewed code. Nothing an agent writes reaches your repository without a person having read it.
  • No smaller team on a promise. We staff a project for the work, not on the assumption that a model will fill the gaps.

What to ask any partner who uses coding agents

  1. Which tools do your engineers use on our project, and through which accounts?
  2. Does any of our data, production or test, go into prompts? Who decides, and how is it enforced?
  3. Do the model providers you use train on our code?
  4. Who reviews agent-written code before it is merged, and does it get the same review as anything else?
  5. Who is the named person accountable for the architecture and data model?
  6. Where does the code live, and could another team maintain it without your tools?
  7. If you claim a productivity gain, how was it measured, and on what kind of work?

A partner who answers these clearly is using agents responsibly. A partner who answers with a percentage probably is not measuring anything.

If this is your situation

If you are choosing a partner and want to know how AI fits into delivery, How we work sets out our engineering standards and what we guarantee. If you want AI inside your own ERP workflows rather than in our engineering, start with how we deliver and govern AI. Either way, talk to an engineer about your project.

Start a conversation

Want this for your operation?

Tell us which case looked closest to yours.