To know whether an automation or AI workflow actually works, let it run unattended for long enough to fail, then audit the evidence it left behind. Check that it ran on schedule, that failures were reported accurately, that its records agree with every other record of the same events, and that its results are compared against a sensible baseline. When I did this with a small trading bot after 851 scheduled runs, the useful findings were not about the strategy. They were a timeout in the job wrapper, two ledgers that disagreed, and trade exits that the headline number hid.

The bot traded no real money. It was a paper trading system I call Nova Trader, built to watch a small basket of crypto assets, pull market data, read news and sentiment data, score signals and keep a simulated portfolio in SQLite. It had five assets, a simulated $10,000 account, small position sizes, a scheduler that woke it every thirty minutes, a watchdog and a few databases. Nothing glamorous, which is why it made a good test.

What did the audit find?

After 851 scheduled cycles over a few weeks, the simulated account ended at $9,984.82, down 0.15 percent. An equal-weight basket of the same assets was down 11.78 percent over the same window. Holding UNI alone was down 23.60 percent, ARB 13.66 percent and WETH 11.56 percent.

So the bot did something useful: it avoided most of a bad market. That was the least interesting finding.

Why do automations fail in operation but not in the demo?

Out of 851 scheduled runs, 848 completed. Three failed, and all three were timeouts. The watchdog script hit a 120-second limit and the scheduler killed it. A manual run worked, the tests passed and the data was still updating. The trading system was fine. The wrapper around it was impatient.

This kind of failure never appears in a demo, because a demo runs once while everyone watches. In operation, a system has to wake up again and again, wait on slow APIs, write its state, report its own health and keep enough evidence that someone can reconstruct what happened later.

A lot of automation breaks in that middle layer. The model or the logic is fine, but the job wrapper, timeout, retry logic, state handling or reporting is weaker than the thing it is supposed to supervise. Here the monitoring reported failures that were not real, which is almost as bad as missing failures that are. Once people learn to ignore alerts, they ignore the real ones too.

Why is a good headline number not enough?

The overall result looked respectable. The trade detail did not.

A backtest over the same period ended at $9,929.94, down 0.70 percent, with a 44.44 percent win rate on closed trades. A low win rate alone is fine if the winners are large. These were not. The average closed win was $6.57 and the average closed loss was $45.79.

The strategy was good at staying out of the market and poor at making money once it entered. Open positions happened to be in profit on audit day, which flattered the balance. The closed trades told the real story. If I had only looked at the latest balance, I would have learned almost nothing.

The business version is familiar. Revenue is up, so nobody looks at the margin on each order. The inquiry form gets submissions, so nobody checks how many get a reply. A single summary figure can hide a process that is quietly losing on every transaction.

What happens when two records of the same events disagree?

The paper trading database recorded 31 trades. The backtest, replaying the same period, recorded 23. That gap mattered more than any performance figure.

A backtest is meant to be a clean replay of historical signals. The paper trader is what actually ran, with its stop behavior, position state, duplicate handling and whatever happened each time the scheduler woke it. Eight trades of difference means the system did not have one version of the truth. The paper trader might be executing stop exits the backtest does not model, processing some signals differently, or doing something the strategy description does not mention.

I did not want a theory. I wanted a reconciliation: same signals, same account assumptions, same execution rules, and every difference explained trade by trade.

Every business with automation meets some version of this. The CRM says one thing, the spreadsheet another, the payment processor a third, and the dashboard shows everything is fine because it reads from the wrong source. The question is not whether the company has data. It is whether its records agree with each other when it matters. I wrote about the same problem in sales systems in why CRM rollouts fail at Japanese SMEs.

How do you audit an automation or AI workflow?

These are the questions I now ask of any automated process, whether it is a trading bot, a Zapier or Make workflow, an AI assistant drafting replies, or a nightly data sync.

Did it run, and on time?

Keep a log of every scheduled run with start time, end time and result. Missing runs are as important as failed ones.

Did it fail cleanly?

When something goes wrong, does it stop safely, report what failed and leave the data in a consistent state? Or does it half-finish and carry on?

Is the monitoring measuring health?

Check that alerts reflect real problems. A timeout on a slow API call is not the same as a broken system, and treating them the same trains people to ignore alerts.

Do the records agree?

Compare the automation’s own records against every other record of the same events: the source system, the accounting data, the CRM. Every mismatch should have an explanation.

Is it better than the alternative?

Compare results against a baseline. For the bot, that meant buy-and-hold for each asset and the basket. For a business process, it might be the manual method it replaced, measured by time and error rate.

Is anyone responsible for it?

Someone should read the logs, handle exceptions and decide when to change it. An automation nobody owns is only working until the day it quietly stops.

What does this mean for Japanese SMEs adopting AI?

The same pattern appears in small businesses in Japan, with fewer crypto assets and more shared folders. A company adds a tool, then another. A spreadsheet survives because it is familiar. A customer record lives partly in email, partly in LINE, partly in a shared drive and partly in one person’s memory. A dashboard exists, but nobody can say which source it reflects. The process works because capable people repair it by hand every day.

From the outside it looks like a working system. Under audit, the gaps show. Usually the problem is not a lack of effort; people are working hard enough to hide the weaknesses. That holds until the business asks for something harder, such as a handover when someone leaves, investor reporting, due diligence, or AI. Then the system has to prove what happened, and it cannot.

This is why I put infrastructure before automation. If a business cannot explain its own state, automation makes the confusion faster. I make the wider case in why AI adoption stalls in Japanese SMEs, and the hidden cost of good enough systems covers what manual repair costs over time.

What I would fix next

The scheduler timeout. Either give the watchdog more time or make it report partial progress, so a slow API is not reported as a failure.

The ledger mismatch. Make the paper trader and backtest agree trade by trade, or explain every difference.

The exits. The strategy does not need to take more risk. It needs to stop accepting $45 losses in exchange for $6 wins.

Automatic baselines. Every run should compare results against holding each asset and the equal-weight basket, so absolute and relative performance are never confused.

None of these fixes is exciting, and that is the point. Nova Trader did better than I expected and worse than I would want. It did not prove the strategy works. It produced enough evidence, after 851 runs, to show exactly what to fix next. That is when an automation starts to become something a business can rely on.

If your company runs automations, AI workflows or dashboards and you are not sure they would survive this kind of audit, start with the free Technology Risk Self-Check. For a written review of how your systems and data actually behave, a Diagnostics review covers it, and I can rebuild the weak parts with your team.


Further reading: why AI adoption stalls in Japanese SMEs · what to check before connecting business apps to ChatGPT · what ChatGPT is actually good for in small business work · Zapier vs Make for Japanese businesses