Atomicity · our key innovation

AI agents that are predictable and repeatable.

LLMs respond differently each time for the same prompt; they are not predictable and repeatable. This is one reason agents are not doing mission critical work without human supervision. We solved this problem.

An atomic workflow is an AI agent built as a chain of small (indivisible) steps, each of which finishes validly, or stops and says so. Every LaunchAI agent works this way. A step is an action, a decision, or a human review, with named inputs, an expected result, and its own record.

Origins

Where the word “atomic” comes from

Atomos is Greek for uncuttable. In 1808 John Dalton published the first part of A New System of Chemical Philosophy, which gave each element its own atom with its own weight. Chemistry got a unit it could count.

Computing borrowed the word for the same reason. In 1981 Jim Gray described the transaction: a unit of work that commits entirely or not at all, and stays committed. In 1983 Theo Härder and Andreas Reuter named the full set of guarantees ACID. The A is atomicity. It is the reason a bank transfer cannot debit one account and forget to credit the other.

LaunchAI applies the same discipline to agents. The unit is the step. A step does one thing, checks that it happened, and writes its own record. A half-finished step never hands its output to the next one.

The problem

Why one big instruction fails in production

Hand a model a large job, such as "reorder whatever is low," and it starts well. Then it meets what nobody mentioned: a second Save button, a quantity in the wrong unit, a warning dialog. With no instruction left, it improvises. Sometimes harmlessly. Sometimes by converting the wrong recommendation into a purchase order, with full confidence.

The public benchmark agrees. On OSWorld, which scores agents on open-ended tasks on real computers, the best agents in September 2026 still missed roughly one task in seven, by their own self-reported numbers. Remarkable progress since 2024, and still a miss rate no controller would accept on a posting.

Back-office work multiplies the misses. Chain 30 steps that each succeed 99 percent of the time and the run finishes clean 74 percent of the time. The misses that cost money are the silent ones: a run that reports success after posting half a batch.

~1 in 7
OSWorld tasks still missed by the best agentsSeptember 2026, self-reported
0.99³⁰
Thirty 99% steps in a row≈ 74% clean finish
0
Silent partial successes allowed in an atomic chain

Silent partial success has four familiar shapes

Silent

The short read

A long report is only partly captured, and the part travels on as if it were the whole.

Silent

The phantom save

Correct values, Save clicked, and a validation message opens behind another window. Nothing was saved. The log says done.

Silent

The half batch

Twenty of forty receipts post, the session drops, and a rerun from the top posts the first twenty again.

Silent

The confident branch

A decision made on an incomplete list: nothing looks low because half the items were never read.

The answer is structural. Make every step small enough to check, then check it.

Prompt-driven agentFails and leaves a transcript and a guess.
Recorded RPA scriptFails when a selector stops matching, then waits for a developer.
Atomic chainFails at a named step, with the screenshot it saw and a clean place to resume.

Building blocks

The three kinds of step

Every workflow is a chain of three step types. Each has a narrow purpose, named inputs, an expected result, and a record of its own.

The three step types
TypeWhat it doesDone whenExample
Action Operates software: click, type, read, open, save, send The expected screen or value appears Open the ERP and wait for the home page
Decision Examines the current state and picks a branch, by exact rule or model judgment A branch is chosen and logged with its reason Is available quantity below safety stock?
Human review Stops the run and waits for a person A person approves, rejects, or corrects Approve this week's replenishment recommendations

Three sorting questions place any piece of work in the right type

  1. 1Does it touch a screen, a file, or a message? It is an action.
  2. 2Does it choose between paths? It is a decision.
  3. 3Does policy say a person must decide? It is a review, on the map where everyone can see it.

Which decisions get an exact rule and which get judgment is the job of the Dual-Path Rule Engine. How the review step pauses, routes, and resumes is on Human review.

Anatomy of a step

The contract each step carries

A step is a contract. It says what it needs, what it will do, and what the screen must show when it is done. The runtime holds it to that contract on every run.

The expected result is mandatory. A step with no proof of done is not a step. If the inquiry returns no rows, step 3000 fails. It does not report "zero items below safety stock" and let the run conclude that nothing needs ordering.

The outputs are held back. A value exists for later steps only after the step that produced it has verified. That single rule is what keeps a short read from turning into a confident branch.

Anatomy of a step
PartWhat it declaresExample: step 3000, Monitor inventory levels
Name and purpose One sentence a manager can read Read available quantity and safety stock for every stocked item
Inputs Named, typed values the step needs before it starts The warehouse to read, set when the run starts
Actions The ordered moves on screen, each checked before the next Navigate to the balance inquiry, extract the grid, open the report, save a snapshot, send a notice
Expected result What must be true when the step is done A results grid with at least one row, and every row carrying all five fields
Outputs Named values later steps may use, held back until the step verifies {{inventory_balances}}
Boundary What the step may touch Only the applications and folders the workflow declares
Record What it leaves behind Before and after screenshots for each action, the extracted values, timing, outcome

Execution

How an atomic step runs, from inputs to commit

The same sequence runs when the Studio validates a step during building and when the agent runs alone at night.

A step approved in the Studio has already passed through the engine that will run it in production.

  1. Check inputs

    Every named input exists and has the right type. A missing warehouse code fails the step before the agent touches a screen.

  2. Act, then check

    One action at a time through the five moves of the Vision Layer, ending with a fresh capture. No target, or two candidates? It does not guess.

  3. Hold outputs aside

    Extracted values are staged, not handed forward.

  4. Verify

    The expected result is checked against the screen, the file, or the value.

  5. Commit

    Staged outputs become available, the record closes, and the connector's condition picks the next step.

  6. Or fail

    Staged outputs are discarded, the failure is recorded with what the agent saw, and policy decides: retry, send to a person, or stop.

Guarantee

What all-or-nothing means for one step

One boundary, stated plainly: LaunchAI does not roll back your ERP. Atomicity governs the agent's own work, meaning its outputs, its record, and its place in the chain. If one step posted a receipt and the next step failed, the receipt stays posted and the record says so.

Completes or flags

A step completes and verifies, or fails and flags. There is no third outcome called "mostly worked."

A retry starts clean

If a read captured 212 of 300 rows before a dialog stole focus, the 212 are discarded. The rule holds in the Studio and at night.

The run knows where it is

Fix the cause and it resumes at the failed step. Steps that already committed do not run again.

Stop means stop

The record shows which steps committed and which never started.

Effect states

Did the click land?

Reading is safe to repeat. Writing is not. Click Post twice and an ERP may post twice. So every step that changes something in another system ends in one of three plain states.

It landed.

The agent saw the proof: a new receipt number, a changed status, the posted line in the grid. The step commits.

commit

It did not land.

Clear evidence of no change: a validation message, a record-locked warning, the form still open with no document number. The step may try again if policy allows.

retry allowed

Nobody knows yet.

The screen froze after Post. The session dropped. A confirmation flashed and vanished. The agent does not click again. It stops, holds its place, and asks.

hold & ask

Idempotency: where duplicate postings go to die

The third state is where duplicate postings are born. A retry that assumes the first attempt failed turns one receipt into two. LaunchAI treats an unknown result as unknown until something settles it: a person looking at the before and after screenshots and the ERP itself, or a business key the target application checks on its own.

An operation is idempotent when doing it twice has the same effect as doing it once. Stripe's API accepts an idempotency key and, when the same key arrives again, returns the saved result instead of charging twice. A screen has no such field, so the key lives in the business data: a vendor invoice number the ERP rejects as a duplicate, a packing slip number stamped on a receipt. With such a key, "did it land?" becomes a lookup. Without one, a person answers it, because a guess is how the same invoice gets paid twice.

The same thinking guards the front door: a run request carries its own key, so a double-clicked Run button cannot start two runs.

ERP · RECEIPT Packing slip PS-88412 POST ? frozen LOOKUP BY KEY PS-88412 posted 6:31 "Did it land?" becomes a lookup

Resume semantics

What runs again, and what never does

A resumed run is the same run. Its identity, its committed steps, and the values those steps produced carry forward.

Values extracted before the failure are not read again. If step 3000 captured the balances at 6:02 and step 12000 failed at 6:40, the resumed run still works from the 6:02 snapshot, the same numbers the approved recommendations were based on. The record says which snapshot fed which decision.

Resume is not a new run with a shortcut. A new run starts at step 1000 and reads everything fresh. A resumed run continues the story the record has already told.

Resume semantics
SituationWhat runs againWhat never runs again
A read step failed partway through a report The whole read step, from its first row Every step before it
A write step did not land (validation error, locked record) That step, once the cause is fixed The committed writes before it
A write step's result is unknown Nothing, until a person or a business key settles it The write itself, once confirmed as landed
A reviewer rejected an item Nothing; a rejection is a branch, not a failure Not applicable; the run follows the rejection path
The machine lost power mid-step The interrupted step, its last action treated as unknown Anything already committed

Loops

Every item gets the same discipline

Back-office work arrives as lists: 140 count lines, 60 receipts, 300 recommendations. A loop in LaunchAI names the collection it walks and a maximum number of passes. There is no unbounded loop anywhere in the runtime.

  • Each pass runs through the same verified steps, so a failure is pinned to one item at one step, not to "somewhere in the batch."
  • Reading the list uses the viewport ledger of the Vision Layer: no row skipped, no row read twice, however long the table.
  • The maximum is a tripwire. A loop that normally walks 140 lines and suddenly finds 1,400 is reporting that something upstream changed: a flag at 6:05 instead of a long night.

Set the maximum from real volumes, with room for a heavy month.

Worked example

An 18-step inventory workflow

An inventory workflow as drafted in the builder for an Infor ERP. Eighteen nodes read left to right on one screen.

  1. 1000

    Start

    Endpoint→ 2000

  2. 2000

    Open the ERP

    Action→ 3000

  3. 3000

    Monitor inventory levels

    Action→ 4000

  4. 4000

    Configure replenishment and run MRP

    Action→ 5000

  5. 5000

    Review replenishment recommendations

    Human review→ Approved: 8000 · Rejected: 6000

  6. 6000

    Record the rejection

    Action→ 7000

  7. 7000

    End: recommendations rejected

    Endpoint

  8. 8000

    Convert approved recommendations

    Action→ 9000

  9. 9000

    Receive and post materials

    Action→ 10000

  10. 10000

    Manage work in progress

    Action→ 11000

  11. 11000

    Control finished goods

    Action→ 12000

  12. 12000

    Conduct cycle counts

    Action→ 13000

    Failed step: earlier steps stay committed.
  13. 13000

    Approve inventory adjustments

    Human review→ Approved: 16000 · Rejected: 14000

  14. 14000

    Escalate the rejected adjustment

    Action→ 15000

  15. 15000

    End: adjustment escalated

    Endpoint

  16. 16000

    Post approved adjustments

    Action→ 17000

  17. 17000

    Publish the summary

    Action→ 18000

  18. 18000

    End

    Endpoint

The count-and-adjust portion is inventory adjustment work on Infor CloudSuite Industrial (SyteLine), which can run to hours a day for busy plants. Screen names vary by version.

Three things to notice

01

A node is not a click

Step 2000 holds three actions: launch, wait for the home page within a timeout, set the process name. Step 3000 holds six, and its extraction names five fields (item number, warehouse, location, available quantity, safety stock) written to {{inventory_balances}}.

02

People sit where inventory value moves

Nothing converts to a purchase order until a person approves at 5000, and no count adjustment posts until a person approves at 13000. Each rejection ends on its own recorded path.

03

Failure stays local

If step 12000 fails on a mislabeled bin, the fix is step 12000. Materials posted at 9000 stay posted and are not posted twice.

Recovery

Three more failures, three local recoveries

Each of these stops one box and leaves the rest of the chain alone.

In each case the record answers the auditor's question before it is asked: which step, what the agent saw, who intervened, what they found, and what ran afterward.

  1. Step 3000 finds no rows

    A warehouse filter was left on the balance inquiry the evening before. The grid comes back empty, the step fails on its expected result, and MRP at 4000 never runs on an empty picture. A person clears the filter; the run retries from 3000.

  2. Step 9000 hits a locked record

    A buyer has the PO open in another session and the ERP refuses the receipt: clear evidence it did not land. Two policy retries; the lock persists; the step flags. The buyer closes the order, the run retries 9000, and the materials post once.

  3. Step 9000 freezes after Post

    No receipt number, no error. The result is unknown, so the agent does not click Post again. A person gets both screenshots, searches by packing slip, finds the receipt posted at 6:31, and settles it as landed. One receipt exists.

The workflow builder

Box by box

The builder is where a process becomes a map the agent follows and a person can read. It is also where an agent spends most of its working life: one thing changed, the rest untouched.

Workflow builder · Inventory · v14 draft 1000Start2000Open ERP3000Monitor levels 8000 Convert 6000 Reject review_status = Approved

The canvas

Steps run left to right, numbered 1000, 2000, 3000. Boxes mark actions, decisions, human reviews, and endpoints. Connectors carry their conditions in plain text, such as human_review_status = Approved. A published change governs the next run, so the map is never a picture of what the agent used to do.

3000 Monitor inventory levels ACTIONS Navigate · balance inquiryExtract grid · 5 fieldsOpen reportSave snapshotSend notice def hook(ctx): ctx["ok"] = True return ctx

Inside a node

  • A name and a description.
  • The ordered action sequence, each action editable in place.
  • A Python hook that takes the workflow context in and hands it back.
  • Attachments pinned to the step: templates, reference documents, screenshots.
  • The step's own run logs, once the agent has run.
40+ ACTIONS EXTRACT sourcebalance gridfieldsitem · qty · …output{{inventory_balances}}

Inside an action

  • A type, from a palette of more than forty: pointer and keyboard, apps and windows, navigation and extraction, file work, variables, notifications, screenshots, retry, waits, option picking.
  • A source: where the data comes from.
  • Fields: the named values to extract.
  • An output variable, available to any later step as {{name}}.
  • A description, for the record.

Versions

From draft to archive

Two people editing at once are detected, not silently overwritten. Any earlier version can be restored. Each run records the version it executed, and a run cannot execute against workflow material that changed after it was queued. Building and approving for production are separate permissions, and publishing can demand step-up authentication.

When a single step changes, the process owner re-shows it in the Studio and it is validated again on the live application before the new version goes up for approval. The change types are on the Run console page.

Workflow version states
StateWhat it means
Draft Being built or changed. Copilot edits land here.
Pending approval Submitted and waiting for a person with the right to approve for production.
Approved Signed off and read-only. Changing it starts a new draft.
Published The version that scheduled and on-demand runs execute.
Archived Retired from use, kept, and restorable.

Why the ceremony matters

It is written into a federal enforcement order. On August 1, 2012, Knight Capital's order router sent more than 4 million orders into the market in 45 minutes while trying to fill 212 customer orders. The SEC traced the damage to a deployment: a technician did not copy new code to one of eight servers. No second technician reviewed it. No written procedure required one. The old, dormant code on the eighth server woke up.

An agent that posts to your ERP deserves the discipline that router lacked: a second person before anything goes live, and no way for an approved version to drift.

4M+
orders sent
45 min
to send them
$460M+
lost
1 of 8
servers not updated

Copilot

Edits proposed in plain words, applied on approval

Type a change. The Copilot proposes the branch, shows the difference on the canvas, validates it against your permissions and the current version, and applies it only when you approve.

It also answers questions about the map: which steps write to the ERP, what step 4000 does, where the reviews sit.

If a colleague published a newer version while you were typing, the proposal is checked against theirs, not yours. And an approved proposal lands in a draft, which follows the same approval path as any other change. The Copilot shortens the typing, not the controls; it cannot edit an approved version in place.

Copilot examples
You typeWhat the Copilot proposes
"Add a check for items on quality hold before converting recommendations" A decision step between 5000 and 8000, with a branch that skips held items and records them
"Send adjustments above the dollar limit in the approvals table to the controller" A decision before 13000 that reads the limit from a table and routes by amount
"Save the count sheet as a PDF in the inventory folder" A file action appended to step 12000, with the folder named
"Rename step 14000 to 'Send the rejected adjustment to the controller'" A renamed node, with nothing else on the map changed

Reuse

Validated steps and shared maps

A validated step, such as signing in to a supplier portal, serves every agent that needs it. The portal's quirks are learned once.

Reuse runs deeper than steps. When LaunchAI trains on an application, the navigation map it builds is shared by every agent in the company. Rules and tables are shared too: the approvals table that routes inventory adjustments can route purchase requisitions. One edit, one owner, every agent that reads it.

Failure-mode catalog

Every failure stops at one step and leaves a record

None ends in silent success. The failures already walked through above are not repeated here.

Failure-mode catalog
FailureWhat the step doesWhat the owner seesRecovery
Readiness check fails before the start The run does not begin Which check failed: version, machine, or inputs Fix the gap; start the run
Target not found (a button moved) Does not guess; fails The screen with the missing target Re-show the step in the Studio, or retrain the map
Many targets missing after a vendor release Fails at the first affected step A screen that no longer matches the map Retrain the application map; workflows stay as they are
Two plausible targets (Approve and Approve All) Refuses the ambiguity; fails Both candidates marked Tighten the action's wording; re-validate
Validation error on save Records "did not land" with the message The message text on screen Correct the data; retry the step
Session expired Fails at the login screen The login page Sign in again, or use the credential vault entry
One-time code requested Waits for the named person A request for the code The person types it; the run continues
Extracted value fails its type Rejects the value The field and what was read Goes to human review
Model read below its confidence threshold Routes to a person The source region and the reading Reviewer confirms or corrects
A table has no row for this case The decision cannot evaluate; fails The lookup that found nothing Add the row, or give the rule a no-row branch
Review waits too long Escalates; the run holds its age The waiting item and assignment Next person on the path decides
Portal down for maintenance Fails to reach its page The maintenance notice Pause for the window; resume after
Action aimed at an undeclared application Refused before it reaches the screen The refusal in the record Declare the application in a new version, if it belongs

Governance and security

Controls built into the chain itself

Atomic steps make controls enforceable because every control has a place to attach.

Every change is a version

Drafts, approvals, and publication are recorded with who and when.

The chain is fenced

A workflow declares its applications, folders, and databases; the runner refuses an action aimed anywhere else, even if a document on screen asks for it.

Python hooks are checked

Generated code is audited for forbidden imports and runs in a restricted interpreter. That is hardening, not a formal sandbox, and we say so.

Records exclude secrets

Credentials, one-time codes, and model chain of thought are never logged.

The map is the procedure

An auditor reads the same numbered boxes the agent executes, not a description written after the fact.

What to measure

Clean numbers, pinned to a step

Atomicity produces clean numbers because every outcome is pinned to a step. These measures tell you whether a chain is healthy.

Chain health measures
MeasureHow to compute itWhat it tells you
Clean-finish rate Runs finished clean ÷ runs started Overall health; the run console shows it per workflow
Failures by step number Count of flags per step, per month Where the chain is weak; one step with most flags is one fix
Retries per step Retry attempts ÷ step executions Whether a timing or locking problem is building
Unknown results Steps ending "unknown" per month Whether writing steps need a business key
Time to resume Resume time minus flag time Whether flags reach someone who can act
Steps changed per version Nodes edited ÷ versions published Whether changes stay surgical
Model calls per run Model-path steps executed per run Cost drift; exact-rule steps make no model call

Limits

Stated as tradeoffs

Proof after every step costs something. Here is what, plainly.

Verification adds time

Every action is checked before the next, and a vision-driven step is slower than an API call. Where a clean API exists, the agent can use it inside the same atomic contract. For a high-volume, API-only subset, an integration platform may run faster and cost less; LaunchAI also reaches the screens those platforms cannot.

Small steps take more design than one prompt

The Studio carries that load by asking only questions that change what the agent does, but someone still has to say what "done" looks like.

Covers the agent's work, not the ledger

A posting that committed stays committed. Reversing it is a business decision, and a reversal can be built as its own workflow with its own review step.

Unknown results cost a person a minute

That minute is the price of never posting twice.

Recorded RPA can be faster on a stable subset

Per transaction, on licenses you already own. LaunchAI runs the same kind of steps with a proof after each one, and reaches the screens a recorded selector cannot.

FAQ

Questions

Is an atomic step the same as a single click?

No. A step is the smallest unit you can verify, fix, and read on its own. Opening an ERP takes three actions: launch, wait, confirm. They live in one node because they succeed or fail together. Each action inside is still checked against the screen before the next one runs.

How small should a step be?

Small enough that you can say what "done" looks like on screen. If the expected result needs a paragraph, split the step. A good test: could a new hire check this step's outcome in five seconds from its screenshot?

What does a failed step look like to the person who owns the workflow?

A flag on one numbered box. Open it and you see the screen the agent saw, the value it expected, and what it found instead. Fix the cause, then retry from that step. Everything before it stays done.

Can a decision step change data in the ERP?

No. A decision reads the current state and picks a branch. Writing is an action's job. Keeping them apart means the record always shows which step chose and which step changed something, which is the first thing an auditor asks.

Is a human review step atomic too?

Yes. Its inputs are the agent's record and the source document. Its expected result is a decision from an assigned person. Its record holds who decided, what they saw, and what they changed. It cannot half-complete: a decision arrives and the run takes that branch, or nothing moves.

What happens when a step fails at 2 a.m. with nobody watching?

The run stops at that step, or routes the item to a review if the workflow says so, and an alert goes out by the Portal, email, or chat. Nothing downstream runs on bad data, and nobody has to reconstruct the night from a log file.

No prompts. No babysitting. No dumb questions.

Get a demo with our specialist team.

Get demo