Vision Layer Interface

VLI: eliminate integrations!

Software built for humans uses a screen. Our Agents do the same.

VLI is how your Agent can see and operate everything on your computer. It captures the screen, finds the target, acts with real mouse and keyboard events, and captures again to confirm the change.

Coverage

What the screen reaches that an API does not

An API is an entrance the vendor controls. The vendor has to build it, maintain it, and agree to open it for you. Many never built one. Some sell it only above a certain license tier. Some expose half the screens and none of the reports.

The screen is the one interface every vendor ships, because people need it.

Legacy desktop ERPs

Including thick clients delivered over remote desktop, where the local machine receives only pixels.

Terminals and green screens

The basic IBM 3270 display showed 24 rows of 80 characters, and emulators still support that layout.

Portals you don't own

Supplier, customer, government, bank.

PDFs and scans

Read on screen as an ordinary step, with no separate OCR pipeline to build.

Spreadsheets

Including the workbook with twenty years of macros in it.

"Is there an API?"

"Can a person do this with a screen, a mouse, and a keyboard?"

The test for whether an agent can do a job changes.

How it works

One action, five moves

If the change did not happen, the step does not pass. A click that landed on nothing is caught at move five, the moment it happens.

What each move leaves behind, and what a failure at each move does

  1. 01

    Capture

    Windows, tabs, forms, tables, dialogs.

    Produces
    The before screenshot, stamped with the time.
    If it fails
    Nothing has been touched; the step fails before acting.
  2. 02

    Interpret

    Which pixels are a button, a disabled field, a table row, and exactly where each sits.

    Produces
    One resolved target, with its position.
    If it fails
    The target is unresolved (see refusing ambiguity, below).
  3. 03

    Decide

    By exact rule when the step is exact, with a model when it needs judgment.

    Produces
    The branch, with the rule that fired or a summary of the model's reasoning.
    If it fails
    The Dual-Path Rule Engine's guardrails decide what happens.
  4. 04

    Act

    Through the operating system, with real mouse and keyboard events.

    Produces
    The pointer or keyboard events, with coordinates and typed text.
    If it fails
    Verification still runs, so a lost click is caught next.
  5. 05

    Verify

    Capture again and confirm the expected change happened.

    Produces
    The after screenshot, with the target marked.
    If it fails
    The step fails and flags, and its half-made output is discarded.

Every row lands in the run's record. What the record holds, and how to trace a value back through it, is covered on the Audit Trail page.

Grounding

How the agent finds the right button

Grounding turns “the Approve button” into a precise spot on the screen. The VLI works through four layers, cheapest and most exact first.

The model comes last. Most targets are settled by cheap, direct evidence before any model is asked anything.

The layers are ordered by cost and exposure. The first three run on the machine: nothing leaves it. The fourth sends the screen region to the model account your company configured, and it is the slowest to answer. Putting the model last keeps most steps fast, keeps most screens off the network, and makes the common case repeatable, because an element name read from an accessibility tree does not vary from run to run.

The four grounding layers
LayerWhat it readsWhere it earns its place
1. Accessibility tree The element structure an application publishes for screen readers (Microsoft UI Automation on Windows, the Accessibility API on macOS, AT-SPI on Linux) Native applications and browsers that expose one
2. OCR Text read directly from the pixels Labels, values, and anything printed on screen
3. Local element detector A small model on the machine that finds buttons, fields, and rows Custom-drawn controls, terminal screens, remote desktop windows
4. Vision model The screenshot and the target description, sent to the configured model Only when the first three cannot resolve the target

Safety

Refusing ambiguity

Ambiguity is never settled by guessing.

If no candidate qualifies, or two do, the target is unresolved. The workflow then re-plans, routes the step to a person, or stops, as its policy says. The agent does not pick the first match and hope.

The cost of the alternative is concrete. “Delete” and “Delete filter” can sit on the same toolbar. “Post” and “Post and print” can share a dropdown. A tool that takes the first match will eventually take the wrong one, on the one record where it matters.

Coordinates

Display scaling, multiple monitors, and remote sessions

Screens lie about coordinates.

Windows defaults to 96 dots per inch, and at a 120 dpi setting everything grows by 25 percent. A button designed at (100, 48) then sits at physical pixels (125, 60), while its logical coordinates stay (100, 48). UI Automation reports physical coordinates, so a tool that mixes the two systems clicks beside the target. Multiple monitors with different scaling multiply the problem.

The VLI resolves each target on the screen it just captured and confirms the result on the screen it captures next. Whatever the scaling, a click that lands beside its target produces no expected change, and move five fails it on the spot.

Real workdays

The messy parts: popups, loading screens, disabled buttons, crooked scans

Scrolling is the famous test case, and it has its own worked example below. The same discipline covers the rest of the mess an ordinary workday throws at a screen.

How the VLI handles messy screens
SituationWhat the VLI doesWhere it is set
An expected popup, such as a confirmation dialog A step handles it like any other screen Taught in the Studio
A popup or session warning nobody expected The capture sees it; verify fails because the expected change never happened; the step routes to re-plan, review, or stop instead of clicking through The workflow's exception policy
A report or grid still loading wait_for_element holds the step until the named element appears, with a timeout The action's wait condition
A disabled button Interpretation reads the disabled state; a press would change nothing, so the step fails rather than moving on Verify, on every action
Two elements that match one description Treated as unresolved, as described above The workflow's exception policy
A crooked scan or a photo of paper Read on screen as a model-path step with typed outputs Dual-Path Rule Engine
A terminal screen No element tree, so OCR and the local detector read it; navigation uses keys, because the screen was built for a keyboard The actions taught in the Studio
A session that expired overnight The sign-in screen appears instead, so verify fails; a workflow taught to sign in draws on a vault entry, and any one-time code is requested from a person Security

Application training

A map before the first run

An application can be trained once, before any workflow depends on it. In an automated session, a small model explores the software: menus, dialogs, pages, and the moves between them. The result is a navigation tree, a map of every route the model found.

  • Agents arrive knowing the way. Navigation becomes route-following instead of a fresh decision on every screen: fewer model calls, faster runs.
  • The map belongs to the company. Train the ERP once, and every agent that uses it inherits the routes.
  • A vendor redesign means retraining the map. The workflows built on top of it stay as they are.

Training is separate from teaching. The map knows how to reach a screen. The workflow, taught in the Studio, knows what to do once there.

When the vendor changes the screens

Cloud ERPs ship releases on the vendor's calendar, and portals change without notice. Small changes, such as a button that moved, are absorbed, because a target is found by what it is, not where it used to be. A large change fails at the step that meets it, with the screenshot of what the agent found, instead of wandering.

The fix is scoped: retrain the part of the map that changed, and have the process owner re-show any step whose screen moved.

Where to train

Menus differ by role in most ERPs; a clerk's account and an administrator's account see different trees. Train under an account with the same rights the agents will use, so the map holds the routes the agent can actually take.

The action language

Instructions read the way a person would say them

More than fifty actions cover pointer, keyboard, reading, scrolling, waiting, verifying, file dialogs, uploads, and annotation. The Studio writes them. Anyone can read them, which is the point: a controller reviewing a workflow should not need a developer to translate it.

Values in double braces are variables. A scrape writes what it read into a named variable, and any later step can use it. A demonstration value never hides inside an action; it is always a named variable, so the same action serves tomorrow's record.

Worked example

A 2,400-line open purchase order report

A buyer runs the open purchase order report every Monday. It returns 2,400 lines. About 30 fit on the screen at once.

Reading it is harder than it looks. Many modern grids render only the rows in view and swap them out as you scroll; AG Grid's documentation, for one, describes rendering just 10 extra rows past each edge by default. There is no whole table sitting in memory to read. The only way through is to scroll and read, roughly 80 screens of it.

A naive bot fails in three quiet ways:

  • +1It scrolls 31 rows on a 30-row screen and skips a line.
  • −1It scrolls 29 and reads a line twice.
  • ∅It captures a screen before the grid has painted and records blanks.

The stakes are concrete: a skipped line is a late purchase order nobody chases, and a duplicated line inflates open commitments.

The job as atomic steps

  1. Open the report

    Open the report screen in the ERP and set the buyer and date filters from variables. ERP screen and menu names vary by version.

  2. Run and wait

    Run the report, then wait_for_element "Report ready," so the first capture is never a half-painted grid.

  3. Read the total

    Read the report's printed line total into {{report_line_total}}.

  4. Scroll and scrape

    Scroll and scrape the grid through the viewport ledger, writing PO number, line, release, supplier, due date, and open quantity for every row into {{open_po_lines}}.

  5. Verify the count

    verify_value: the ledger's row count must equal {{report_line_total}}. Equal passes; anything else fails the step.

  6. Hand off

    Hand {{open_po_lines}} to the expediting steps that follow.

Viewport ledger

Every row once, in tables of any length

The VLI keeps a viewport ledger:

  • What is visible now.
  • What has already been captured, keyed by each row's content (PO number, line, release), not by its position on screen.
  • How far the last move scrolled, and how far to move next.

A row that appears at the bottom of one screen and the top of the next is recognized as one row. Each new screen is checked against what the ledger already holds before the agent moves again. When a scroll brings no new row into view, the grid has ended, and the count goes to step 5.

How the ledger reads the first screens
ScreenRows in viewAlready in the ledgerNew rows capturedLedger total
Screen 1 1 to 30 None 30 30
Screen 2 29 to 58 29 and 30, matched by PO number, line, and release 28 58
Screen 3 57 to 86 57 and 58 28 86
Last 2,371 to 2,400 Rows carried over from the screen before The remainder 2,400

The figures are illustrative; the scroll distance depends on the grid and the row height.

Containment

When the count does not match

A mismatch fails step 5, and the failure is contained. The partial list is discarded, so the expediting steps never see 2,391 lines dressed as 2,400. The retry reads the report again from the top. If the count still disagrees, the workflow's policy sends it to a person, with screenshots of the grid and of the printed total, because a report that disagrees with itself twice is a question for someone who knows the ERP.

The row key is chosen when the step is taught. PO number, line, and release identify an open line uniquely; a key that could repeat (supplier and due date, say) would merge two real lines into one, and the count check would catch it.

Shortcut

When the ERP offers an export

If the ERP can export the same report to a file, the agent can run the export and read the file instead; files are ordinary work for an agent that uses the desktop.

The ledger exists for grids that offer no export, exports that leave out a column the buyer needs, and portals that show data they will not let you download.

Your part

What you see and do

Most of the VLI's work is visible during the run and after the fact, not configured by hand.

  • While building, the Studio performs each action on your screen before you approve it.
  • While a run is live, the run console shows the agent's screen as it works.
  • After a run, every action opens to its before and after screenshots, with the target marked in red and the coordinates recorded.
  • When policy routes an unresolved target to a person, it arrives in the review queue with the screenshot that stopped it.

You do not position clicks or write selectors. The choices you make are the ones a manager would make: which workflows run on which machines, which applications to train, and what a step should do when it meets something it does not recognize.

Architecture

How the VLI connects to the rest of LaunchAI

The VLI is the agent's hands and eyes. Everything else decides what the hands do and keeps the record.

Governance and security

Fences at the screen

An agent that can operate any application needs fences that do not depend on the application.

  • It works with the rights of the signed-in session and no others.
  • Each workflow is fenced to the applications it declares. The Runner refuses to operate an application the workflow did not declare.
  • Grounding layers one to three run on the machine. Layer four sends only the screen region being read, and only to the model account your company configured.
  • Text on a screen is treated as data, never as an instruction, so a document that says “approve this invoice” is read, not obeyed.
  • A credential drawn from the vault is typed at the operating system's input layer at the moment of use and masked in every screenshot and recording.
  • No browser extension, injected script, or automation protocol is attached to the browser; the VLI uses the same mouse and keyboard a person does.

Metrics

What to measure

These numbers come from run history and the per-action record. Read them per application and per screen, because a problem on one ERP form hides inside a fleet average.

VLI metrics
MeasureWhat it tells you
Unresolved targets per screen Where look-alike controls or cramped layouts need a clearer step description
Verify failures per application Which applications changed, lag, or throw unexpected dialogs
Model calls per run Cost and latency; it falls as maps are trained and routes are reused
Seconds per step, minutes per run The speed side of the tradeoff, measured against the person's time for the same work
Rows read against report totals Whether long reads reconcile, run after run
Retraining events per vendor release The maintenance cost of each application

Tradeoffs

Speed and cost: slower per step, wider reach

LaunchAI gives up latency to win coverage.

Vision costs more per step than an API call. The agent looks, understands, acts, and verifies, every time. Three facts keep that cost in proportion:

  • Training removes most of the thinking from navigation.
  • The agent only has to outpace the person it frees.
  • Where a clean API exists for a step, the step can use it.

Through the screen, LaunchAI can do the work every narrower method does, plus the work none of them reach. Where a narrower method covers a subset faster or is already paid for, that is a legitimate reason to use it there.

Methods compared
MethodPer-step speedWhat it reachesWhen it can make sense
Vendor API or integration platform Fastest per call The objects and actions the vendor chose to expose High-volume transfers between systems with good APIs; a LaunchAI step can call the same API and use the screen for the rest of the job
RPA with selectors Fast on stable screens Applications with stable, exposed selectors Licenses and developers already in place; LaunchAI covers the same screens plus remote desktops, terminals, and custom controls, verifying every action
Cloud browser agents Fast in parallel Web pages Public-web collection at scale; LaunchAI covers the same pages from your own session, plus every desktop application
LaunchAI VLI Slower per step Any application a person can use Work that crosses systems, lacks APIs, or needs per-action evidence

Questions

Questions

How accurate is visual grounding?

LaunchAI promises no accuracy percentage, and a vendor that does is quoting a benchmark, not your screens. Perception fails at the edges: low contrast, tiny fonts, two nearly identical labels. The design answer is to refuse ambiguity, verify every action, and stop at the step that failed. Test it on your own screens during a pilot.

Can it operate an ERP delivered over Remote Desktop or Citrix?

Yes, as another window on the machine. No element tree crosses the connection, so the accessibility layer sits out and the other three layers do the work. Display scaling and color settings vary between remote sessions, so test on your own during a demo.

Can it read scanned or handwritten documents?

It reads them on screen, like any other window. Crooked scans and handwriting go down the model path, which returns typed values checked against a schema. A read below the confidence threshold goes to a person with the document beside the value, and is not posted until that person confirms it.

Who retrains an application after a vendor redesign?

The process owner or IT starts it, and LaunchAI's support engineers help on request; retraining after a redesign is part of what support covers beyond break-fix. The process owner then re-shows any step whose screen moved. Rules and tables need no change, because a redesign moves screens, not policy.

How many model calls does a run make?

It depends on the workflow. Exact-rule decisions make none. Targets resolved by the accessibility tree, OCR, or the local detector make none. Model calls come from reading messy documents, judging unfamiliar screens, and targets the local layers cannot settle. Training an application removes most of the calls that navigation would otherwise need.

Does the VLI work the same on Mac, Windows, and Linux?

The design is the same on all three: capture, interpret, decide, act, verify. What differs is the accessibility layer each operating system publishes, and how much of it a given application fills in. Confirm your operating system version before a pilot.

No prompts. No babysitting. No dumb questions.

Get a demo with our specialist team.

Get demo