Deleting RAG from our web agent made it 2.3x faster

    TL;DR - We rewrote Skyvern from scratch using Pi as inspiration, making it 2.3x faster and score 90.5% on Odysseys benchmark. Try it out the via Open Source or Cloud.

    💡 Recap: What is Skyvern? It’s an open source framework that helps non-technical teams prompt to build browser automations. Skyvern is model agnostic, and can be run with closed models like Claude, or open ones like Deepseek. We have helped thousands of companies automate things like interacting with healthcare portals, fetching invoices or utility bills, filling out government forms, applying to jobs, and more.

    RAG? What are you talking about?

    For those unfamiliar with agents, RAG stands for Retrieval Augmented Generation. It’s a way to pass information to an LLM so they make smarter decisions. For example, before asking questions to an LLM about a document, you could parse that document and feed it into an LLM. You can read more about it here.

    Diagram of a RAG pipeline: knowledge base, vector database, retriever, augmentation, and LLM generation
    RAG-based LLM Frameworks work by fetching information from a Knowledge base, transforming it into context, and passing it to an LLM to make decisions on

    The first version of Skyvern was built using RAG. We would load up a website, screenshot and annotate it, condense it into a smaller list of elements, and send it to an LLM to take actions on.

    Diagram of the first Skyvern loop: draw bounding boxes, parse HTML, extract interactable elements, call an LLM, execute actions, repeat
    First version of Skyvern mimicked RAG based frameworks: It fetched information from the website, transformed it into actionable elements, and asked the LLM what action to take

    This worked great if you could identify every interactable element on a website with 100% accuracy. We thought we could eventually get there. We were wrong.

    Except.. the web is messy. Accurately identifying every element on a website proved to be a nightmare. What started off as a nice little utility to identify interactable elements essentially turned into building a new DOM interpreter.

    Parsing things like Shadow DOMs, iframes, select2 dropdowns, Kendo UI widgets, Canvas elements, …. this became an endless battle.

    The worst part was that agent performance did not improve with better LLMs. Instead, it was bottlenecked by this DOM parsing logic. If it failed to correctly identify an element, the entire agent run failed.

    This is a sign that we need to give the LLM more freedom.

    Give the LLM freedom to do what it wants

    RAG systems, if you aren’t careful, can assume that the pre-processing layer always happens ahead of any LLM calls, and don’t let the LLM decide when it is necessary. This was a good idea when LLMs weren’t good at following instructions, which meant that you couldn’t trust them to reliably pull the information when needed, but instruction following hasn’t been a big problem since the models released after Opus 4.5.

    We noticed 2 moves in the market that highlighted this change in thinking:

    1. Coding agents like Claude code were moving away from embedding style code search to grep style code search because the models were increasingly getting better at using primitive tools to fetch the context they needed
    2. Minimalist coding agents like Pi were becoming increasingly popular, with extensions of Pi like Prime-intellect hitting SOTA on the ARC-AGI-3 benchmark by tuning the harness and letting the agent generate and execute python code as needed

    So we asked ourselves: would the same approach work for browser agents? What if the agent started with the minimum context required, and gave it the ability to pull additional context as needed?

    In practice, it means giving the agent the following tools:

    1. Scan the HTML with model-generated javascript
    2. Take any rustwright-supported action on the website, for example Click, Type, Scroll, Switch tabs, etc
    3. Take a screenshot
    4. Mark the run as final (success, or failed with reason)
    Diagram of Skyvern 3.0: the LLM calls four tools (Read HTML, Take action, Take screenshot, Finish) to fill out a Geico insurance quote form

    The results were better than we expected

    Our hypothesis was that we would see an increase in performance, in exchange for added cost and a slower processing speed. The new architecture should result more calls to the LLM to get the same task done.

    To test this hypothesis, we split the test up into two branches:

    1. The leading public browser agent benchmark: Odysseys to test for accuracy
    2. A live A/B test across ~100,000 customer runs to test for speed / cost

    Odysseys benchmark results

    Skyvern 3.0 topped the Odysseys benchmark. It reported a 90.5% perfect rubric success rate, pretty much exactly as we expected. What surprised us though was that it was highly efficient in its runs, at only 65.4 average steps.

    Odysseys leaderboard, accuracy vs. average steps per task. Skyvern 3.0 (Opus 5) is on the Pareto frontier at 90.5% and 65 steps

    Production A/B Test results

    We ran the production A/B test on approx 100,000 runs and measured the impact

    Skyvern 2.0Skyvern 3.0Improvement
    Mean run time593 s262s2.3x
    Total cost per run$0.039$0.0302-22.5%

    The production A/B test results were surprising. The new architecture was.. faster and cheaper? We expected at least one of these dimensions to be a net negative

    Why was that the case? Digging into it, we found a few interesting things:

    1. Screenshot utilization dropped by 90% (27% —> 2.7% of all LLM calls). LLM calls with images go through a vision encoder to help the LLM “see” the image, which tend to be 1.7x slower than text-only encoders
    2. Token consumption increased by 52% (188K tokens / run —> 286K tokens per run) Average number of turns (ie LLM Calls) increased by 3.59x (6.4 —> 23), but tokens per turn decreased by 59% (29K tokens —> 12K tokens)
    3. Prompt Cache hit rates increased by 195% (27% —> 79.7%) because we stopped loading the screenshot or HTML on every turn, which meant a larger % of the prompt was cacheable

    Despite the increased token usage, cached tokens carry a 70% discount, which explains the net reduction in cost.

    So… are browser agents solved?

    The answer is still no. They can still be improved across all 3 dimensions (cost, speed, accuracy).

    Our team is now exploring the role of ”system one models” like Jev for browser agents. The current generation is still highly limited: 64K context window, no image processing, which causes unacceptable losses in accuracy, but we expect this to change over the next 6-12 months 🙂

    Give Skyvern a try!

    Run it via our Skyvern Open Source or Skyvern Cloud versions and let us know what you think!