Completed experiment · tool catalogue
79 tools, each with a verdict.
A dated map of practical browser-agent and iOS-agent testing choices. Every row has either local execution evidence or a concrete fit boundary; there are no placeholder “not run” rows.
- Catalogue reviewed
- 20 September 2026
- Last experiment
- 20 September 2026
Coverage is exhaustive within the seven named categories as of the review date, using maintained official projects and materially relevant reference tools. It is a dated map, not a claim that every testing product in existence belongs here.
Evidence key
- benchmarked
- Repeated locally against the shared Vaultwealth journeys or faults.
- screened
- A bounded local probe was completed, but the tool was not fully qualified.
- setup-blocked
- A dated setup or run was attempted, but a named prerequisite or bounded failure prevented a fair result.
- source reviewed
- Official sources were reviewed and a concrete fit boundary was recorded; no runtime ranking is claimed.
Saved web automation 12
| Tool | Evidence | Cost | Version or boundary | What we know |
|---|---|---|---|---|
| Playwright | benchmarked | Free / local | 1.62.1 | Qualified web baseline; fixed assertions and visual comparisons. |
| agent-browser | benchmarked | Free / local | 0.38.1 | Batch reduced its own login overhead; batch edit failed 0/5. |
| Playwright CLI | screened | Free / local | 0.1.21 | Login and edit passed 5/5 but were slower than the saved baseline. |
| Puppeteer Core | benchmarked | Free / local | 24.43.1 | Search passed 5/5 at 2.081 s workflow median and detected stale search; fastest added screen, but not a 2x win. |
| Selenium WebDriver | benchmarked | Free / local | 4.49.0 / ChromeDriver 151.0.5 | Search passed 5/5 at 3.633 s workflow median and detected stale search. |
| Cypress | setup-blocked | Free core / paid cloud | 16.1.0 package | Package pinned; the isolated experiment had no Cypress application binary. No tool-quality claim. |
| WebdriverIO | benchmarked | Free / local | 9.31.9 | Search passed 5/5 at 2.958 s workflow median and detected stale search. |
| Nightwatch | benchmarked | Free / local | 3.16.0 | Search passed 5/5 at 5.425 s workflow median and detected stale search. |
| TestCafe | setup-blocked | Free / local | 3.7.6 | Proxy run did not return a test result before the bounded 60 second timeout on this service-worker fixture. |
| Taiko | benchmarked | Free / local | 1.5.0 | Search passed 5/5 at 23.850 s workflow median and detected stale search; isolated launches were required. |
| CodeceptJS | source reviewed | Free / local | 4.1.0 package inspected | Installed and pinned, but it delegates to Playwright/WebDriver/Puppeteer; no separate driver speed claim after those engines were measured directly. |
| Testplane | source reviewed | Free / local | WebDriver runner alternative | Applicable WebDriver test runner, but it does not change the underlying driver result measured with Selenium and WebdriverIO. |
Adaptive browser agents 22
| Tool | Evidence | Cost | Version or boundary | What we know |
|---|---|---|---|---|
| Jev Ultrafast | benchmarked | Cheap paid selector + local text model | 0.1.0 @ 1231850a0bf1a0c0341fe408ef1668dbbfdfac46 | Authenticated search passed 5/5; password fields are excluded upstream. |
| Stagehand | screened | Free code / model-dependent | 3.4.0 | Cached replay reached 2.201 s median, 4/4; authoring took 25.544 s. |
| Browser Use | screened | Free code / model-dependent | 0.13.10 | Worked with local Bonsai but passed 4/5 at 90.730 s among successes. |
| Chrome DevTools MCP | screened | Free / local | 1.9.0 | Fast diagnostic companion; not a journey executor. |
| Playwright MCP | source reviewed | Free / local | 0.0.82 package inspected | Pinned locally; an agent transport over Playwright, not a saved-executor speed replacement. |
| Codex coding agent | benchmarked | Subscription usage | 0.155.1 / gpt-6-astra | Stepwise login passed 5/5; the batched instruction arm qualified 0/5. |
| Skyvern | source reviewed | Open source / model-dependent | Model or agent orchestration layer | Computer-vision and LLM browser workflow platform. |
| LaVague | source reviewed | Open source / model-dependent | Model or agent orchestration layer | Web-agent framework with browser actions and retrieval. |
| AgentQL | source reviewed | Free tier / paid service | Model or agent orchestration layer | Natural-language element and page-data queries. |
| Magnitude | source reviewed | Open source / vision-model dependent | Model or agent orchestration layer | Vision-first browser agent; text-only Bonsai is not a fair model. |
| Browser MCP | source reviewed | Free / local browser | Personal-profile boundary | Controls an existing browser through an extension; excluded because the experiment forbids personal browser profiles. |
| OpenAI computer use | source reviewed | Paid API | Model or agent orchestration layer | General screenshot-and-action computer-use model, not a test oracle. |
| Anthropic computer use | source reviewed | Paid API | Model or agent orchestration layer | General computer-use tool requiring an external execution environment. |
| Browserable | source reviewed | Open source / provider-dependent | Agent framework | Self-hostable agent framework; model/orchestration choice rather than a fixed local driver benchmark. |
| Drisp Browser MCP | source reviewed | Free / local | MCP interaction layer | Local browser-control MCP with snapshots and actions; use as an agent interface, not a correctness oracle. |
| Midscene.js | source reviewed | Open source / model-dependent | Vision model required | AI-driven web and Android automation; the selected Bonsai text model is not a fair vision backend. |
| TestZeus Hercules | source reviewed | Open source / model-dependent | Agent framework | Agentic browser testing framework; requires model configuration and independent outcome verification. |
| Shortest | source reviewed | Open source / model-dependent | Model-dependent test runner | Natural-language Playwright tests; model reasoning remains the dominant variable for a fair comparison. |
| UI-TARS Desktop | source reviewed | Open source / local or hosted model | Desktop GUI agent | General desktop GUI agent with browser control; broader than the local web-driver question and still needs an external oracle. |
| Agent TARS | source reviewed | Open source / model-dependent | General multimodal agent | General multimodal agent framework, not a deterministic test executor. |
| Harness | source reviewed | Free / model-dependent | Experimental agent harness | Experimental browser and mobile agent harness; useful for exploration research, without a qualified Vaultwealth run. |
| Playwright Test Agents | source reviewed | Free code / model-dependent | Authoring and healing layer | Planner, generator and healer agents author Playwright tests; repeat execution remains Playwright. |
Agent evaluation suites 5
| Tool | Evidence | Cost | Version or boundary | What we know |
|---|---|---|---|---|
| BrowserGym | source reviewed | Free / model-dependent | Evaluation suite, not driver | Unifies reproducible web-agent environments; it evaluates agents rather than replacing product test execution. |
| AgentLab | source reviewed | Free / model-dependent | Research harness | Reproducible web-agent experimentation on BrowserGym; not an application regression runner. |
| WebArena | source reviewed | Free / local infrastructure | Benchmark environment | Self-hosted realistic web tasks for agent research; not a substitute for product-specific assertions. |
| VisualWebArena | source reviewed | Free / local infrastructure | Multimodal benchmark | Visual web-agent benchmark; measures general capability rather than Vaultwealth defect detection. |
| WorkArena | source reviewed | Free code / ServiceNow instance | Enterprise benchmark | ServiceNow task benchmark with an application-specific infrastructure prerequisite. |
Remote browser infrastructure 5
| Tool | Evidence | Cost | Version or boundary | What we know |
|---|---|---|---|---|
| Browserbase | source reviewed | Paid service | Remote infrastructure outside local-only scope | Managed remote browsers; Stagehand integration is first-party. |
| Browserless | source reviewed | Self-hosted or paid | Remote infrastructure outside local-only scope | Remote Playwright/Puppeteer browser infrastructure. |
| Steel | source reviewed | Open source or paid | Remote infrastructure outside local-only scope | Open-source browser API for agents and apps. |
| Hyperbrowser | source reviewed | Paid service | Remote infrastructure outside local-only scope | Managed browsers and agent infrastructure. |
| Anchor Browser | source reviewed | Paid service | Remote infrastructure outside local-only scope | Cloud browsers for agents and automation. |
Visual verification 11
| Tool | Evidence | Cost | Version or boundary | What we know |
|---|---|---|---|---|
| Vitest browser visual assertions | source reviewed | Free / local | Visual oracle or diff layer | Screenshot matching inside Vitest browser mode. |
| Percy | source reviewed | Paid service | Visual oracle or diff layer | Managed cross-browser visual review. |
| Applitools Eyes | source reviewed | Paid service | Visual oracle or diff layer | Managed visual-AI comparison and review. |
| Chromatic | source reviewed | Free tier / paid service | Visual oracle or diff layer | Storybook-oriented component visual testing. |
| Argos | source reviewed | Open source / paid service | Visual oracle or diff layer | Visual and accessibility-tree diffs for pull requests. |
| Lost Pixel | source reviewed | Open source | Visual oracle or diff layer | Sunsetting; useful comparison history, not a fresh adoption recommendation. |
| BackstopJS | source reviewed | Free / local | Visual oracle or diff layer | Self-hosted browser screenshot regression testing. |
| Loki | source reviewed | Free / local | Visual oracle or diff layer | Storybook screenshot regression testing; verify maintenance before adopting. |
| reg-suit | source reviewed | Free / local or object storage | Screenshot diff pipeline | Screenshot regression workflow; requires a capture source and does not explore the application. |
| TestivAI OSS | source reviewed | Open source / model-dependent | AI visual test framework | Open-source visual testing framework; model and review configuration are additional variables. |
| Odiff | source reviewed | Free / local | Image diff primitive | Fast image comparison primitive, not a journey runner or visual-understanding agent. |
iOS and mobile 20
| Tool | Evidence | Cost | Version or boundary | What we know |
|---|---|---|---|---|
| Maestro | benchmarked | Free CLI / optional cloud | 2.6.1 / Java 21 | Retained iOS saved-flow baseline. |
| AXe | benchmarked | Free / local | 1.8.0 | Fast warm runs, but cold accessibility and text-entry reliability blocked adoption. |
| XcodeBuildMCP | screened | Free / local | 2.7.0 | Useful installed-app probe; full save/relaunch journey was not qualified. |
| XCUITest | setup-blocked | Included with Xcode | Native interaction or simulator layer | Vaultwealth checkout lacked an app-owned Xcode UI-test target. |
| Detox | setup-blocked | Free / local | Native interaction or simulator layer | Vaultwealth checkout lacked Detox configuration and an app-owned target. |
| Appium | source reviewed | Free / local | Native interaction or simulator layer | Cross-platform W3C WebDriver automation; iOS uses the XCUITest driver. |
| Appium MCP | source reviewed | Free / local | Native interaction or simulator layer | Official MCP layer for Appium mobile automation. |
| Patrol | source reviewed | Free core / optional cloud | Native interaction or simulator layer | Flutter-focused native integration testing. |
| EarlGrey 2 | source reviewed | Free / local | Native interaction or simulator layer | Google iOS UI automation integrated with XCUITest. |
| idb | screened | Free / local | fb-idb 1.1.7 | Isolated runtime found the simulator and returned its accessibility tree through ios-simulator-mcp; the broken user-level install was not modified. |
| idb-mcp | source reviewed | Free / local | Native interaction or simulator layer | MCP wrapper around idb for iOS Simulator control. |
| AppleSimulatorUtils | source reviewed | Free / local | Simulator utility, not journey runner | Simulator state and permission utility; upstream deprecates overlapping operations in favor of simctl. |
| SimPilot | source reviewed | Free / local | macOS simulator agent | Agent-oriented iOS Simulator automation; requires a separate product-state verifier for defect claims. |
| Wand | source reviewed | Free / local | Private-framework risk | Fast simulator automation built around private Apple frameworks; unsuitable as the default long-lived baseline without accepting that maintenance risk. |
| ios-simulator-mcp | screened | Free / local | 2.1.0 | Started with 17 tools, found the dedicated simulator and returned a 6,940-character accessibility description using isolated idb. |
| ios-simulator-mcp by Yael Gilboa | source reviewed | Free / local | Alternative MCP fork | Alternative simulator MCP focused on screenshots, taps and accessibility; overlaps XcodeBuildMCP and AXe. |
| iosef | source reviewed | Free / local | Accessibility automation CLI | Accessibility-first iOS Simulator control; a candidate for a future native screen, but no claim is made beyond interface fit. |
| Mobile MCP | screened | Free / local | 1.0.4 / device agent 0.0.26 | Started with 32 tools, found the simulator and listed 27 elements; its temporary device agent was removed after the probe. |
| AutoMobile | source reviewed | Free / model-dependent | Mobile exploration agent | Agentic mobile exploration project; model cost and nondeterminism require separate accounting. |
| Mobilewright | source reviewed | Free / local | 0.0.57 package inspected | Installed transitively with Mobile MCP; early project, so baseline maturity and API stability remain boundaries. |
Local model helpers 4
| Tool | Evidence | Cost | Version or boundary | What we know |
|---|---|---|---|---|
| Ternary Bonsai 2 27B PTQ1_0 | screened | Free / local | revision 6ed5e12bf84b7a63069882c91dd9e9218647d17b | Text helper passed 25/25; no-thinking median 1.082 s at 7.305 GB sampled peak RSS. |
| Prism llama.cpp | screened | Free / local | 0.2.0-dev build 10709 @ 9a9394a895b96003ca842a6041cb28ac49a108f7 | Pinned runtime used to serve Bonsai locally. |
| Ollama | source reviewed | Free / local | Model runtime, not UI executor | Convenient local model runtime; not benchmarked in this experiment. |
| MLX LM | source reviewed | Free / local | Model runtime, not UI executor | Apple-silicon local model runtime; not benchmarked here. |