Vaultwealth web and iOS Simulator, seeded local fixtures, one Mac, bounded runs, and an independent correctness oracle.
Last experiment
20 September 2026
Current decision
Keep Playwright + Maestro
Bottom line
Keep Playwright and Maestro as regression baselines. Use adaptive tools for exploration with an independent verifier.
Promising results remain screening evidence. None completed the planned 20 warm and three cold qualification set across every journey.
Coverage in numbers
Tools mapped79
Tools executed21
Measured journey arms26
Verified journey attempts118/129
Seeded defect types5
Accepted paid API spend$0.00480
Executed means benchmarked or bounded-screened locally. The other 54 catalogue entries retain source-backed fit boundaries; they are not assigned synthetic scores.
Measured journey matrix 26 arms
Median and observed p95 include the timing basis shown under each mode. At these sample sizes, observed p95 is usually the maximum. A dash means the report did not establish that number.
Platform
Journey
Mode
Median
Observed p95
Verified
Model / cost
Seeded fault
Web
Login
Playwright savedverified passes
1.178 s
1.400 s
5/5
0 calls · $0
clipped control detected
Web
Login
agent-browser directverified passes
2.714 s
3.846 s
5/5
0 calls · $0
clipped control detected
Web
Login
agent-browser batchverified passes
1.323 s
1.734 s
5/5
0 calls · $0
clipped control detected
Web
Edit + reload
Playwright savedverified passes
1.898 s
1.950 s
5/5
0 calls · $0
3/3 write faults detected
Web
Edit + reload
agent-browser directverified passes
1.762 s
1.795 s
5/5
0 calls · $0
3/3 write faults detected
Web
Edit + reload
agent-browser batchno verified completion
—
—
0/5
0 calls · $0
not eligible after clean failure
Web
Search
Playwright savedverified passes
2.768 s
2.796 s
5/5
0 calls · $0
stale result detected
Web
Search
agent-browser directverified passes
2.826 s
2.833 s
5/5
0 calls · $0
stale result detected
Web
Search
agent-browser batchverified passes
2.649 s
2.836 s
5/5
0 calls · $0
stale result detected
Web
Search
Puppeteer Coreverified passes
2.081 s
2.146 s
5/5
0 calls · $0
stale result detected
Web
Search
WebdriverIOverified passes
2.958 s
3.087 s
5/5
0 calls · $0
stale result detected
Web
Search
Selenium WebDriververified passes
3.633 s
3.651 s
5/5
0 calls · $0
stale result detected
Web
Search
Nightwatchverified passes
5.425 s
5.700 s
5/5
0 calls · $0
stale result detected
Web
Search
Taikoverified passes
23.850 s
23.921 s
5/5
0 calls · $0
stale result detected
Web
Search
Stagehand cachedverified replays
2.201 s
2.233 s
4/4
0 replay calls · $0
stale result detected
Web
Search
Stagehand uncachedverified passes
26.951 s
31.401 s
5/5
3 local calls · $0
stale result detected
Web
Search
Jev + Bonsaiverified passes
11.101 s
11.391 s
5/5
7 paid selector + 3 local calls · ~$0.00063/run
stale result detected
Web
Search
Browser Use + Bonsaisuccessful verified passes
90.730 s
93.623 s
4/5
8 local calls on successes · $0
stale result detected
Web
Login
Playwright CLIverified passes; p95 unavailable
3.235 s
—
5/5
0 calls · $0
no fault sample
Web
Edit + reload
Playwright CLIverified passes; p95 unavailable
3.076 s
—
5/5
0 calls · $0
no fault sample
Web
Adaptive login
Codex stepwiseverified passes
155.971 s
225.686 s
5/5
918,042 median input tokens
clipped control detected separately
Web
Adaptive login
Codex batchedall attempts; 2/5 functional
123.684 s
139.657 s
0/5
398,585 median input tokens
clipped calibration noticed; unqualified
iOS
Login
Maestro saved flowverified warm passes
32.964 s
34.775 s
5/5
0 calls · $0
clipped control detected
iOS
Login
AXe batchverified warm passes
12.612 s
12.892 s
5/5
0 calls · $0
clipped control detected
iOS
Search
Maestro saved flowverified warm passes
34.673 s
41.146 s
5/5
0 calls · $0
no fault sample
iOS
Search
AXe batch, one-char typingverified warm passes
15.279 s
15.317 s
5/5
0 calls · $0
no fault sample
Other measured probes 7
These numbers are useful, but they are not comparable end-to-end application journeys.
gh repo clone sarthakagrawal927/agent-testing
cd agent-testing
npm test
# Validate the real-product adapter
npm run validate:vaultwealth
# Read exact setup and replay commands
open adapters/vaultwealth/runtime/README.md
The source repository is public. Product adapters must still use local or disposable seeded targets only.
Experiment 01 · complete screening
Saved web workflows
Arms: Playwright · agent-browser direct · agent-browser batch
Playwright remained the baseline. agent-browser batching improved its own login time, but no candidate met the 2x adoption threshold.
Login: Playwright 1.178 s; direct 2.714 s; batch 1.323 s medians; all 5/5.
Edit: Playwright 1.898 s and direct 1.762 s medians; batch failed 0/5.
Search: Playwright 2.768 s; direct 2.826 s; batch 2.649 s medians; all 5/5.
Stagehand cache was the most promising replay result, but missed the 2x gate. Browser Use was slow and 4/5. Chrome DevTools MCP was useful for diagnostics only.
Stagehand uncached 26.951 s median; 5/5. Cached 2.201 s median; 4/4 after 25.544 s authoring.
Browser Use 90.730 s median among successes; 4/5.
Chrome DevTools MCP diagnostic bundle 0.826 s median; Lighthouse snapshot 1.243 s.
XcodeBuildMCP replaced the formatted cash field in 5/5 probes but no save/relaunch journey was qualified. XCUITest and Detox were setup-blocked by checkout structure.
XcodeBuildMCP action median 3.834 s; semantic observation median 1.343 s.
Maestro/debug setup still took about 35 s.
No claim is made that XCUITest or Detox failed as tools.
All five candidates completed five independently verified clean searches and detected the stale-search fault. Puppeteer was fastest, but did not meet the 2x adoption gate against Playwright.
Both MCP servers started and found the dedicated simulator. ios-simulator-mcp returned the accessibility tree through isolated idb; Mobile MCP listed 27 elements and its temporary device agent was removed afterward.
ios-simulator-mcp 2.1.0 exposed 17 tools and returned a 6,940-character accessibility description.
Mobile MCP 1.0.4 exposed 32 tools and listed 27 elements with device agent 0.0.26.
These are readiness probes, not saved-flow latency or defect-detection results.
These pins make the historical result reproducible. They are not recommendations to avoid newer versions.
What would justify rerunning
Rerun when the application journey changes, a candidate has a material new release, the browser or simulator changes, or a five-run screen beats the current reliability and verified-feedback result. Do not rerun the whole catalogue merely because another tool exists.