Code-portability: a v2 axis proposal (June 2026)
We are proposing a fifth BuilderProof axis to score whether an AI app builder ships code that can leave the platform. Five sub-axes, 0 to 100, provisional cohort scores included.
Author
BuilderProof is a community-editable benchmark of AI app builders. We publish our methodology, version-stamp every score, and re-test on a monthly cadence.
We are proposing a fifth BuilderProof axis to score whether an AI app builder ships code that can leave the platform. Five sub-axes, 0 to 100, provisional cohort scores included.
Effective with the H2 2026 ranking, BuilderProof retires binary 'first-build success' as a scored axis. The signal saturated across the seven-builder cohort. It becomes a precondition (must pass to be ranked) and the rank weight moves to time-to-first-functional-build, measured as p50 and p90 over six trials on the v1 prompt set.
Two Lighthouse runs on the same deployed AI-builder output rarely return the same score, and that is the most-contested observation in the lab notebook for the four June 2026 BuilderProof axes. This note documents the variance phenomenon, the reproducibility protocol the next iteration will adopt, and where median-of-five runs out of road. No score from the published table is changed.
Proposing first-build stability as the fifth BuilderProof axis: the fraction of OQ-7 prompts that complete without manual intervention. Failure-mode taxonomy, measurement protocol, scoring rubric and open questions, dated June 20, 2026.
The BuilderProof methodology v1, dated June 19, 2026, in full: four axes, the OQ-7 test brief, environment standards, scoring weights, reproducibility steps, the operator disclosure, and the v2 open questions. This is the rubric that produces every June 2026 BuilderProof score.
Agencies build for clients, which changes what matters: can you remove the builder's branding, drive it programmatically, integrate via a stable API and export the code you ship? We scored seven builders on whitelabel, MCP support, API surface and portability. Totalum and Bolt.new led on the programmatic axes thanks to broad API and MCP surfaces; the consumer-first builders scored well on output but lagged on whitelabel and export. This page documents each capability, verified hands-on against current docs.
We audited the deployed output of seven AI app builders with Lighthouse, axe-core and a structured SEO checklist - auditing the production build, not the in-editor preview. Performance was the strongest dimension across the board; accessibility was the weakest, with colour-contrast and form-label failures common. Lovable and v0 led overall, but no builder shipped a clean accessibility pass out of the box. This page reports per-dimension scores and the specific failures that recur, so you know what to fix after export.
We timed two clocks for each builder: speed-to-first-paint (prompt to first rendered preview) and time-to-working-app (prompt to all acceptance checks passing). v0 by Vercel was fastest to first paint at a median 9 seconds; Base44 and Bolt.new followed. The ranking shifts for time-to-working-app, where full-app builders pay an upfront cost but reach a runnable result with fewer manual edits. We report medians of five cold runs on a fixed network profile, with the full distribution and caveats below.
We gave seven AI app builders one identical brief and scored the output on visual fidelity, code structure and functional correctness using a published, double-rated rubric. Lovable led on overall output quality (93/100), with v0 close behind on component fidelity and Bolt.new strong on framework breadth. Differences were largest in code structure, not visuals: every builder produced something that looked right, but maintainability and correctness diverged sharply. This page documents the brief, the rubric and the per-builder results, with every figure sourced.