Fun 1.0
10 quick tests. Reasoning, creativity, and simple traps. Best for a vibe check in any chat UI.
Vibe Coder's Life
Stop trusting benchmark slides. Run it yourself.
Practical, reproducible AI model comparisons. No blended intelligence score. No LLM judge. This Space is a browser explorer — official scoring is clone-and-run on GitHub with your own key.
10 quick tests. Reasoning, creativity, and simple traps. Best for a vibe check in any chat UI.
10 practical coding and debugging prompts. Mix of deterministic and exploratory. Best for vibe coders.
12 original JavaScript tasks. Hidden unit tests on GitHub. Best for a unit-test pass count — not an IQ score.
Want the interpretation, not just the numbers? Vibe Coder's Life publishes practical model comparisons and reruns after notable releases.
Scored items use deterministic asserts (or hidden unit tests on Score). Exploratory items are side-by-side judgment only. There is no blended “best model” score across Fun, Dev, and Score.
The GitHub README omits some Dev paste so a casual chat session does not quietly contaminate a CLI matrix. Prompts are still public in the repo and in this explorer. Official fair runs: fork the repo, set your key, run npm run eval:fun / eval:dev / eval:score.
Score hidden tests are not published here. Clone GitHub to grade generated functions.
VCL VibeBench is built by Viktor at Vibe Coder's Life (Budapest). GitHub is the canonical source. Hugging Face is the discovery layer. The newsletter owns the audience relationship.
Dataset: kondasviktor/vcl-vibebench (Fun, Dev, Score, results).