Pelican Benchmark
Submit your result

How this collection works

Pelican Benchmark is a collection of results from the pelican-riding-a-bicycle test, submitted by visitors and, for a starting set, added by the site. It shows what models have produced. It does not score, rank or measure them.

Where results come from

Anyone can submit a result without an account. A contributor uploads the source file and fills in:

  • the provider, model and thinking level, picked from a list (models not on the list can be typed in and are added after a manual review of the name);
  • the prompt they actually used — the standard prompt is offered, but never assumed;
  • the generation mode: single reply, several rounds of feedback, human-edited, or not sure;
  • optionally a generation date, a name, a source link and notes.

Results added by the site itself are credited to “admin”. Submitted results can be credited to any name except reserved ones such as “admin”.

What we do not verify

We cannot see how a result was produced, so the model, settings, prompt and generation mode are the contributor’s own account. A “single reply” label means the contributor reports one reply was used, not that no other attempts were made. Pages say so: “Model and generation details are provided by the contributor.”

The two tracks

  • Classic SVG uses the prompt “Generate an SVG of a pelican riding a bicycle”. The SVG is shown as an image, which stops any script inside it from running.
  • Animated HTML uses the prompt “Create a self-contained HTML file with a 2D SVG animation of a pelican riding a bicycle”. The file runs unchanged in a sandbox on a separate domain, with network requests blocked and outside resources limited to a short list of public CDNs.

We store and show the source exactly as submitted. We never rewrite a contributor’s code to make it look better or to make it safe; isolation does that job.

What happens after you submit

  1. Automatic checks: file size, text encoding, valid structure, complexity limits and exact duplicates. A file that fails is not stored, and the form says what to fix.
  2. Content screening: a screenshot of the result goes to a vision model, which checks that it shows a pelican and a bicycle and has no obvious sexual, violent, hateful or promotional content. If the check passes, the result is published; if it fails or cannot run, a person looks at it first.
  3. Ongoing checks: published results are re-rendered and re-screened about once a day. A result that no longer shows a pelican on a bicycle — for example because something it loads has changed — or that turns out to be unsuitable is taken down until a person has looked at it.

This screening decides whether a result may be shown. It does not judge how good the drawing is: a clumsy or odd-looking pelican is still a valid result.

What this collection cannot tell you

  • Which model is best. There are no scores, and results were made under different conditions.
  • How consistent a model is. One result is one run; the collection is not a controlled experiment.
  • Whether a result is exactly as described. See “What we do not verify” above.
  • What the models do today. Hosted models change over time, and a result reflects the day it was made.

Corrections and takedowns

To report a result, fix a wrong model name, or ask for your work to be taken down, email contact@pelicanbenchmark.com with the link. See the Terms of Use for details.

Community-submitted results for the Pelican test proposed and popularized by Simon Willison. Not affiliated with, and not endorsed by, any model provider.

Contact / report: contact@pelicanbenchmark.com

How this collection works · Terms of Use · Privacy Policy