GET BUG FOR REAL / MODEL STUDY 01
One scene.
Five models.
No alibis.
A controlled head-to-head test of element/reference-to-video generation using the same live-action cast, the same glass-divided geography, and the same multi-shot dramatic chunks.
CONTROL THE INPUT
Every supplied pixel belongs to the scene.
V2 removes the contradictory “style only” and illustrated references. One live-action master owns the world; every identity and chunk frame is derived from it and can be interpreted literally.
One literal scene system. Eight coherent live-action stills cost $1.848 to build and are reusable shared preparation—not charged again to individual model lanes.
Geography survives the medium change. David’s bed is beside a single large plate-glass divide. Youxin occupies the clinical side. They initially feel co-present; the reveal makes the separation legible.
The glass has a dramatic function. It separates, receives both hands, then becomes Youxin’s mirror when David’s side goes black.
THE UNIT OF COMPARISON
Multi-shot chunks, not isolated shots.
Each model must solve editorial thought: a beginning, an internal transition, and a resolved end image.
SYNCHRONIZED REVIEW
Five lanes. One playhead.
Every lane exposes its exact submitted prompt, ordered inputs, settings, task ID, and output asset.
FAIRNESS PROTOCOL
What stays locked.
- 01Same dramatic brief
Identical story beats, spoken intent, geography, casting, aspect ratio, and requested duration.
- 02Native syntax only
The content is locked; prompt syntax is adapted to each provider’s documented reference mechanism.
- 03First viable result
Record the first complete render before any retries. A retry becomes a separately costed attempt.
- 04Receipts, not estimates
Cost and render time come from actual generation receipts and timestamps, then remain editable for audit.
- 05Two distinct tracks
True multi-reference video models rank together. FLUX.3 and LTX 2.5 remain visible but are not misrepresented as equivalent R2V APIs.
DECISION SURFACE