Reverse-engineer any video into a production-ready prompt. The vision AI describes every frame, and the language model writes the prompt — entirely on your device.
Your ideas and prompts never leave your device. No uploads, no logs, no one watching. True for every single generation.
The model runs on YOUR hardware, so there's no server bill for anyone. No API keys, no quotas, no paywalls — generate 100 prompts a day.
After the one-time download, the model lives in your browser cache. Turn off Wi-Fi and it still generates — airplane mode, no problem.
No account, no email, no analytics on your prompts. The model has no idea who you are — and neither does anyone else.
No. Frames are extracted and described entirely in your browser. Nothing is uploaded — there is no server that could receive it. Load the page, switch on airplane mode, and it still works.
No — it reconstructs a prompt from what is visible: subject, actions, camera, lighting, style. The original prompt cannot be recovered from footage, so the result is an honest interpretation.
Yes: a compact vision model describes the frames — Standard (~560MB) or High Precision (~1.5GB) — and the text model you already use writes the final prompt.
Any format your browser can play (MP4, WebM, MOV and more) — and no size limit, because the video never leaves your device. The page extracts frames locally, so a longer video simply takes a little longer to analyze.
The AI reconstructs a prompt from what the frames visibly show: subject, actions, camera, lighting, style. It cannot recover the original prompt used to create the video — but it produces an honest, production-ready interpretation you can refine.