Video → Prompt · Local AI

Video to Prompt

Reverse-engineer any video into a production-ready prompt. The vision AI describes every frame, and the language model writes the prompt — entirely on your device.

Vision Engine

✨ Synthesize prompt — Qwen2.5 · Balanced
Choose video
Analyze frames
Synthesize prompt
🎬
Drag a video here, or click to select
🔒 Analyzed 100% on your device — the video never leaves this browser. Clips under 30s work best.
📝 Have an idea instead? Write a prompt from a description →
🎬
🎬
Your generated prompt will appear here.

Why local generation

🔒

Total privacy

Your ideas and prompts never leave your device. No uploads, no logs, no one watching. True for every single generation.

Free & unlimited

The model runs on YOUR hardware, so there's no server bill for anyone. No API keys, no quotas, no paywalls — generate 100 prompts a day.

📡

Works offline

After the one-time download, the model lives in your browser cache. Turn off Wi-Fi and it still generates — airplane mode, no problem.

💎

Your data is yours

No account, no email, no analytics on your prompts. The model has no idea who you are — and neither does anyone else.

Frequently Asked Questions

Q: Does my video leave my device?

No. Frames are extracted and described entirely in your browser. Nothing is uploaded — there is no server that could receive it. Load the page, switch on airplane mode, and it still works.

Q: Does “video to prompt” recover the original prompt?

No — it reconstructs a prompt from what is visible: subject, actions, camera, lighting, style. The original prompt cannot be recovered from footage, so the result is an honest interpretation.

Q: Do I need to download another model?

Yes: a compact vision model describes the frames — Standard (~560MB) or High Precision (~1.5GB) — and the text model you already use writes the final prompt.

Q: What video formats and sizes are supported?

Any format your browser can play (MP4, WebM, MOV and more) — and no size limit, because the video never leaves your device. The page extracts frames locally, so a longer video simply takes a little longer to analyze.

Q: How accurate is the result?

The AI reconstructs a prompt from what the frames visibly show: subject, actions, camera, lighting, style. It cannot recover the original prompt used to create the video — but it produces an honest, production-ready interpretation you can refine.