Dataset Captioner
AI video captioning for training datasets
Auto-generate consistent, structured captions for video clips and still images - built for LoRA and diffusion-model training. Bulk captioning with format control and trigger words. Not subtitles.
Captioner - live output
Captioning is the bottleneck in dataset prep
Bulk captions, in the format your pipeline expects
Not transcription
FAQ
Is this the same as subtitles or transcription?
No. Captioning describes what happens in a clip as training-ready text - spoken dialogue is only quoted inside the caption when an audio-aware format (LTX-2.3, MiniMax H3) has audio enabled. If you need a standalone transcript or subtitle file, use the Transcription tool instead.
What caption formats does it produce?
Structured formats tuned for training pipelines - character-focused, image-to-video, and general description - as LTX-2.3 tags, MiniMax H3 cinematic prose with dialogue quoted inline, or plain Wan-style prose. You can steer the output with a custom instruction and inject a trigger word.
Does it caption still images too?
Yes. Image datasets get the same per-trainer treatment as clips: LTX-2.3 stills are tagged [VISUAL]/[TEXT], H3 and Wan-style stills stay natural prose, and in image-to-video mode a still is captioned as the conditioning frame - the starting state, with no invented motion.
What is it built for?
Preparing datasets for LoRA and diffusion-model training (for example WAN 2.2, LTX-2.3 and MiniMax H3), where hundreds of clips need consistent, well-structured captions. It pairs with the LoRA Dataset Builder.
What does it cost?
1 credit per 3 clips, or 1 credit per 10 images (each rounded up). See the pricing page for credit packs.
Can AI write an automatic video description?
Yes - that is exactly what the Captioner does. It watches each clip and writes a text description of what happens in it, in a structure you choose, across hundreds of clips with a consistent voice.
How do I write video descriptions for a LoRA dataset?
Upload your clips, pick a caption format (character-focused, image-to-video, or general), optionally set a trigger word for your concept, and export. The descriptions drop straight into common LoRA training pipelines such as WAN 2.2, LTX-2.3 and MiniMax H3.
Caption your dataset
Consistent, structured captions in bulk. 1 credit per 3 clips or 10 images.
Get started