Dataset Captioner

AI video captioning for training datasets

Auto-generate consistent, structured captions for video clips and still images - built for LoRA and diffusion-model training. Bulk captioning with format control and trigger words. Not subtitles.

Captioner - live output

Captioning is the bottleneck in dataset prep

A good training run needs hundreds of clips and images, each described in a consistent format. Doing that by hand is slow, and inconsistent captions quietly degrade the model you are training.

Bulk captions, in the format your pipeline expects

Point it at your clips and images and it writes structured, training-ready captions - character-focused, image-to-video, or general - with a consistent voice across the whole set. Stills are captioned by the same per-trainer rules (LTX-2.3 stills get [VISUAL]/[TEXT] tags, prose formats stay prose). Add a custom instruction to steer emphasis, and a trigger word for your concept.

Not transcription

This describes what is in the clip for a model to learn from - it does not transcribe spoken words. If you need speech turned into text, use Transcription instead.

FAQ

Is this the same as subtitles or transcription?

No. Captioning describes what happens in a clip as training-ready text - spoken dialogue is only quoted inside the caption when an audio-aware format (LTX-2.3, MiniMax H3) has audio enabled. If you need a standalone transcript or subtitle file, use the Transcription tool instead.

What caption formats does it produce?

Structured formats tuned for training pipelines - character-focused, image-to-video, and general description - as LTX-2.3 tags, MiniMax H3 cinematic prose with dialogue quoted inline, or plain Wan-style prose. You can steer the output with a custom instruction and inject a trigger word.

Does it caption still images too?

Yes. Image datasets get the same per-trainer treatment as clips: LTX-2.3 stills are tagged [VISUAL]/[TEXT], H3 and Wan-style stills stay natural prose, and in image-to-video mode a still is captioned as the conditioning frame - the starting state, with no invented motion.

What is it built for?

Preparing datasets for LoRA and diffusion-model training (for example WAN 2.2, LTX-2.3 and MiniMax H3), where hundreds of clips need consistent, well-structured captions. It pairs with the LoRA Dataset Builder.

What does it cost?

1 credit per 3 clips, or 1 credit per 10 images (each rounded up). See the pricing page for credit packs.

Can AI write an automatic video description?

Yes - that is exactly what the Captioner does. It watches each clip and writes a text description of what happens in it, in a structure you choose, across hundreds of clips with a consistent voice.

How do I write video descriptions for a LoRA dataset?

Upload your clips, pick a caption format (character-focused, image-to-video, or general), optionally set a trigger word for your concept, and export. The descriptions drop straight into common LoRA training pipelines such as WAN 2.2, LTX-2.3 and MiniMax H3.

Caption your dataset

Consistent, structured captions in bulk. 1 credit per 3 clips or 10 images.

Get started