/watch — Give Claude the Ability to Watch Any Video
A Claude Code skill that gives Claude the ability to watch any video. Paste a URL or local path, ask a question, and it fetches captions first, downloads only what's needed, extracts scene-aware frames, pulls a timestamped transcript (free captions first, Whisper API fallback), and reads every frame as an image — so Claude actually sees the video and hears the audio before answering.
Spec
/watch — Give Claude the Ability to Watch Any Video
You don't have a video input; this skill gives you one. A Python script gets captions first, optionally downloads the video, extracts frames as JPEGs (scene-aware, or fast keyframes at efficient detail), gets a timestamped transcript (native captions first, then Whisper API as fallback), and prints frame paths. You then Read each frame path to see the images and combine them with the transcript to answer the user.
When to use
- User pastes a video URL (YouTube, Vimeo, X, TikTok, Twitch clip, most yt-dlp-supported sites) and asks about it.
- User points at a local video file (
.mp4,.mov,.mkv,.webm, etc.) and asks about it. - User types
/watch [question].
Recommended limits
- Best accuracy: videos under 10 minutes. Frame coverage scales inversely with duration.
- Universal rate cap: 2 fps. The script never samples faster than 2 fps.
- The frame ceiling is set by the detail mode (
WATCH_DETAILin~/.config/watch/.env, or--detail):transcript→ no framesefficient→ up to 50 (keyframes)balanced(default) → up to 100 (scene-aware)token-burner→ uncapped (scene-aware; soft warning past 250 frames)--max-frames Noverrides whichever cap the mode would otherwise use.
- Full-video frame budget by duration:
- ≤30s → ~12-30 frames
- 30s-1min → ~40 frames
- 1-3min → ~60 frames
- 3-10min → ~80 frames
-
10min → up to the detail cap, sparsely spaced (warning printed)
- If the user hands you a long video, consider asking whether they want a specific section before burning tokens on a sparse scan.
How to invoke
Step 1 — parse the user input. Separate the video source (URL or path) from any question the user asked. Example: /watch https://youtu.be/abc what language is this in? → source = https://youtu.be/abc, question = what language is this in?.
Step 2 — run the watch script. Pass the source verbatim:
python3 "${SKILL_DIR}/scripts/watch.py" "SOURCE" --json
Step 3 — Read the output JSON. The script prints a JSON object with frames (array of {path, t}), transcript (timestamped text), and working_dir. Read each frame path with the Read tool — JPEGs render directly as images in Claude's context.
Step 4 — Answer the question. Combine what you see in the frames with what the transcript says. You saw the video. You heard the audio. Answer the way someone who watched the video would.
Detail modes
| Mode | Engine | Frames | Cap | Use when |
|------|--------|--------|-----|----------|
| transcript | none (captions) | 0 | — | Cheapest; text-only questions |
| efficient | keyframe | 50 | 50 | Speed tier (~0.5s extraction) |
| balanced | scene-change | 100 | 100 | Default; good visual fidelity |
| token-burner | scene-change | uncapped | — | Maximum fidelity; high token cost |
Frame deduplication
A dedup pass drops near-identical frames before they reach Claude (runs by default; --no-dedup turns it off):
- One
ffmpegcall scales each extracted JPEG to a 16×16 grayscale thumbnail. - For each frame, compute the mean absolute difference against the last frame that was kept.
- If that difference is ≤ threshold (
2.0), the frame is a near-duplicate and is dropped. - The frame-budget cap applies after dedup, so the budget is spent on distinct frames.
Comparing against the last kept frame (not the previous one) catches slow fades that never trip a frame-to-frame threshold.
How it works (end to end)
- You paste a video and a question. URL (anything yt-dlp supports) or a local path.
yt-dlpchecks captions first. Attranscriptdetail, captioned URLs return without downloading video.ffmpegextracts frames at the chosen detail. JPEGs are 512px wide by default.- The transcript comes from one of two places: native captions (free, instant) or Whisper API fallback (Groq's
whisper-large-v3preferred, OpenAI'swhisper-1as alternative). - Frames + transcript are handed to Claude. The script prints frame paths with
t=MM:SSmarkers and the transcript with timestamps. - Claude
Reads each frame in parallel and answers grounded in what's actually on screen and in the audio. - Cleanup: the script prints a working directory at the end. If not asking follow-ups, Claude removes it.
Requirements
- Python 3 (python3 on macOS/Linux;
pythonon Windows) - ffmpeg + ffprobe + yt-dlp (auto-installed on macOS via Homebrew; exact commands printed on Linux/Windows)
- Optional: Groq API key or OpenAI API key (for Whisper fallback when videos lack captions)
Install
Claude Code (auto-updates via marketplace):
/plugin marketplace add bradautomates/claude-video
/plugin install watch@claude-video
Codex, Cursor, Copilot, Gemini CLI, or any Agent Skills host:
npx skills add bradautomates/claude-video -g
What people use it for
- Analyze someone else's content ("what hook did they open with?")
- Diagnose a bug from a screen recording
- Summarize a video
- Cut the hype out of an update video
- Turn a playlist into notes

