Local-first Codex plugin that indexes long videos with OCR/ASR and answers with timestamped, SHA-verified visual evidence.
Turn long local videos into searchable, evidence-linked memory — including events that exist only in the picture.
English · 繁體中文 · Changelog / 變更紀錄 · Security · License
This is not “upload a video and ask for a summary.” It is a local media pipeline, a Codex-host visual inspection loop, and a versioned evidence store designed to answer: what happened, when did it happen, and which exact frame proves it?
Long-video assistants often reduce a video to its transcript. That loses the information people usually need most:
- a silent title card, diagram change, gesture, UI transition, object movement, or on-screen error; - repeated events that must be counted instead of vaguely summarized; - the exact frame and timestamp behind an answer; - continuity across tasks without re-watching or re-uploading the whole video; - a way to correct one fact without rebuilding unrelated evidence.
Codex Native Long Video Memory solves this by splitting the job into two explicit halves:
1. Local deterministic work: FFmpeg/ffprobe, scene detection, frame extraction, local OCR, local ASR, SQLite, hashing, indexing, resume, and invalidation. 2. Codex-host semantic work: the current Codex model inspects timestamped frames/contact sheets, writes structured observations, and reopens the evidence before verifying important claims.
The result is a reusable local memory graph with chapters, scenes, events, entities, actions, OCR/ASR spans, corrections, verification state, hashes, timestamps, and reopenable evidence.
From the project README.
Add the radar badge to your README — it shows your project was picked up by MCP Radar and links to this page:
[](https://mcp.liqiwa.com/s/NitcanKen--Codex-native-long-video-memory.html)
MCP server: turn narrated screen recordings into agent-ready data — local Whisper transcript, scene keyframes, OCR, wall-clock anchoring. Record your screen, talk — your AI agent files the bugs.
ozankasikci/global-agent-memoryLocal-first, project-aware memory MCP server for Claude Code, Codex, and other AI agents, with Obsidian and an owner dashboard.
KuaaMU/mcp-vision-bridgeMCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude Code, Codex, Kimi, opencode, PI.
sosoj92/jarvis-assistant-vocalAssistant vocal local en francais : Claude ou Ollama (offline), domotique Hue, OBS, agenda, navigateur, appels Twilio, serveur MCP. Python.
Riccardo8888/agent-linkAn end-to-end encrypted channel between coding agents, anywhere
zouyuanqing/vision-primitives-mcpVision-primitives MCP server for text-only LLMs: describe, locate (coordinates), OCR with bbox, annotate, crop, zoom, automated anomaly scanning, computer use — powered by Xiaomi MiMo V2.5. Single-fil
The top new MCP servers of the week, every Monday. No spam, unsubscribe anytime.