Video transcripts for RAG and LLM pipelines
Turn YouTube and TikTok videos into chunks an LLM can search and cite: fetch timed segments, chunk by time, keep timestamp links.
Updated
Short answer: fetch the transcript as JSON, not plain text. The segments array carries a start and end time for every caption line, so you can chunk by time and store the start as metadata. Then every answer your model gives can link to the exact moment in the video.
Why timed segments beat plain text#
- Citations: a chunk that knows it starts at 754 seconds becomes a link to that moment.
- Natural boundaries: caption lines end at pauses, so chunks rarely cut a sentence in half.
- Provenance:
source_origintells you whether the text came from the creator, the platform's automatic captions or AI transcription, andlanguagegives its language.
Fetch and chunk#
This fetches one transcript and groups segments into windows of about a minute. The transcript JSON shape is documented in the API reference.
import requestsAPI = "https://www.transcriptdock.com"H = {"Authorization": "Bearer td_live_YOUR_KEY"}def transcript(url: str) -> dict:job = requests.post(f"{API}/v1/jobs", headers={**H, "Idempotency-Key": f"rag-{url}"[:128]},json={"source": {"url": url}, "mode": "captions_only"}).json()while job["status"] not in ("succeeded", "failed", "cancelled"):job = requests.get(f"{API}/v1/jobs/{job['id']}", params={"wait": 25}, headers=H).json()if job["status"] != "succeeded":raise RuntimeError(job["error"]["code"])return requests.get(f"{API}/v1/transcripts/{job['result_id']}", headers=H).json()def chunks(t: dict, seconds: float = 60.0):"""Group caption segments into windows of about `seconds`, never splitting a segment."""url = t["source"]["canonical_url"]buf, start = [], Nonefor seg in t["segments"]:if start is None:start = seg["start"]buf.append(seg["text"])if seg["end"] - start >= seconds:yield {"text": " ".join(buf), "start": start, "url": f"{url}&t={int(start)}s" if "youtube" in url else url}buf, start = [], Noneif buf:yield {"text": " ".join(buf), "start": start, "url": url}t = transcript("https://www.youtube.com/watch?v=dQw4w9WgXcQ")for c in chunks(t):print(c["start"], c["url"], c["text"][:80]) # embed c["text"], store start and url as metadata
Many videos#
- Find the videos with video discovery: YouTube search, a channel, a playlist or a TikTok profile.
- Submit them as a batch and let a webhook tell your ingester when each job finishes, instead of polling.
- Re-reading and re-exporting a transcript is free, and submitting a link you already transcribed returns the stored result, so re-running an ingest does not cost twice.
Agents that fetch on demand#
Not every pipeline needs an index. An agent connected to the TranscriptDock MCP server can call get_video_transcript when a user mentions a video and read the text straight into its context. For frameworks, see LangChain and n8n.