Skip to content
IntroductionHow it works, guides, and what you can transcribe.
QuickstartSubmit a link, wait for the job, export the transcript. Three requests.
AuthenticationAPI keys, scopes, Idempotency-Key and X-Request-Id.
JobsCreate, wait for, list, cancel and retry jobs.
TranscriptsThe transcript object and txt, srt, vtt, json exports.
BatchesUp to 50 videos in one request.
UploadsTranscribe your own audio and video files.
WebhooksGet called when a job or batch finishes; verify the signature.
ErrorsEvery error code, what it means and what to do.
Rate limitsSubmits, jobs in progress and reads per plan; the 429 response.
Pricing and credits1 credit per caption transcript, 2 per minute of AI transcription. Plans and extra credits.
SourcesYouTube, TikTok, your files and direct links: accepted URLs and modes.
MCPFind and transcribe videos from Claude, Cursor, Windsurf or your own agent.
OpenAPI SpecificationThe OpenAPI 3.1 spec for codegen and typed clients.
Claude CodeOne command adds TranscriptDock to Claude Code.
Claude appAdd TranscriptDock as a custom connector on claude.ai, Claude Desktop or mobile.
CursorAdd TranscriptDock as an MCP server in Cursor.
WindsurfAdd TranscriptDock to Windsurf so Cascade can find and transcribe videos.
OpenClawConnect TranscriptDock to OpenClaw autonomous agents.
19 results
Guides

Video transcripts for RAG and LLM pipelines

Turn YouTube and TikTok videos into chunks an LLM can search and cite: fetch timed segments, chunk by time, keep timestamp links.

Updated


Short answer: fetch the transcript as JSON, not plain text. The segments array carries a start and end time for every caption line, so you can chunk by time and store the start as metadata. Then every answer your model gives can link to the exact moment in the video.

Why timed segments beat plain text#

  • Citations: a chunk that knows it starts at 754 seconds becomes a link to that moment.
  • Natural boundaries: caption lines end at pauses, so chunks rarely cut a sentence in half.
  • Provenance: source_origin tells you whether the text came from the creator, the platform's automatic captions or AI transcription, and language gives its language.

Fetch and chunk#

This fetches one transcript and groups segments into windows of about a minute. The transcript JSON shape is documented in the API reference.

rag_ingest.py
import requests
 
API = "https://www.transcriptdock.com"
H = {"Authorization": "Bearer td_live_YOUR_KEY"}
 
def transcript(url: str) -> dict:
job = requests.post(f"{API}/v1/jobs", headers={**H, "Idempotency-Key": f"rag-{url}"[:128]},
json={"source": {"url": url}, "mode": "captions_only"}).json()
while job["status"] not in ("succeeded", "failed", "cancelled"):
job = requests.get(f"{API}/v1/jobs/{job['id']}", params={"wait": 25}, headers=H).json()
if job["status"] != "succeeded":
raise RuntimeError(job["error"]["code"])
return requests.get(f"{API}/v1/transcripts/{job['result_id']}", headers=H).json()
 
def chunks(t: dict, seconds: float = 60.0):
"""Group caption segments into windows of about `seconds`, never splitting a segment."""
url = t["source"]["canonical_url"]
buf, start = [], None
for seg in t["segments"]:
if start is None:
start = seg["start"]
buf.append(seg["text"])
if seg["end"] - start >= seconds:
yield {"text": " ".join(buf), "start": start, "url": f"{url}&t={int(start)}s" if "youtube" in url else url}
buf, start = [], None
if buf:
yield {"text": " ".join(buf), "start": start, "url": url}
 
t = transcript("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
for c in chunks(t):
print(c["start"], c["url"], c["text"][:80]) # embed c["text"], store start and url as metadata

Many videos#

  • Find the videos with video discovery: YouTube search, a channel, a playlist or a TikTok profile.
  • Submit them as a batch and let a webhook tell your ingester when each job finishes, instead of polling.
  • Re-reading and re-exporting a transcript is free, and submitting a link you already transcribed returns the stored result, so re-running an ingest does not cost twice.

Agents that fetch on demand#

Not every pipeline needs an index. An agent connected to the TranscriptDock MCP server can call get_video_transcript when a user mentions a video and read the text straight into its context. For frameworks, see LangChain and n8n.