Skip to content
IntroductionHow it works, guides, and what you can transcribe.
QuickstartSubmit a link, wait for the job, export the transcript. Three requests.
AuthenticationAPI keys, scopes, Idempotency-Key and X-Request-Id.
JobsCreate, wait for, list, cancel and retry jobs.
TranscriptsThe transcript object and txt, srt, vtt, json exports.
BatchesUp to 50 videos in one request.
UploadsTranscribe your own audio and video files.
WebhooksGet called when a job or batch finishes; verify the signature.
ErrorsEvery error code, what it means and what to do.
Rate limitsSubmits, jobs in progress and reads per plan; the 429 response.
Pricing and credits1 credit per caption transcript, 2 per minute of AI transcription. Plans and extra credits.
SourcesYouTube, TikTok, your files and direct links: accepted URLs and modes.
MCPFind and transcribe videos from Claude, Cursor, Windsurf or your own agent.
OpenAPI SpecificationThe OpenAPI 3.1 spec for codegen and typed clients.
Claude CodeOne command adds TranscriptDock to Claude Code.
Claude appAdd TranscriptDock as a custom connector on claude.ai, Claude Desktop or mobile.
CursorAdd TranscriptDock as an MCP server in Cursor.
WindsurfAdd TranscriptDock to Windsurf so Cascade can find and transcribe videos.
OpenClawConnect TranscriptDock to OpenClaw autonomous agents.
19 results
Guides

Video transcripts in LangChain

Load YouTube and TikTok transcripts as LangChain Documents with timestamps, or give a LangChain agent the TranscriptDock MCP tools.

Updated


Short answer: there are two ways. For indexing, a small loader turns a transcript into LangChain Document objects, one per caption segment with its start time. For agents, the TranscriptDock MCP server plugs into LangChain through langchain-mcp-adapters and the agent fetches transcripts itself.

A document loader#

Twenty lines, no extra package. Each Document keeps the video URL and timing, so retrieval results can cite the moment they came from. Group segments into larger chunks before embedding, as shown in transcripts for RAG.

loader.py
# pip install langchain-core requests
import requests
from langchain_core.documents import Document
 
API = "https://www.transcriptdock.com"
H = {"Authorization": "Bearer td_live_YOUR_KEY"}
 
def load_video(url: str, mode: str = "captions_only") -> list[Document]:
"""One Document per caption segment, with the video URL and start time as metadata."""
job = requests.post(f"{API}/v1/jobs", headers={**H, "Idempotency-Key": f"lc-{url}"[:128]},
json={"source": {"url": url}, "mode": mode}).json()
while job["status"] not in ("succeeded", "failed", "cancelled"):
job = requests.get(f"{API}/v1/jobs/{job['id']}", params={"wait": 25}, headers=H).json()
if job["status"] != "succeeded":
raise RuntimeError(job["error"]["code"])
t = requests.get(f"{API}/v1/transcripts/{job['result_id']}", headers=H).json()
meta = {"source": t["source"]["canonical_url"], "title": t["source"]["title"], "language": t["language"]}
return [Document(page_content=s["text"], metadata={**meta, "start": s["start"], "end": s["end"]})
for s in t["segments"]]
 
docs = load_video("https://www.youtube.com/watch?v=dQw4w9WgXcQ")

Agent tools over MCP#

MultiServerMCPClient connects to the remote server and turns its tools into LangChain tools: get_video_transcript, YouTube search, channel and playlist listing and TikTok profile listing.

agent.py
# pip install langchain-mcp-adapters langgraph langchain-anthropic
import asyncio
from langchain_mcp_adapters.client import MultiServerMCPClient
from langgraph.prebuilt import create_react_agent
 
async def main():
client = MultiServerMCPClient({
"transcriptdock": {
"url": "https://www.transcriptdock.com/mcp",
"transport": "streamable_http",
"headers": {"Authorization": "Bearer td_live_YOUR_KEY"},
}
})
tools = await client.get_tools()
agent = create_react_agent("anthropic:claude-sonnet-5-5", tools)
out = await agent.ainvoke({"messages": "Summarize https://www.youtube.com/watch?v=dQw4w9WgXcQ"})
print(out["messages"][-1].content)
 
asyncio.run(main())

Tool names, arguments and credit costs are on the MCP page. LlamaIndex and other frameworks that speak MCP connect the same way.