We help AI labs and data teams collect YouTube video, audio and structured metadata for model training. Choose 1–40 Gbps unmetered proxies to run your own pipeline, or let us deliver curated datasets straight to your cloud storage.
From raw video streams to clean, structured metadata — collected at scale, deduplicated and organized to your specification.
Full-resolution video (up to 4K/8K), audio-only tracks, and selected formats. Delivered as original containers or transcoded to your target codec.
Title, description, tags, channel, publish date, duration, view/like counts, category, language and more — normalized to JSONL or Parquet.
Auto-generated and manual captions in all available languages, with timestamps aligned to the video for speech and multimodal training.
Top-level comments, replies, like counts and reply threads, useful for conversational, sentiment and ranking datasets.
All thumbnail resolutions plus optional keyframe extraction at custom intervals for vision-language pretraining.
Collect by keyword, channel list, playlist, topic, language, region, upload window or engagement thresholds. You define the scope.
Both options are built for sustained, high-volume collection without rate limits or bandwidth caps.
Dedicated 1–40 Gbps unmetered proxy capacity for your own collection stack.
Tell us what you need. We collect, process and deliver directly to your storage.
Every record is validated and normalized. Custom fields and schemas are available on request.
| Field | Type | Description |
|---|---|---|
video_id | string | Unique YouTube video identifier |
title / description | string | Original title and full description text |
channel_id / channel_name | string | Uploader channel identifiers and subscriber count snapshot |
published_at | datetime | Upload timestamp (UTC, ISO 8601) |
duration_sec | integer | Video length in seconds |
view_count / like_count / comment_count | integer | Engagement metrics at collection time |
tags / category | array / string | Creator tags and YouTube category |
language | string | Detected primary language (BCP-47) |
formats | array | Available resolutions, codecs, bitrates and file sizes |
subtitles | object | Available caption tracks by language, manual vs. auto-generated |
thumbnails | array | Thumbnail URLs and local paths for all resolutions |
file_path / sha256 | string | Location of delivered media in your bucket and integrity checksum |
Share your target: keywords, channels, languages, regions, time range, formats and estimated volume.
We deliver a free sample batch with metadata so you can verify quality, schema and format before committing.
Collection runs on dedicated 1–40 Gbps capacity, with progress dashboards and daily reports.
Data lands in your cloud storage with manifests and checksums. Incremental updates available on schedule.
Pay for bandwidth or pay for delivered data — whichever fits your pipeline. Volume discounts available on both.
Unmetered traffic. Scale from 1 Gbps to 40 Gbps.
Priced by delivered volume and processing scope. Contact us for a quote.
All prices in USD. Custom enterprise agreements, annual contracts and hybrid plans (proxies + delivery) are available — contact sales for details.
Share your target scope and volume. We'll respond with a sample plan, timeline and quote — typically within one business day.