Zero-Config Architecture

Targeted Compression

A standalone pipeline targeting strict platform upload limits. It intelligently probes, calculates optimal Bits-Per-Pixel, splits the file, and runs multithreaded hardware encoding.

Download Latest Release
$ ./20mb-hevc-win64.exe input.mp4 # Fast CLI Usage
20MB
Free Tier Users

Discord DMs, Group DMs, and Servers with less than Level 2 boosts.

50MB
Server Lvl 2

Optimized for Discord Server Level 2 Boost environments.

100MB
Server Lvl 3

Maximized for Discord Server Level 3 Boost limits.

500MB
Nitro Users

Designed for the expanded limits of Discord Nitro subscribers.

Drag & Drop Demonstration

The Execution Pipeline

1. Input

Drag & drop a video onto the preset executable.

2. Probe

ffprobe extracts duration, resolution, and fps.

3. Smart Split

Divides video based on 50% bit density mark to balance thread load.

4. Dual Encode

Two concurrent FFmpeg instances process chunks in parallel.

5. Stitch

Concatenates encoded chunks with +faststart for instant progressive streaming.

// HARDWARE TELEMETRY & MEMORY TOPOLOGY

Pipeline Dataflow Architecture

BUS PROTOCOL: PCIE 4.0 x16 / NVDEC / NVENC

Physical tracing of compressed NAL packets, raw pixel frames, and presentation timestamps across Storage, Host System Memory, the PCIe interconnect, and dedicated GPU silicon.

Benchmark Reference Baseline & Mathematical Assumptions
GROUND TRUTH METRICS
Test Asset Profile
1920x1080 @ 60.00 fps
H.264 High Profile (Progressive)
Sample Duration
236.95 s (03:56.95)
14,217 total frames (236.95 x 60)
Storage Payload
457.00 MB Container
Bitrate: 16,178 kbps compressed
Target Container Limit
< 100.00 MB
Rate: 3,446 kbps (78.6% reduction)
// 1080p Uncompressed Payload: 1920 x 1080 x 1.5 bytes (NV12/YUV420p) = 3,110,400 bytes/frame (3.11 MB). Total = 14,217 x 3.11 MB = 44.22 GB.
Rate: 186.6 MB/s per 60fps worker
// 720p Scaled Payload: 1280 x 720 x 1.5 bytes = 1,382,400 bytes/frame (1.38 MB). Total = 14,217 x 1.38 MB = 19.65 GB.
Rate: 82.9 MB/s per 60fps worker
Why the PCIe variance? In Native 1080p, frames never leave GPU VRAM; only compressed bitstream (457 MB in + 88 MB out) traverses PCIe (~545 MB). In Downscaled 720p, FFmpeg downloads raw 1080p frames to host RAM for CPU Lanczos scaling (44.22 GB D2H) and re-uploads 720p frames to NVENC (19.65 GB H2D), consuming ~64.41 GB total PCIe bandwidth.
// MULTITHREADED EXECUTION ENGINE

Smart Split Worker Processing

Zero-Decode Fast Seek & Byte-Load Balancing
// ALGORITHMIC FLOWCHART: ZERO-DECODE FAST SEEK & WORKLOAD BALANCING
Smart Split 3-Phase Workload Partitioning Flowchart
PHASE 01 // PROBE & ACCUMULATE
ffprobe CSV Packet Stream
ffprobe -select_streams v:0 -show_entries packet=pts_time,size,flags -of csv=p=0

Line-by-line iterator tracks cumulative packet bytes and keyframe timestamps (flags=K). Zero pixels are decoded.

O(1) MEMORY (< 10 MB) Target = 50% Bytes
PHASE 02 // PARTITION AT KEYFRAME
Smart Split Point T
Worker 1: [0.0s → T]
50% Byte Load • Bitrate br1
Worker 2: [T → EOF]
50% Byte Load • -avoid_negative_ts make_zero

Keyframe alignment prevents corrupt B-frames or video freezing at cut boundaries.

BALANCED PARALLEL WORKLOAD
PHASE 03 // CONCURRENCY & MUX
Dual Encode & Faststart Concat
ffmpeg -f concat -safe 0 -i list.txt -c copy -movflags +faststart

Worker 1 and 2 run simultaneously on separate subprocesses. Concat demuxer stitches slices without re-encoding, placing the moov box at byte 32.

ZERO-COPY STITCH moov guide →
STAGE 01 // PROBE
Packet Stream

ffprobe streams compressed packet headers line-by-line via CSV pipe (packet=pts_time,size,flags). Zero frames are decoded, keeping RAM under 10 MB.

O(1) MEMORY FOOTPRINT
STAGE 02 // BALANCE
50% Byte Load

The accumulator tracks cumulative byte weight and selects the nearest keyframe at exactly 50% data load, giving both workers equal computational work and preventing thread starvation.

BALANCED WORKLOAD
STAGE 03 // EXECUTE
Dual Workers

Worker 1 encodes [0.0s → Split], Worker 2 encodes [Split → EOF]. Worker 2 resets start timestamps to 0.0s via -avoid_negative_ts make_zero for clean stitching.

INDEPENDENT SLICES
STAGE 04 // MUX
Zero-Copy Concat

FFmpeg concat demuxer (-c copy) joins bitstreams in sub-seconds without GPU re-encoding, relocating the index atom (moov box) to byte 32 for immediate web streaming.

FASTSTART COMPLIANT moov guide →
// SELECT EXECUTION PIPELINE
// HARDWARE TOPOLOGY RACK PHYSICAL DATA BUS
BAY 01 // STORAGE SUBSYSTEM
NVMe M.2 / DISK
Input Bitstream
457.00 MB
input.mp4 (H.264)
Muxed Output
< 100.00 MB
output.mp4 (+faststart)
BAY 02 // HOST SYSTEM MEMORY & CPU
DDR4 / DDR5 RAM
Demux Buffers
O(1) Streaming
< 15 MB RAM Footprint
CPU Lanczos Scale
Bypassed (Native)
libswscale filtergraph
BAY 03 // PCIE 4.0 x16 INTERCONNECT
Minimal (~2 MB/s)
Host → Device (H2D)
~457 MB (Stream)
Inbound to VRAM
Device → Host (D2H)
~88.5 MB (Muxed)
Outbound to Host
BAY 04 // GPU SILICON & GDDR VRAM
HARDWARE ASICS
NVDEC ASIC
Hardware Decode
Direct VRAM Surface
NVENC ASIC
Hardware Encode
VBR Multipass Fullres
// LIVE TELEMETRY READOUT
Direct Hardware Path
Total PCIe Volume
~545 MB Total
99.1% Bus Reduction
Frame Pixel Format
pix_fmt: cuda
VRAM Surface
Scaling Location
Bypassed
Zero Pixel Copies
Host CPU Load
< 5% Load
Demux / Telemetry
Zero-copy VRAM execution with zero software scaling copies.

Hardware Native Unscaled (1080p60)

Frames are demuxed from storage and streamed as compressed NAL packets across PCIe into GPU VRAM. NVDEC decodes into CUDA surfaces, the frame scaling filter is completely bypassed (zero software pixel copies), and NVENC encodes directly on-chip. Only compressed bitstream traverses PCIe.

// ARCHITECTURAL IMPLICATION

When downscaling is triggered by the BPP calculator (e.g. 1080p to 720p), keeping the filter on the CPU causes a massive 64 GB PCIe roundtrip. Transitioning to scale_cuda allows the GPU to resize frames in VRAM, recovering native 545 MB PCIe efficiency while preserving quality.

// PHYSICAL DATAFLOW SCHEMATICS & BUS ROUTING

Physical PCIe Transfer Paths

Physical Routing Analysis: In Native 1080p, frames remain in GPU VRAM. However, when downscaling (e.g. 1080p to 720p), the filtergraph execution location dictates PCIe bandwidth. Below is the physical diagrammatic comparison across the architectural iterations: Path 1 (CPU decode bottleneck), Path 2 (current NVDEC + CPU scale), and Path 3 (scale_cuda zero-copy potential).
CURRENT IMPLEMENTATION // SINGLE-PASS DUAL WORKER

Path 2: Current Hybrid Downscale Pipeline (NVDEC Decode + CPU Lanczos Scaling)

PCIE TRAFFIC: ~64.41 GB ROUNDTRIP
// INTERCONNECT DATAFLOW DIAGRAM (PCIE PING-PONG ROUNDTRIP)
Path 2 Current Hybrid Downscale Pipeline Flowchart
01 // HOST SYSTEM RAM DDR4 / DDR5
Demux Inbound Stream
457.00 MB Bitstream
Streamed to GPU NVDEC
Decoded 1080p Frames (D2H)
44.22 GB Raw NV12
14,217 frames downloaded
CPU Lanczos Filtergraph
scale=-2:720 (libswscale)
~30% - 40% CPU utilization
Scaled 720p Frames (H2D)
19.65 GB Scaled NV12
Uploaded back to NVENC
Host I/O: Heavy RAM bus contention (~64 GB buffered)
02 // PCIE 4.0 x16 INTERCONNECT DUAL ROUNDTRIP
LANE 1: H2D INBOUND 457.00 MB
Host → GPU NVDEC (Compressed)
LANE 2: D2H DOWNLOAD 44.22 GB
NVDEC → Host RAM (Raw 1080p60)
LANE 3: H2D RE-UPLOAD 19.65 GB
Host RAM → NVENC (Scaled 720p60)
LANE 4: D2H MUX STREAM 88.50 MB
NVENC → Disk (Compressed bitstream)
Total PCIe Volume: 64.41 GB (1.6 GB/s sustained)
03 // GPU SILICON & VRAM HARDWARE ASICS
NVDEC Hardware Decoder
Active (~88% Engine)
Hardware decode into VRAM surface
Surface Ejection Point
VRAM → D2H Download
Decoded frames dumped to Host RAM
NVENC Hardware Encoder
Active (~91% Engine)
Encodes 720p stream from Host RAM
Limitation: Lacks scale_cuda filter to keep frames in VRAM
Architectural Tradeoff: Offloading decoding to NVDEC completely eliminated the 91% CPU decode bottleneck from intermediate v1.1.6-pre (Path 1). However, because scaling is performed in CPU software, the 44.22 GB decoded stream and 19.65 GB scaled stream must still ping-pong across PCIe.
// EMPIRICAL HARDWARE BENCHMARKS

Quantitative Architectural Matrix

1080p60 Source (236.95s, 14,217 frames, 457.00 MB)
Metric v1.1.5 (Downscaled) Intermediate v1.1.6-pre Current Implementation Potential scale_cuda
Passes Executed 2 passes (fake pass 1) 1 pass 1 pass 1 pass
Video Decode Location GPU NVDEC (x2) CPU Software (avcodec) GPU NVDEC (x1) GPU NVDEC (x1)
Scaling Location CPU Software (x2) CPU Software CPU Software (flags=lanczos) GPU CUDA Cores (scale_cuda)
PCIe Bus Volume ~128 GB (2x roundtrips) ~19.74 GB (H2D scaled) ~64.41 GB (1x roundtrip) 0.55 GB (545 MB // -99.1%)
Host System RAM I/O > 200 GB > 60 GB ~64 GB < 1 GB (Near-Zero I/O)
CPU Utilization 43% 91% (Starved NVENC) ~30% - 40% < 5% (Idle Host)
GPU Decode Engine ~88% 0% (Idle) ~88% ~88%
GPU 3D / CUDA Engine ~91% ~4% ~5% - 10% ~35% - 50% (scale_cuda)
Wall-Clock Encode Time 37.81 s 23.22 s (CPU bound) 22.4 s - 23.0 s ~15 s - 18 s (est.)
Architectural Takeaway: Current single-pass hybrid (Path 2) delivers a 49.6% PCIe reduction and cuts wall-clock encoding from 37.8s down to 22.4s by eliminating Pass 1 and using NVDEC. Transitioning to zero-copy scale_cuda (Path 3) unlocks an additional 99.1% PCIe reduction down to 545 MB, freeing host CPU/RAM completely.

Parallel Processing Engine

By dividing the video at a calculated keyframe and utilizing Python's threading to run multiple FFmpeg subprocesses simultaneously, the tool significantly reduces encoding times on modern multi-core systems and hardware encoders.

Progress Aggregation

A thread-safe ProgressTracker class parses stderr from both FFmpeg instances, merging ETA and speed metrics into a unified console output.

Parallel Split Encoding

Hardware encoders execute a dual-worker parallel split pipeline with on-chip rate control and hardware lookahead for maximum speed and bitrate compliance.

Safe Concatenation

Chunks are stitched using FFmpeg's concat demuxer with +faststart, relocating the moov atom to the front for progressive streaming without re-encoding.

Parallel vs Serial Efficiency

Abstract relative duration showing multithreaded efficiency.

BPP Scaling Logic Simulator

To prevent extreme pixelation, the script ensures the target bitrate yields a Bits-Per-Pixel (BPP) ≥ 0.04. If it fails, the script automatically tests lower resolutions and framerates. Use this calculator to see how the script decides whether to scale down a video based on your inputs.

Source Video Profile

Logic Candidates

Threshold: BPP >= 0.04
Resolution FPS Calculated BPP Status

Hardware Priority Matrix

The tool uses a dynamic fallback chain. It tests hardware availability with a dummy encode and selects the fastest valid encoder before falling back to CPU-based rendering. Select an OS to view its priority mapping.

Windows Priority