Automatically transform 16:9 widescreen YouTube podcasts, webinars, and vlogs into high-converting 9:16 portrait clips for TikTok, Instagram Reels, and Shorts. Powered by local AI face tracking & smart scene cutting.
Creating vertical content manually consumes 80%+ of an editor's time. VerticalX eliminates repetitive keyframing.
Tracking speaker faces frame-by-frame in Premiere Pro or CapCut takes 2 to 3 hours per 10-minute video clip.
Computer vision automatically tracks speaker movements and fast action, generating 9:16 cuts in seconds.
Proprietary cloud SaaS platforms limit export minutes, lock key features behind paywalls, and slap watermarks on exports.
Zero monthly subscription fees, zero watermark locks, and unlimited video processing forever on your local hardware.
Uploading gigabytes of unreleased podcasts or raw footage to third-party cloud servers risks privacy leaks and bandwidth delays.
Your video files never leave your machine. Process videos locally at maximum hardware speeds on Mac or PC.
No editing expertise required. Upload your video and let VerticalX handle the reframing pipeline.
Drag and drop any widescreen 16:9 MP4 or MOV file into the VerticalX web editor interface.
PySceneDetect identifies shot boundaries while computer vision centers active speaker faces dynamically.
Batch render 9:16 vertical shorts complete with thumbnails and automatic Whisper subtitle captions.
Comprehensive video reframing architecture for creators, podcast hosts, and developers.
Tracks human faces, animals, or fast action across video frames, keeping subjects centered automatically.
Select optimal framing for any shot: Smart Focus, Center Crop, Blurred Background, Black Fit, Zoom-Fit, and Dual Stack Split Screen.
Integrated OpenAI Whisper engine automatically transcribes spoken audio into frame-accurate WebVTT vertical subtitles.
Probes Apple Silicon GPU (`h264_videotoolbox`) and NVIDIA GPUs (`h264_nvenc`) for blazing fast batch rendering.
Visual browser editor for timeline tuning + a headless Python CLI engine for batch automation.
Detects when speakers sit far apart and warns you to select Dual Stack split-screen layout.
Never miss a speaker transition. VerticalX calculates frame-by-frame face coordinates and applies exponential moving average (EMA) smoothing for jitter-free panning.
Select the framing technique tailored to podcasts, gaming highlights, or vlog interviews.
Dynamically repositions the 9:16 crop window across horizontal space to keep active speakers centered in every frame.
Stacks two widescreen subjects vertically (Top Speaker / Bottom Guest), perfect for 2-person podcast interviews where subjects sit apart.
Places the original landscape video in the center while filling top and bottom bars with a scaled gaussian-blurred background.
Directly crops the exact center 1080x1920 viewport from the landscape source video without tracking movement.
Preserves 100% of the horizontal video frame by scaling it down to fit in the 9:16 frame with clean black padding top and bottom.
Scales the landscape video with custom zoom percentage adjustments (0% to 100%) to balance framing and letterbox padding.
Clone the repository and launch the Web Studio editor or Python CLI on Mac or PC.
Yes. VerticalX is completely open-source under the MIT license. There are no paid tiers, no subscription fees, no export minute limits, and no video watermarks.
No. VerticalX runs entirely on your local machine (100% offline). Your video files, audio transcripts, and metadata never leave your hardware.
VerticalX automatically probes your hardware capabilities, supporting Apple Silicon VideoToolbox (`h264_videotoolbox`) on macOS and NVIDIA NVENC (`h264_nvenc`) on Windows/Linux, with fallback to CPU `libx264` encoding.
VerticalX utilizes OpenCV facial cascades, MobileNet object detection, and optical motion flow to track human subjects, applying exponential moving average (EMA) smoothing for fluid camera panning.