
Overview
Cadence is a self-hostable, full-spectrum generative audio, music, speech, and media platform. It brings voice synthesis and cloning, transcription, music and sound generation, image and video creation, viral clipping, and real-time voice agents into one private production workspace.
Key Features
Voice & Speech Studio: Generate natural speech, transcribe recordings with timestamp-level navigation, clone voices, and transform voice tone with dedicated creative tools.
Music & Media Creation: Produce instrumental tracks, vocal songs, Foley sound effects, images, and video through an integrated studio experience.
Viral Video Clipper: Finds strong moments in long-form video, reframes speakers, and prepares short-form clips with accurate captions.
Real-Time Voice Agents: Build low-latency duplex voice experiences with voice activity detection, fast synthesis, and conversation inspection.
Built-In Observability: Tracks AI and model activity through a dual-stream Langfuse and Convex trace engine with latency, token, and credit insights.