STATUS: COMPLETED PROTOTYPEMODE: Local Operator Workspace RESTRICTED_ACCESS

Public Demo

Local-first streaming system for TikTok Live. Transforms live comments and gifts into contextual AI speech, local TTS, and synchronized animations.

CapyOwlCat

Real-time Autonomous AI Character for TikTok Live · Local Prototype

Next.js 16React 19TypeScriptZustandxAI GrokTikTok Live ConnectorPythonFastAPIPiper TTSFFmpegTailwind CSS
[SYSTEM_OVERVIEW]

CapyOwlCat is an interactive AI character system built for live streaming content. It receives TikTok Live comments and gifts, prioritizes incoming events, generates in-character responses with Grok, synthesizes speech through a local Piper TTS service, and coordinates the character's visual behavior through a timeline-driven animation state machine.

The system includes a dedicated operator workspace for managing animations, reactions, voices, visual layers, subtitles, background audio, and broadcast composition. Built as a completed local-first prototype demonstrating a high-performance real-time AI and media pipeline.

Core Loops
TikTok Event Ingestion (Comments & Gifts)Priority Queueing & Combo HandlingGrok Character Response GenerationLocal Piper TTS Audio SynthesisTimeline FSM State TransitionsFFmpeg Chroma-Key & Virtual 9:16 Broadcast

/// DEPLOYMENT_SCOPE

  • ::TikTok Live Connector Ingestion
  • ::Bounded Priority Event Queue
  • ::Grok Persona LLM Core
  • ::FastAPI + Piper Local TTS Engine
  • ::Timeline Animation FSM
  • ::Virtual 9:16 Operator Workspace

Engineering Architecture

Local-First FastAPI + Piper TTS Engine

Replaced cloud TTS services with a locally hosted FastAPI Python wrapper around Piper TTS. Reduces voice generation latency to milliseconds and ensures zero external bandwidth dependency during live broadcasts.

Bounded Priority Queue & Gift Combo Engine

Implements bounded queues in Zustand to handle live stream spam. High-value gifts trigger immediate combo aggregation and queue preemption over generic chat messages.

Timeline-Driven Animation FSM

An extended Finite State Machine synchronizes audio phonemes, subtitles, background audio, and sprite visual layers, ensuring lip-sync alignment and smooth animation transitions without frame drops.

Virtual 9:16 Operator Broadcast Workspace

A built-in Next.js operator interface offering real-time 9:16 virtual monitor previews, live chroma-key adjustment, layer stacking, and manual emotion/reaction triggers.

AI Engine

Grok In-Character Persona Engine

Custom prompt engineering and context sliding window for xAI Grok. Keeps responses concise, humorous, and strictly in-character while processing live chat dynamics.

Multi-Modal Event Parsing

Differentiates between user comments, viewer gifts, sub triggers, and system events, shaping the AI persona's cognitive reaction accordingly.

Synchronized Subtitle & Speech Dispatcher

Parses synthesized audio duration and streams word-by-word subtitles matching character speech cadence in real time.

Admin & Security

Operator Reaction Workspace

Full manual control dashboard to override AI responses, trigger custom gift animations, tweak voice pitches, or pause stream events on demand.

FFmpeg Media & Chroma-Key Pipeline

Handles real-time video compositing, transparent background keying, and audio channel mixing for OBS broadcast output.

Reliability Principles

  • Stream Comment FloodingHigh chat volume can overwhelm LLM rate limits. Solved with bounded queue dropping and combo aggregation.
  • TTS Latency SpikesCloud speech synthesis causes awkward live stream pauses. Mitigated by running local Piper TTS on GPU/CPU.
  • Animation DesynchronizationAudio length variations can desync visuals. FSM relies on audio duration callbacks to trigger state resets.

Invariants (Strict Rules)

  • ::Strict Priority Queue: Gift combos and direct interactions jump ahead of general chat
  • ::Deterministic State Transitions: FSM guarantees seamless morphing between idle, talking, gift, and emotion animations
  • ::Local-First Processing: Audio synthesis via local FastAPI Piper TTS avoids third-party API latency and usage fees
  • ::Bounded Latency Anchor: Persona context window is bounded to maintain sub-second response cadence
  • ::Operator Override: Workspace controls can instantly preempt AI state or adjust visual layers

Data Entities

event_queue (bounded fifo + priority queue)character_persona (grok system prompt + context window)animation_states (idle, speaking, gift_reaction, emotion_burst)tts_jobs (fastapi piper synth payload + audio buffer)stream_composition (layer stack, chroma key specs, subtitle buffer)

UX Philosophies

  • Real-time 9:16 Broadcast Preview
  • Instant Operator Manual Override
  • Visual Layer & Chroma Key Adjustment
  • Stream-Safe Content Boundaries

Future Signals

  • Multi-character Live Duo Streaming
  • VTuber / 3D Model Blendshape Support
  • Twitch & YouTube Live Multi-Connector