PikaStream 1.0: Pika Labs Gives Every AI Agent a Face, a Voice, and a Seat in Your Meeting
Pika Labs just launched PikaStream 1.0, a real-time video model that lets AI agents join Google Meet calls with rendered avatars, cloned voices, and full task execution. Here is what it means for the future of AI interaction.

The Shift Nobody Expected
Pika Labs built its reputation on creative tools. Pika 2.5 made text-to-video accessible. Pika Effects turned static images into motion. The trajectory seemed clear: prettier clips, longer durations, better physics.
Then they shipped PikaStream, and the trajectory bent sideways.
Instead of competing on clip quality (a race that Runway, Kling, and Google are already running), Pika built a real-time video engine designed for live interaction. The target user is not a filmmaker. It is an AI agent.
What PikaStream Actually Does
The core product is a self-contained skill module that plugs into AI coding agents like Claude Code, OpenAI assistants, or anything supporting the Pika Developer API. Once installed, the agent can:
- Join a Google Meet call with a fully rendered avatar
- Speak with a cloned voice (record a short sample, the agent talks like you)
- Maintain memory of previous conversations and context
- Execute tasks during the call (write code, search data, trigger actions)
The avatar is not a static image with a lip-sync overlay. PikaStream generates personalized video at up to 30 FPS and 480p resolution, powered by a real-time diffusion pipeline running on a single H100 GPU.
Under the Hood: FlashVAE and Streaming DiT
Two architectural innovations make the real-time performance possible.
FlashVAE is a full Transformer-based Variational Autoencoder trained from scratch. It provides its own latent space and reconstructs video through streaming decoding at hundreds of frames per second with minimal GPU memory. This is the component that makes real-time output feasible without needing a GPU farm.
The 9B Diffusion Transformer (DiT) uses a clever training trick: a bidirectional teacher model is distilled into a causal autoregressive student through optimized self-forcing. This enables chunk-by-chunk streaming at real-time frame rates, much like how language models stream text token by token.
Why This Matters More Than Another Pretty Video Model
The AI video space in April 2026 is crowded. Runway Gen-4.5 handles cinematic quality. Kling 3.0 does multi-shot storyboards. Seedance 2.0 generates audio and video together. Google Veo 3.1 is everywhere.
PikaStream does not compete with any of them.
It competes with Zoom avatars, Synthesia-style presenters, and the entire concept of "you must be present on camera." The difference is that PikaStream avatars are not pre-recorded or scripted. They react, respond, and act in real time.
The "AI Self" Concept
Pika is also pushing a consumer angle they call Pika AI Self. The idea: create a digital version of yourself that can attend meetings on your behalf. Your avatar looks like you, sounds like you (via voice cloning), and has access to your calendar, notes, and conversation history.
This is not science fiction anymore. The beta works on Google Meet today, with Zoom and FaceTime support coming soon.
The implications for remote work are immediate:
- Skip the meeting, send your AI Self with context about what you need
- Let the agent take notes, make decisions within defined boundaries, and report back
- Multiple meetings at the same time, each with your consistent presence
Pricing and Access
PikaStream is priced at $0.20 per minute of meeting participation. For a 30-minute meeting, that is $6. Compared to sending a human representative or missing the meeting entirely, the economics work for certain use cases: status updates, information gathering, routine check-ins.
The skill is available through the Pika Developer API and can be installed as a module in compatible agent frameworks. Google Meet is the only supported platform at launch.
What This Signals for the Industry
PikaStream represents a category split in AI video. On one side: offline generation (Runway, Kling, Veo) focused on content creation. On the other: real-time generation for live interaction, telepresence, and agent embodiment.
Text-to-Video Era
4-10 second clips, impressive demos, limited practical use
Production-Grade Clips
60-second coherent videos, native audio, character consistency
Real-Time and Interactive
Live video generation for agents, meetings, and interactive applications
The models that generate beautiful 2-minute clips and the models that power live 30 FPS avatars are solving different problems. PikaStream is the first serious entry in the second category, and it will not be the last.
For creators and developers, the question is no longer just "what can AI video look like?" It is "where can AI video show up, and what can it do when it gets there?"
Sources

Creative technologist from Lausanne exploring where AI meets art. Experiments with generative models between electronic music sessions.
View profile →Related Articles
Continue exploring with these related posts

Luma Agents and the Rise of Agentic Video Editing: AI Learns to Direct
Luma just launched creative AI agents that coordinate text, image, video, and audio. Combined with a16z backing the thesis, agentic video editing is the biggest shift since text-to-video.

Pika 2.5: Democratizing AI Video Through Speed, Price, and Creative Tools
Pika Labs releases version 2.5, combining faster generation, enhanced physics, and creative tools like Pikaframes and Pikaffects to make AI video accessible to everyone.

AI Video Tools Pricing 2026: How to Budget Without Burning Credits
AI video tools are no longer hard to find. The hard part is choosing the right pricing model for your workflow. Here is a practical budgeting guide for creators and teams.