Interaction models
Living Thread: continuous interaction models as the realtime System 1/control layer around slower reasoning/agentic System 2.
Current hypothesis
Interaction models are emerging as a distinct model/system layer for continuous human–AI interaction. The important shift is from turn-based generation to a continuously running interaction policy that can perceive, wait, speak, interrupt, backchannel, react nonverbally, and route work while the interaction continues.
A likely durable architecture has three different time scales:
- fast perception/generation: audio/video codecs, speech, gaze, gesture, avatar/body generation;
- realtime Interaction Model: continuous perception, timing, turn-taking, interruption, affect, routing and delegation;
- background System 2 / agent: slower reasoning, search, coding, computer use and long-running tasks.
In this framing, full-duplex is necessary but not sufficient. The more important capability is that actions such as wait, speak, interrupt, react, call_tool, and dispatch_background_task belong to one continuous policy. Once delegation is one of those actions, an Interaction Model can serve as System 1 while a stronger background model or agent supplies System 2.
Current evidence / exemplars
- Thinking Machines Lab Interaction Models: a general continuous-interaction architecture using short micro-turns and an explicit background model for slower reasoning and tools.
- Tavus Griffin: a closely related continuous conversational controller extended with streaming speech/video generation and embodiment; it makes gaze, facial expression, gesture and body behavior part of the realtime loop.
- GPT-Live and Gemini Live: important adjacent systems for native realtime/full-duplex multimodal interaction and delegation/orchestration, useful for testing whether this becomes a general model category rather than a specialized avatar stack.
- Earlier systems such as Moshi and SeedDuplex remain useful reference points for the evolution from native speech-to-speech/full-duplex toward broader interaction modeling.
Current synthesis
TML and Griffin look less like competing ideas than complementary realizations of the same architecture. TML emphasizes generic interaction intelligence plus a background reasoning branch; Griffin emphasizes interaction intelligence plus an audiovisual embodiment/output layer. A combined system would keep both: the Interaction Model decides what to do now, the background agent handles slow cognition and long tasks, and the audiovisual generator realizes the behavior continuously.
The main research question is therefore broader than latency or avatar realism: whether Interaction Model becomes a stable model category and the human-facing control/orchestration plane for persistent personal/team agents.
Open questions
- Will frontier labs converge on a dedicated Interaction Model layer, or absorb these capabilities into general multimodal foundation models?
- Which interaction decisions need to be learned jointly in the model, versus remaining in a harness/controller?
- How should realtime System 1 and background System 2 share mutable context while both continue running?
- Can user corrections during an ongoing interaction update, redirect or cancel already-dispatched background work cleanly?
- What is the right benchmark for continuous interaction beyond end-to-end latency, turn-taking accuracy and short-term human-likeness?
- How far should embodiment extend: voice and face only, or also computer/device/robot control as part of the same interaction policy?
- Does the interaction layer naturally become the persistent personal-agent control plane that routes models, tools, devices and sub-agents?
Maintenance rule
Update this Thread only when new evidence changes, strengthens, weakens, or materially complicates the current architecture or category-level judgment. Routine realtime-model releases, incremental latency gains, avatar-quality improvements, or isolated benchmarks can remain Signals unless they change the synthesis.