AI 时代的新交互
Living Thread: AI is changing how intent is expressed, grounded, negotiated, executed and rendered—from semantic pointing and generative UI to continuous Interaction Models and embodied agents.
Core question
What becomes the basic interaction grammar of computing when AI can understand context, negotiate intent, generate interfaces and act across applications?
The useful unit is broader than a model or modality. Track the whole interaction stack: how the user indicates what they mean, how the system decides what to do now, how intent becomes action, and how results are presented or embodied.
A layered view
1. Intent capture and grounding
The first change is how users refer to things. AI reduces the need to translate intent into app-specific commands, file paths, copied text or carefully written prompts.
Important patterns include:
- natural language and voice;
- pointing, hovering, lasso/select and gaze;
- screen/context awareness;
- multimodal combinations such as “this + do X”.
Google Magic Pointer is a useful exemplar. A cursor gesture summons Gemini at the object already under attention; selected text, images and local context become grounded input, after which the user can ask for a semantic operation. The important primitive is therefore not the cursor itself but deictic interaction: point to “this”, then express the desired transformation or action without re-encoding the context.
This suggests a broader shift from open app → locate command → supply parameters toward indicate object/context → state intent.
2. Realtime interaction policy
Once interaction is continuous, the system also needs to decide what to do at each moment rather than only answer completed turns.
Interaction Models sit here:
- continuous perception;
- micro/mini-turn timing;
- wait / speak / interrupt / backchannel;
- gaze, affect and nonverbal response;
- routing and delegation while interaction continues.
Current exemplars include Thinking Machines Lab Interaction Models, Tavus Griffin, GPT-Live / realtime models, Gemini Live, Moshi and SeedDuplex.
Current hypothesis: this layer can become a realtime System 1 / control plane. If dispatch_background_task is one of its possible actions, slower reasoning, search, coding, computer use or other agents can serve as System 2 while the interaction loop remains live.
3. Semantic action and direct manipulation
AI can turn a local reference plus intent directly into an operation. This is distinct from conversation: the system changes the object or environment itself.
Examples to track:
- selected content → schedule, translate, summarize, transform or verify;
- natural-language editing of documents, images, video or code in place;
- cross-app actions without manual copy/paste or navigation;
- agentic direct manipulation where the UI object becomes both context and action target.
Magic Pointer belongs partly here as well: Google’s examples move from contextual understanding directly into Calendar scheduling, email analysis and image composition.
The important question is whether AI introduces a general semantic action layer above application-specific commands, similar to how graphical direct manipulation once abstracted away many textual commands.
4. Adaptive output and generative interface
AI also changes the output surface. Instead of returning only text, a system can choose or generate the representation best suited to the task.
Track:
- generative UI and temporary task-specific controls;
- custom widgets and dashboards;
- artifacts/workbenches that support direct editing and collaboration;
- rich multimodal explanation, including diagrams, audio and video;
- interfaces that progressively materialize around an ongoing task rather than requiring a fixed application beforehand.
This layer intersects with the separate workbench/sidebar discussion: the strategic question is whether future software is primarily a set of fixed applications or a reusable capability substrate over which agents generate task-specific interaction surfaces.
5. Embodiment, presence and form factor
The same interaction logic can be embodied in different surfaces:
- voice-only agents;
- audiovisual digital humans such as Griffin;
- desktop pointers and context-local assistants;
- glasses, wearables and spatial interfaces;
- robots and other physical agents.
The form factor matters because it changes available grounding signals and actions. A desktop has cursor, selection and application state; glasses add gaze and world perception; a digital human adds face, gesture and social timing; a robot adds physical action.
Current synthesis
Several apparently separate developments may be converging on a common change: AI is moving interaction from explicit command construction toward context-grounded intent negotiation and execution.
The stack can be summarized as:
ground what I mean → negotiate what happens now → act on the target → render the right interface/behavior → persist across context and devices
Interaction Models are therefore one important layer, not the whole topic. They solve the temporal/control problem. Magic Pointer-like interaction solves the grounding/direct-manipulation problem. Generative UI solves the representation problem. Agents and tools solve the execution problem. New form factors expand the available context and action space.
A stronger long-term hypothesis is that the dominant AI interface may not be “chat” at all. Chat is one way to express intent when shared context is weak. As systems gain richer grounding, continuous perception and direct action, interaction can become increasingly indexical and local: point, speak, glance, gesture, correct, interrupt, and let the agent act in the currently shared context.
Boundaries with adjacent Threads
- Interaction Models remains a technical subtopic inside this Thread: model architecture, full duplex, timing and System 1/System 2 orchestration.
- Proactive AI / triggers is adjacent but distinct: it asks when the agent wakes up; this Thread asks how human and agent interact once a shared interaction context exists.
- Agent workbench / sidebar overlaps mainly with adaptive output and editing surfaces.
- Conversation disappearing overlaps at the object-model level: richer grounding and persistent agents reduce the need for users to treat individual chat sessions as the primary interaction container.
Open questions
- Does a stable set of AI-native interaction primitives emerge—point, select, speak, gaze, interrupt, delegate—analogous to click, drag, type and scroll in GUI computing?
- Is “this + intent” / deictic interaction the most important bridge between GUI direct manipulation and language-based AI?
- Which interactions should remain explicit for control and trust, and which can become inferred or proactive?
- Will semantic actions become portable across applications, or remain proprietary features tied to each operating system/app ecosystem?
- How should an Interaction Model, UI state and background agent share the same evolving context without losing provenance or user control?
- Does generative UI replace parts of the fixed application model, or mostly augment existing applications and workbenches?
- Which form factors materially change interaction rather than simply moving the same assistant to another screen?
- What benchmarks can measure interaction quality beyond task success and latency: grounding accuracy, correction cost, interruption handling, action reversibility, attentional burden and continuity?
Maintenance rule
Update this Thread when a new model, product, interface primitive, operating-system feature, form factor, benchmark or empirical result materially changes the layered picture above. Prefer developments that introduce a genuinely new interaction primitive or alter the boundary between grounding, realtime control, action, representation and embodiment. Routine feature additions or cosmetic UI changes can remain Signals without changing the synthesis.