Persona-Based Assistant Output with Dynamic Visual Cues

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated assistants lack visual richness in their dialog sessions, with current visual representations being limited in conveying information, failing to replicate the diverse and dynamic facial expressions and body language humans use in communication.

Innovation Solution

Implementations enable automated assistants to dynamically adapt their outputs based on assigned personas, incorporating personalized textual content and visual cues, such as animated gestures and display animations, through the use of large language models (LLMs) to enhance visual content during dialog sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If hard-coded rules are used to control visualized representation, then the implementation is simple, but the visual richness and adaptability are limited

Engineering Contradiction:
Improvevisual richnessVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transitions from static hard-coded rules to dynamic persona-based control. Visualized representations are now controlled by persona assignments that can be dynamically selected and changed, allowing the same automated assistant to exhibit different visual behaviors (e.g., waving, nodding, facial expressions) based on the assigned persona, thereby achieving visual richness without permanent system complexity

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces persona as a controllable parameter that changes the behavior of the visualized representation. By modifying the persona parameter, the system can switch between different visual styles, gestures, and expressions, enabling adaptability while keeping the underlying control mechanism relatively simple through parameter-driven behavior selection

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If static visual content is provided, then the implementation is simple, but the visual content lacks richness and fails to mirror natural human communication

Engineering Contradiction:
Improvevisual content adaptabilityVSAvoidimplementation simplicity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent makes visual content dynamic by linking it to persona assignments. Instead of static visual elements, the system now generates visual content that adapts to the current persona, allowing visualized representations to display context-appropriate gestures, expressions, and animations that mirror natural human communication patterns

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces persona as an intermediary layer between the automated assistant's response and the visual content generation. This intermediary translates the assistant's textual or auditory response into appropriate visual manifestations based on the assigned persona, enabling rich and adaptive visual content while maintaining implementation simplicity through the mediating persona framework

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12400634B2Dynamically adapting given assistant output based on a given persona assigned to an automated assistant
Publication Date: 2025.08.26 GOOGLE LLC
  • US12400634B2 patent drawing
  • US12400634B2 patent drawing
  • US12400634B2 patent drawing

AI summary

Implementations relate to dynamically adapting a given assistant output based on a given persona, from among a plurality of disparate personas, assigned to an automated assistant. In some implementations, the given assistant output can be generated and subsequently adapted based on the given persona assigned to the automated assistant. In other implementations, the given assistant output can be generated specific to the given persona and without having to subsequently adapt the given assistant output to the given persona. Notably, the given assistant output can include a stream of textual content to be synthesized for audible presentation to the user, and a stream of visual cues utilized in controlling a display of a client device and/or in controlling a visualized representation of the automated assistant. Various implementations utilize large language models (LLMs), or output previously generated utilizing LLMs, to reflect the given persona in the given assistant output.