Interactive Interface for Embodied Agent Configuration
The method addresses the impersonality of AI assistants by enabling real-time customization and preview of digital avatar animations and behaviors, improving interaction through natural language processing and machine learning, thus enhancing user experience and efficiency.
Patent Information
- Application Number
- JP2025542354
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-24
- Filing Date
- 2024-01-24
- Publication Date
- 2026-02-03
AI Technical Summary
Existing AI assistants lack visual representation and fail to analyze user mood, posture, or movements, relying solely on verbal input, making interactions impersonal and lacking personal elements, while traditional digital avatar creation tools lack interactive preview of animations and behaviors during the creation process.
An autonomously animated digital avatar configuration method that allows real-time customization and preview of animation, behavior, and appearance through an interactive interface, utilizing natural language processing and machine learning to adapt to user inputs and environmental conditions.
Enhances the interaction experience by providing personalized and efficient customization of digital avatars with real-time feedback, reducing cognitive load and power consumption on electronic devices.
Smart Images

Figure 2026504130000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to the field of computer graphics and animation, and more particularly to techniques for autonomously customizing animated digital avatars. More specifically, but not exclusively, embodiments of the present invention relate to interactive interfaces for embodied agent configuration.
[0002] AI assistants on smart devices and platforms are becoming increasingly prevalent and can respond to users' voice requests. While these assistants can provide information and control other devices, they lack visual representation, making the interaction feel impersonal. Furthermore, these AI assistants cannot analyze the user's mood, posture, tone, or movements; they rely solely on verbal input. They also require a structured set of inputs to generate a structured set of outputs, thereby missing the personal element of human interaction.
[0003] Human face and body language are key elements of human interaction and communication, making generating realistic models of human emotional expression and behavior one of the most challenging problems in computer animation.
[0004] Interactive digital character creation tools typically offer a variety of interface designs that allow for adjustable levels of user involvement, although some offer few or no customization options in order to keep the user experience as simple as possible.
[0005] Some avatar configuration approaches allow for avatar graphic customization with previews only (no animations) and are based on a graphical user interface where the user customizes their digital avatar through pre-selected interchangeable features (e.g., facial hair, face shape, wrinkles), sliders (e.g., weight, age, ethnicity), or deformable features (e.g., nose, mouth, eye, chin deformation) modified via a small set of locators. Examples include Unreal Metahuman Creator, MakeHuman, Nintendo Miis, and Reallusion Inc.'s Character Creator.
[0006] It is therefore an object of the present invention to provide a more efficient and / or intuitive technique for creating animation, display and appearance of digital characters that overcomes at least some of the above-mentioned shortcomings of existing techniques.
[0007] Traditional solutions for designing, developing, and deploying digital humans (e.g., "Generating Digital Avatar" US2021201549A1) lack the ability to interactively preview the embodied agent's animations, behaviors, and experiential settings during the creation process. Summary of the Invention
[0008] The autonomously animated avatars are configurable, with configurations updated and / or previewed substantially in real time.It is an object of the present invention to improve interactive interfaces for embodied agent configuration, or at least to provide a useful option to the public or industry. [Brief explanation of the drawings]
[0009] [Figure 1] 1 illustrates a user interface for interactive customization of a digital avatar. [Figure 2]1 shows an application flow diagram for deploying a digital avatar. [Figure 3] 1 shows a system diagram for interactive customization of digital avatars. [Figure 4] 1 illustrates a user interface for interactively customizing the behavior of a digital avatar. [Figure 5] 1 illustrates a user interface for interactively customizing the appearance of a digital avatar. [Figure 6] 1 illustrates a user interface for interactively customizing the angle of a virtual camera. [Figure 7] 1 illustrates a user interface for interactively customizing a digital avatar's voice. [Figure 8] It shows how to deploy a digital avatar. [Figure 9] Demonstrate how to interactively customize your digital avatar. DETAILED DESCRIPTION OF THE INVENTION
[0010] One embodiment of the present invention provides a method for customizing a digital avatar. Digital avatars are also referred to herein as "avatars," "digital characters," "digital humans," "virtual agents," "embodied agents," etc. Such digital avatars can provide digital representations of real or fictional characters. However, the concepts and principles disclosed herein are not limited to digital humans. Accordingly, digital avatars can represent any type of virtual creature, such as a humanoid, animal, alien, organism, or life form with a particular visual appearance. In the broadest sense, digital avatars can include any type of embodied agent, for example, in the form of a virtual object or digital entity. Thus, digital avatars can include large models of humans or animals (e.g., a human face), as well as other models represented in or available in a virtual environment or a computer-generated or computer-implemented environment. In some cases, digital avatars may be limited to only a portion of a real body (e.g., a hand, a face, or other body part), particularly when a complete model is not required. Digital avatars are preferably animated and therefore capable of displaying multiple facial expressions.
[0011] (User interface for graphical representation and customization options of digital avatars) FIG. 1 illustrates an authoring user interface 100 in one embodiment. The user interface 100 includes a display 112 that displays a digital avatar 102. The user interface 100 also includes a microphone 114 and a speaker 116, thereby providing an audio-visual interface for a user 104 to interact with the digital avatar 102. In one embodiment, the user interacts with the avatar through customized dialogue, which is described in more detail below. In the illustrated embodiment, the display 112, microphone 114, and speaker 116 are located on an electronic device (not shown in FIG. 1). The digital avatar can be further configured to receive visual input from the real world. This visual input can include visual input such as a video stream from a camera, depth-sensing information from a camera with range imaging capabilities (e.g., a Microsoft Kinect camera), or other input such as input from biosensors, thermal sensors, or other suitable input devices. Other input received by the agent may include touchscreen input.
[0012] 1, user interface 100 includes, in addition to the audiovisual interface, additional user interface components including a graphical display 106 that displays the text of the customized conversation, user-selectable customization options 108, and a text entry field 110. These additional user interface components can further improve human-machine interaction but may be omitted in certain embodiments.
[0013] The avatar display area displays an avatar, and the customization option display area displays customization options 108. In FIG. 1, the avatar display area is on the left and the customization option area is on the right, although the invention is not limited in this respect. The avatar display area may overlap the customization option area. The avatar may be configured to interact with the customization option area (e.g., touch or point at a customization option) in the context of a conversational interaction to customize the avatar, as described in further detail below. (Script animation preview)
[0014] The user interface 100 shown in Figures 5-8 displays a script preview 400 input field, allowing the user / author to enter input text (including input text with gesture markup) and preview how the avatar will animate the input text and / or markup (where the markup specifies animation commands that define the animation parameters of the digital avatar). (Interactively autonomously animated avatars - configurable parameters)
[0015] Any suitable aspect of the digital avatar 102 may be configurable. Embodiments of the present invention allow for real-time previewing of configurable aspects. Examples include, but are not limited to: (Behavior animation)
[0016] 4 illustrates a user interface 100 that allows a user to configure the behavior animation of a digital avatar 102 by selecting one or more behavior styles from selectable customization options 108. Each behavior style is described and listed on the left side of the configuration user interface 100. The behavior style controls nuances of the animation of the digital avatar 102, such as the type and / or level of emotional expression, the set of mannerisms / gestures used by the digital avatar, or other suitable animations. In other embodiments, individual animations or animation expressions may be configured directly in the user interface 100. (Graphical representation of an avatar)
[0017] 5 illustrates a user interface 100 in which a user can configure the graphical representation of a digital avatar 102 by selecting one or more avatar graphical representations from selectable customization options 108. As different customization options 108 are selected, the digital avatar 102 in an avatar display area on the right side of the user interface 100 is updated to reflect the graphical representation selected from the customization options 108. In other embodiments, certain elements of the avatar's graphical representation may be directly configurable. Any suitable aspect of the avatar's appearance may be customizable in this manner, including, but not limited to, accessories, hairstyle, skin color, clothing, face shape, skin texture, etc. (Camera angle)
[0018] 6 illustrates a user interface for defining parameters of a virtual camera capturing a digital avatar 102 and interactively customizing a virtual camera angle x. For example, a user may select a virtual camera shot from a number of predefined virtual camera settings (close-up, medium distance, wide angle) represented as selectable customization options 108. In other embodiments, certain parameters of the virtual camera settings (e.g., camera zoom and angle) may be specifically defined. (audio)
[0019] 7 shows a user interface for interactively customizing a digital avatar's voice. Customization options 108 allow the user to select from several voice recognition options in different voices with different qualities and characteristics, accents, and / or emotional characteristics.
[0020] Further configurable voice parameters include voice pitch and voice rate.
[0021] Other suitable aspects of the animation, appearance, or interactivity of the digital avatar may be defined in a similar manner (e.g., the appearance of the virtual background or characteristics of the virtual environment in which the avatar is animated, the conversational content / dialogue spoken by the digital avatar, the language / languages understood and / or spoken by the avatar, or other suitable functionality). (language generation) Avatars can be configured to speak based on a language model or large-scale language model (LLM), whereby user requests become prompts to the LLM, which responds turn-by-turn. To tailor the digital avatar's conversational behavior, multiple LLM contextual prompts can be predefined and selectable. These contextual prompts set the scene or context to which the large-scale language model (LLM) should respond, but do not contain explicit questions or requests. They provide background information, frame the conversation, and guide how the LLM should approach the response. Subsequent responses generated by the LLM are influenced by the context established in the initial prompt, ensuring more consistent and contextually appropriate interactions. (Voice Recognition)
[0022] Speech recognition customization options may allow a user to select from speech recognition algorithms trained to recognize spoken words associated with a particular region or accent. (How to trigger customization settings) (Graphical User Interface)
[0023] In one embodiment, the user can change the parameters of the customization options 108 directly on the graphical user interface. In some embodiments, if a user desires a preview update, the user may send an update signal (e.g., an update button, a voice update command, or other suitable method). In other examples, a preview update may be requested automatically when a parameter of a customization option is changed via the user interface. (Natural language conversation creation)
[0024] In certain embodiments, instead of or in addition to a graphical user interface, customization parameters may be changed by the user using a conversational audiovisual interface. In certain embodiments, natural language processing (NLP) techniques are used to understand user intent and control blending parameters through these intents. A combination of NLP and regular expression matching may be used to extract the intent of animation changes. This method can also display possible change options to facilitate discussion, as shown in customization options 108 in FIG. 1. This method notifies the user if a requested functionality change is not possible or outside a defined range. The questions and answers uttered by the avatar to the user may be generated using NLP or similar techniques.
[0025] In certain embodiments, the customization functionality is built on top of existing animation engines, such as the applicant's Digital DNA (DDNA) Blender product, which allows users to customize the animation and interactivity of their digital avatars using slider controls, allowing the avatars to animate autonomously during the design process, i.e., animated in real time, providing real-time feedback to the user on what the avatar's facial expressions will look like as it expresses itself.
[0026] In a practical implementation according to an embodiment of the present invention, we use an English multi-task CNN model from spaCy (en_core_web_sm module), an open-source library for natural language processing in Python that includes features such as NER, POS tagging, dependency parsing, and word vectors.
[0027] In one embodiment, the model is trained in OntoNotes and optimized for CPUs. The OntoNotes project is a collaboration between BBN Technologies, the University of Colorado, the University of Pennsylvania, and the University of Southern California Information Sciences Institute. The goal of this project was to annotate a large corpus of texts in three languages (English, Chinese, and Arabic) across a variety of genres (news, conversational telephone speech, weblogs, Usenet newsgroups, broadcasts, and talk shows) with structural information (syntactic and predicate logic structure) and shallow semantic information (word senses associated with ontologies and coreferences).
[0028] When a user makes a customization request, the NLP model identifies the noun and the corresponding adjective / adverb-adjective. For example, if a user says: "Can you change my voice to female?" The NLP sends the following command to the script: voice / female.
[0029] The script executes these instructions and generates an avatar with the corresponding customizations applied.
[0030] In other embodiments, similar methods may be used to allow an end user to modify the behavior and / or appearance of a digital avatar in real time during an interaction to enhance the interaction experience, rather than just during the creation process to design the behavior and / or appearance of the digital avatar. For example, a user may request that the avatar speak a different language or change its behavior style. (Rule-based triggers)
[0031] In some implementations, one or more rules may automatically require updates without specific instruction from the creator and / or end user.
[0032] One or more listener or user detection algorithms, periodically, continuously, or temporarily (such as at the beginning of a session or when triggered by a specific event), determine if one or more conditions are met and trigger rules that require parameters to be updated in a specific manner according to the rules. For example, a detection algorithm (incorporating machine learning techniques or otherwise) may monitor audio, visual, or other inputs from the user and / or the user's environment. Detection algorithms may include age, gender, ethnicity recognition, object recognition algorithms, emotion recognition algorithms, or other algorithms.
[0033] In one embodiment, rules related to speech detection may use machine learning classification algorithms to determine the language and / or accent of a user conversing with a digital avatar. According to the rules, the determined language and / or accent may automatically trigger an update to modify the digital avatar's speech-to-text detection algorithm to match the end user's detected language and / or accent.
[0034] In another example, one rule may monitor the emotional state of a user and trigger an update of the digital avatar's personality according to a predefined mapping of emotional states to the digital avatar's personality. (Technical elements)
[0035] Various technology elements work together to provide an autonomously animated embodied agent. Elements may come from a single source (e.g., a company or technology provider) or from distributed sources. Examples include, but are not limited to: Facial emotion analysis of end-user video streams · Autonomous animation of digital people using cognitive models 3D rendering cloud architecture WebRTC video and audio streaming front-end SDK Natural Language Processing Speech to text, STT Text-to-Speech (TTS)
[0036] Digital avatars are hosted on a session server. The Video Host, which surrounds the SDK, allows for connecting to the session server to enable a webRTC session. The Video Host typically runs over webRTC or locally in windows. A configuration service can provide an API for configuration operations for the embodied agent server configuration. In one embodiment, this configuration service is a web service whose purpose is to provide a CRUD HTTP API for the configuration of the digital avatar. (How to customize animations for interactive digital avatars)
[0037] 8 illustrates one method for deploying an interactive digital avatar. A user configures form settings using a form setting interface and sends the setting parameters to a setting service 908. A session server 912 then creates a session using the requested setting parameters. The session server 912 establishes a video stream with the end user device 906. (Interactive customization with live feedback)
[0038] 9 illustrates one method for customizing animation of an interactive digital avatar. The digital avatar customization component includes a form user interface 902, a form settings 904, and a video user interface 906. The form settings 904 can be defined in any suitable computer-readable format.
[0039] When the digital avatar is first loaded, the settings service can set default settings. Upon receiving an update trigger, the settings service 908 requests the latest settings from the form settings 904 in step 1. The form settings updates the form user interface to reflect the settings in step 953. In step 954, the form settings module 904 requests a preview deployment from the settings service 906. The settings service 908 returns connection information to the form settings 904. The settings service uploads the preview deployment to the object storage service 910. The object storage service 910 sends the preview settings to the session server 912. The session server 912 establishes a video stream with the video user interface 906. The user may modify the form user interface 902. An update may be triggered again in an appropriate manner. Once the digital avatar is updated, the form settings 904 requests the latest settings from the settings service 908 in step 954. The settings service 908 returns the latest settings in step 952. Steps 954 through 959 are then repeated as described above. (advantage)
[0040] Current methods for creating and editing digital avatars on electronic devices often involve complex, time-consuming user interfaces that are cumbersome and inefficient. These existing methods waste user time and consume device energy, which is particularly problematic on battery-powered devices. This technology provides faster, more efficient methods and interfaces for creating and editing digital avatars on electronic devices. These methods and interfaces can be used in combination with or replace existing methods to improve efficiency while reducing user cognitive load. Furthermore, on battery-powered computing devices, these methods and interfaces reduce power consumption and extend the time between battery charges. The method includes displaying an avatar navigation user interface on a display of the electronic device, detecting a gesture toward the user interface, and displaying a first or second type of avatar in the avatar navigation user interface in response to the gesture. The avatar may be deployed and reconfigured even while the user is interacting with the avatar in another session. (Summary of the Invention) The present invention relates to a computer-implemented method for customizing animation of an interactive digital avatar, which provides an avatar creation user interface with a customization interface for modifying animation parameters, and a digital avatar animation display that displays real-time animation based on those parameters. The user initiates the customization process by sending an update signal indicating a request to change one or more animation parameters of the digital avatar, which triggers the deployment of an avatar session from the avatar cloud server and the receipt of a live video stream from the server. In real time, the live video stream is displayed on a digital avatar animation display and the digital avatar in the display is updated with the modified animation parameters, the update signal may be received via a graphical user interface or through conversational interaction between the user and the digital avatar. The user interface may further include an input field for content that allows the digital avatar animation display to deliver the content based on a specified delivery field. Additionally, the method supports audio input and output, where a user provides a customization request through a microphone and the digital avatar returns audio output through a speaker and simultaneously animates on the display. In certain cases, the customization response may include a query for customization options, and the customization response may consist of at least one customization option. The available customization options may depend on the state of the current customization session. The method includes determining whether the customization request satisfies predefined constraints before customizing the digital avatar. Only if these constraints are satisfied is the digital avatar customized. (interpretation)
[0041] The described methods and systems can be utilized in any suitable electronic computing system. According to the embodiments described below, the electronic computing system utilizes various modules and engines to utilize the techniques of the present invention. The electronic computing system may include at least one processor, one or more memory devices or interfaces for connecting to one or more memory devices, input and output interfaces for connecting to external devices so that the system can receive and process instructions from one or more users or external systems, a data bus for internal and external communication between various components, and a suitable power supply. Additionally, the electronic computing system may include one or more communication devices (wired or wireless) for communicating with external and internal devices, and one or more input / output devices, such as a display, pointing device, keyboard, or printing device. The processor is configured to execute steps of a program stored as program instructions in a memory device. The program instructions enable the implementation of the methods of the invention described herein. The program instructions may be developed or implemented using any suitable software programming language and toolkit, such as, for example, a C-based language and compiler. Furthermore, the program instructions may be stored in any suitable manner that allows them to be transferred to a memory device or read by a processor, including, for example, on a computer-readable medium. The computer-readable medium is any suitable medium for physically storing program instructions, including, for example, solid-state memory, magnetic tape, compact disc (CD-ROM or CD-R / W), memory card, flash memory, optical disk, magnetic disk, or any other suitable computer-readable medium. The electronic computing system is configured to communicate with a data storage system or device (e.g., an external data storage system or device) to obtain associated data. It is understood that the systems described herein include one or more elements configured to perform the various functions and methods described herein.The embodiments described herein are intended to provide the reader with examples of how various modules and / or engines constituting elements of a system can be interconnected to implement functionality. Furthermore, the embodiments described herein explain how steps of the methods described herein are performed, along with system-related details. Conceptual diagrams are provided to show the reader how various data elements are processed at various stages by various modules and / or engines. It is understood that the arrangement and configuration of modules or engines can be varied as appropriate depending on the requirements of the system and the user, and that various functions may be used by modules or engines different from those described herein, or that certain modules or engines may be integrated into a single module or engine.
[0042] It will be understood that the described modules and / or engines may be implemented or provided with instructions using any suitable form of technology. For example, a module or engine may be implemented or created using any suitable software code written in any suitable language, which code may be compiled to generate an executable program and run on any suitable computing system. In place of, or in combination with, the executable program, the module or engine may be implemented using any suitable combination of hardware, firmware, and software. For example, some of the modules may be implemented using an application-specific integrated circuit (ASIC), a system-on-a-chip (SoC), a field-programmable gate array (FPGA), or other suitable adaptable or programmable processing device. The methods described herein may be implemented using a general-purpose computing system specifically programmed to perform the described steps. Alternatively, the methods described herein may be practiced using a specific electronic computing system, such as a data sorting and visualization computer, a database query computer, a graphical analysis computer, a data analysis computer, a manufacturing data analysis computer, a business intelligence computer, an artificial intelligence computer system, or the like, which is specifically adapted to perform the described steps on specific data obtained from an environment related to a particular field.
Claims
1. 1. A computer-implemented method for customizing animation of an interactive digital avatar, comprising: a) providing an avatar creation user interface, the avatar creation user interface comprising: i) a customization interface for customizing one or more animation parameters of the digital avatar; and ii) a digital avatar animation display unit that displays real-time animation of the avatar configured and animated using one or more animation parameters; b) receiving an update signal corresponding to a request to update one or more animation parameters of said digital avatar with modified animation parameters; c) triggering deployment of the avatar session from the avatar cloud server; d) receiving a live video stream from the avatar cloud server; and e) displaying the live video stream of the digital avatar in real time in the digital avatar animation display portion using the modified animation parameters and updating the digital avatar in the digital avatar animation display portion. A method comprising:
2. The method of claim 1 , wherein the update signal is received via a graphical user interface.
3. The method of claim 1 , wherein the update signal is received through a conversational interaction between the user and the digital avatar.
4. The method of claim 1 , wherein the update signal is received from a computer algorithm configured to detect one or more characteristics of the user or the user's environment.
5. The method of claim 1 , wherein the user interface further comprises an input field for content, and the digital avatar animation display is configured to deliver content according to the delivery field.
6. a) receiving a voice input from the user via a microphone of the electronic device indicating a customization request; and b) providing an audio output by the digital avatar through a speaker of the electronic device indicating the customization response, and simultaneously animating the digital avatar on a display of the electronic device in accordance with the audio output. The method of claim 3.
7. the customization request includes a query about customization options, and the customization response includes at least one customization option. The method of claim 6.
8. The method of claim 7 , wherein the at least one customization option depends on a state of a current customization session.
9. a) determining whether the customization request satisfies one or more customization constraints; and b) customizing said digital avatar according to said customization request if, and preferably only if, said one or more customization constraints are satisfied; 9. The method of claim 6, further comprising:
10. A data processing device or system comprising means for carrying out the method according to any one of claims 1 to 9.
11. A computer program, or a computer readable medium having said computer program stored thereon, comprising instructions which, when executed by a computer, cause said computer to carry out the method of any one of claims 1 to 10.