Interactive interfaces for embodied agent configuration
Patent Information
- Application Number
- EP2024747026
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-24
- Filing Date
- 2024-01-24
- Publication Date
- 2025-12-03
AI Technical Summary
Existing digital avatar configuration tools lack the ability to interactively preview and customize the animation and behavior of digital characters, relying on structured inputs and lacking personalization, which limits user experience and efficiency.
An interactive interface that allows real-time customization of digital avatars through an audio-visual interface, using natural language processing and machine learning to analyze user inputs and update the avatar's appearance and behavior, enabling users to customize various aspects such as facial expressions, voice, and behavior styles, with live feedback and conversational interaction.
This solution provides a more intuitive and efficient method for creating and editing digital avatars, enhancing user interaction and reducing cognitive load, while conserving device resources and improving the personalization of digital human representations.
Smart Images

Figure IB2024050659_02082024_PF_FP
Abstract
Description
INTERACTIVE INTERFACES FOR EMBODIED AGENT CONFIGURATIONTECHNICAL HELD[1] The present invention generally concerns the field of computer graphics and animation, and in particular techniques for customizing autonomously animated digital avatars. More particularly but not exclusively, embodiments of the invention relate to interactive interfaces for embodied agent configuration.BACKGROUND ART[2] Al assistants, such as those found on smart devices and platforms, are becoming more prevalent and can fulfill verbal requests from users. These assistants can provide information and control other devices, but they do not have a visual representation and the interaction can feel impersonal. Additionally, these Al assistants can't analyze the user's mood, posture, tone, or movement, and rely only on verbal inputs. They also require a structured set of inputs to produce a structured set of outputs, which removes the personal aspect of human interaction.[3] The human face and body language is a key component of human interaction and communication. For this reason, the generation of realistic models of human emotional expression and behaviour have been one of the most interesting problems in computer animation.[4] Interactive digital character creations tools typically come with various interface designs offering a range of levels of user involvement. Some approaches try make the user experience as simple as possible and thus offer few or no customization options.[5] Some avatar configuration approaches facilitate avatar graphics customization with preview only (not animation), and are based on graphical user interfaces which allow the user to customize a digital avatar by way of certain preselected swappable features (e.g., for facial hair, face shape, wrinkles), sliders (e.g., for weight, age, ethnicity) and / or deformable features modified via a sparse set of locators (e.g., to morph the nose, mouth, eyes, jaw). Examples include Unreal Metahuman Creator, MakeHuman, Nintendo Miis and Character Creator by Reallusion Inc.[6] It is therefore a problem underlying the invention to provide more efficient and / or intuitive techniques for authoring the animation, display and appearance of digital characters which overcome the above- mentioned disadvantages of the prior art at least in part.[7] Prior solutions for designing, developing and deploying digital humans (such as “Generating Digital Avatar” US2021201549A1) lack the capability to interactively preview embodied agent animation, behaviour and experience composition during the authoring process.SUMMARY OF INVENTION[8] Autonomously animated avatars are configurable and configurations are updated and / or previewed substantially in real-time. It is an object of the invention to improve interactive interfaces for embodied agent configuration, or to at least provide the public or industry with a useful choice.BRIEF DESCRIPTION OF DRAWINGSFigure 1 shows a user interface for interactive customization of a digital avatar;Figure 2 shows an application flow diagram for deploying a digital avatar;Figure 3 shows a system diagram of interactive customization of a digital avatar;Figure 4 shows a user interface for interactively customizing digital avatar behaviour;Figure 0 shows a user interface for interactively customizing digital avatar appearance;Figure 5 shows a user interface for interactively customizing virtual camera angle;Figure 7 shows a user interface for interactively customizing digital avatar voice;Figure 8 shows a method for deploying a digital avatar; andFigure 9 shows a method for interactive customization of a digital avatar.DESCRIPTION OF EMBODIMENTS[9] One embodiment of the invention provides a method of customizing a digital avatar. A digital avatar may also be referred to herein as “avatar”, “digital character”, “digital human”, “virtual agent”, “embodied agent” or the like. Such a digital avatar may provide a digital representation of a real or fictious human. The concepts and principles disclosed herein are, however, not limited to digital humans. Accordingly, a digital avatar may likewise represent any kind of virtual organism, e.g., in the form of a humanoid, animal, alien, creature, or any life-like animated entity of a certain visual appearance. In the broadest sense, a digital avatar may comprise any type of embodied agent, e.g. in the form of a virtual object or digital entity. Accordingly, digital avatars may include both largemodels of humans or animals, such as a human face, as well as any other model represented, or capable of being used, in a virtual or computer-created or computer-implemented environment. In some cases, the digital avatar may not be complete, but may be limited to a portion of an entity, for instance a body portion such as a hand or face; in particular where a full model is not required. A digital avatar is preferably animated and thus capable of displaying multiple facial expressions.User interface for Graphical Display of Digital Avatar and Customization OptionsFigure 1 shows an authoring user interface 100 according to one embodiment. The user interface 100 comprises a display 112 which displays a digital avatar 102. The user interface 100 also comprises a microphone 114 and a speaker 116, thereby providing an audio-visual interface allowing a user 104 to interact with the digital avatar 102. In one embodiment, the user interacts with the avatar in a customization conversation, which will be explained in more detail further below. In the illustrated embodiment, the display 112, the microphone 114 and the speaker 116 are arranged in an electronic device (not shown in Figure 1). The digital avatar may be further configured to receive visual Input from the real-world, which may include visual input (such as a video stream from a camera), depth sensing information from cameras with range-imaging capabilities, such as the Microsoft Kinect camera, or other input such as input from a bio sensor, heat sensor or any other suitable input device. Other input the agent receives may be touch-screen input.
[0010] In the embodiment shown in Figure 1, the user interface 100 comprises further user interface components besides the audio-visual interface, namely a graphical display 106 of the text of the customization conversation, user-selectable customization options 108 and a text input field 110. These additional user interface components may further improve the human-machine interaction but may be omitted in certain embodiments.
[0011] An avatar display region may display the avatar and a customization option display region may display the customization options 108. In Figure 1, the avatar display region is to the left, and the customization option region is to the right, however the invention is not limited in this respect. The avatar display region may overlap with the customization option region. The avatar may be configured to interact with the customization option region (for example, touch or point to customization options) in the context of a conversational interaction for customizing the avatar, which is described in further detail below.Script Animation Preview
[0012] The User Interface 100 of Figure 5 to Figure 8 shows a script preview 400 input field. User / s authors may enter input text, which may include input text with gestural mark-up (wherein mark up specifies animation commands defining animation parameters of the digital avatar), to preview how the avatar animates input text and / or markup.Interactive Autonomously Animated Avatar - Configurable Parameters
[0013] Any suitable aspect of the digital avatar 102 is configurable. Embodiments of the invention enable preview of the configurable aspects in real time. Examples include, but are not limited to:Behavioural Animation
[0014] Figure 4 shows a user interface 100 wherein the behavioural animation of the digital avatar 102 is configurable by selecting from one or more behaviour styles from the user selectable customization options 108. Each behavioural styles is described and enumerated on the left hand side in the configuration user interface 100. Behaviour styles may control the nuances of digital avatar 102 animation, for example, the type and / or level or emotional expressivity, set of mannerisms / gestures used by the digital avatar, or any other suitable animation. In other embodiments, individual animations, or animation expressions may be directly configured in the user interface 100.Graphical Representation of Avatar
[0015] Figure 5 shows a user interface 100 wherein the graphical representation of the digital avatar 102 is configurable by selecting from one or more avatar graphic representations, from the user selectable customization options 108. When a different customization option 108 is selected, the digital avatar 102 in the avatar display region on the right hand side of the user interface 100 is updated to reflect the selected graphical representation from the customization options 108. In other embodiments, specific elements of the graphical representation of the avatar may be configurable directly. Any suitable aspect of avatar appearance may be customized in this manner, including, but not limited to, accessories, hairstyle, skin colour, clothing, facial geometry, skin texture.Camera Angle
[0016] Figure 5 shows a user interface for interactively customizing virtual x angle, defining the parameters of the virtual camera capturing the digital avatar 102. For example, a virtual camera shot may be selected from a plurality of predefined virtual camera configurations (close up, medium angle, wide angle), represented as selectable customization options 108. In other embodiments, specificparameters of the virtual camera configuration may be specifically defined, such specific camera zoom and angle.Voice
[0017] Figure 7 shows a user interface for interactively customizing digital avatar voice. Customization options 108 may enable a user to select from several speech-to-text options of different voices with different qualities and characteristics, accents and / or emotional qualities.
[0018] Further configurable voice parameters may include speech pitch and speech speed.
[0019] Any other suitable aspect of the digital avatar animation, appearance, or interactivity may be defined in a similar manner, for example, the appearance of characteristics of the virtual background or virtual environment within which the avatar is animated, the conversational content / dialogue which the digital avatar speaks, the language / s that the avatar understands and / or speaks or any other suitable feature.Language GenerationThe avatar may be configured to speak according to a language model or large language model (LLM), whereby user requests are prompts to the LLM, which the LLM responds to turn by turn. A plurality of LLM context prompts may be predefined and selectable from to tune conversational behaviour of the digital avatar. The context prompt sets the scene or context for how a large language model (LLM) should respond, but is not followed by an explicit question or request and provides background information or frames the conversation, guiding the LLM on how to approach its responses. The subsequent responses generated by the LLM are influenced by the context established in the initial prompt, ensuring a more coherent and contextually relevant interaction.Voice Recognition
[0020] Voice Recognition customization options may allow a user to select from voice recognition algorithms trained to recognize speech associated with a particular region or accent.Triggering Methods for Customization ConfigurationGraphical User Interface
[0021] In one embodiment a user may directly change parameters of customization options 108 on a graphical user interface.In some embodiments, the user may send an update signal (e.g. via an update button, verbal update command, in any other suitable manner) when they desire a preview update. In other embodiments, a preview update may be automatically requested any time a customization option parameter is changed via the user interface.Natural Language Conversation Authoring
[0022] In certain embodiments, instead of and / or in addition to a graphical user interface, customization parameters may be changed by a user using a conversational audio-visual interface. Certain embodiments use natural language processing (NLP) techniques to understand the intent of the user and to drive the blending parameters through these intents. A combination of NLP and / or regular expression matching may be used to extract the animation modification intent. The method may also display a selection of possible modifications to drive the discussion, as illustrated by the customization options 108 in Fig. 1. The method advises the user when a requested feature modification is not possible or outside a defined range. The avatar’s questions and responses to the user may be generated using NLP or other similar techniques.
[0023] In certain embodiments, the customization functionality is built on top of an existing animation engine. One example is the Digital DNA (DDNA) blender product developed by the applicant, which allows users to customize animation and interactive features of the digital avatar using slider controls. This allows the avatar to be autonomously animated during the design process, i.e. in real-time, which provides real-time feedback to the user on how the facial features of the created avatar will look when articulating.
[0024] A practical implementation in accordance with embodiments of the invention uses the English multitask CNN model from spaCy (en_core_web_sm module). spaCy is an open-source library for Natural Language Processing in Python which features NER, POS tagging, dependency parsing and word vectors.
[0025] In one embodiment, the model was trained on OntoNotes and optimized for CPU. The OntoNotes project is a collaborative effort between BBN Technologies, the University of Colorado, the University of Pennsylvania and the University of Southern California’s Information Sciences Institute. The goal of the project was to annotate a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, Usenet newsgroups, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information (syntax and predicate argument structure) and shallow semantics (word sense linked to an ontology and coreference).
[0026] When the user makes a customization request, the NLP model identifies the nouns and the corresponding adjectives / adverb-adjectives. For example, if the user says:“can you change my voice to female ”, the NLP will pass the following command to the script: voice / f emale
[0027] The script will then execute those orders and generate an avatar with the corresponding customizations.
[0028] In addition to during the authoring process, for designing the behavior and / or appearance of a digital avatar, in other embodiments similar methods may be used by end-users to change the behavior and / or appearance of digital avatars in real-time during interaction, to enhance their interaction experience. For example, users may request that the avatar speaks a different language or changes their behavior style.Rule-based Triggers
[0029] In some embodiments, one or more rules may automatically request updates, even without the specific instructions of an author and / or end-user.
[0030] One or more listeners or user detection algorithms may periodically, continuously, or on a one-off basis (such as at the start of a session or triggered by a certain event) determine that one or more condition / s are met, to trigger rule to request update of parameters in a certain manner as defined by the rules. For example, detection algorithms (implanted with machine learning methods or otherwise) may monitor audio, visual and / or other inputs from the user and / or the user’s environment. Detection algorithms may include demographic recognition algorithms, such as age, gender, ethnicity recognition, object recognition algorithms, emotion recognition algorithms, or otherwise.
[0031] In one embodiment, a rule associated with speech detection may employ a machine learning categorization algorithm to determine a language and / or an accent of a user speaking with the digital avatar. According to the rule, the determined language and / or accent may automatically trigger an update to change the speech-to-text detection algorithm of the digital avatar, to match the detected language and / or accent of an end-user.
[0032] In another example, a rule may monitor for an emotional state of a user, and according to a predefined mapping of emotional states to digital avatar personalities, may trigger an update of the digital avatar’s personality.Technology Components
[0033] Various technological components work together to deliver an autonomously animated embodied agent. Components may be provided from a single source (e.g. company or technology provider) or from distributed sources. Examples include, but are not limited to:• Facial Sentiment Analysis of End User Video Stream• Autonomous Animation of Digital Person with Cognitive Models• 3D Rendering Cloud Architecture• WebRTC Video and Audio Streaming Front-end SDK• Natural Language Processing• Speech to Text Conversion, STT• Text to Speech Synthesis, TTS
[0034] The digital avatar is hosted on a session server. A Video Host wrapper around the SDK may enable webrtc sessions by connecting with the Session Server. Video Host typically runs via webrtc or run locally within a window. A configuration service may provide an API for operations for configurations for building an embodied agent server. In one embodiment, the configuration service is a web service intended for providing CRUD HTTP API for digital avatar build configurations.Method of customizing animation of an interactive digital avatar
[0035] Figure 8 shows a method of deploying an interactive digital avatar. A user may configure a form configuration using a form configuration interface, which send the configuration parameters to an Configuration service 908, triggering the session server 912 to create a session with the requested configuration parameters. The session server 912 establishes a video stream to an end-user device 906.Interactive Customization with Live Feedback
[0036] Figure 9 shows a method of customizing animation of an interactive digital avatar. A digital avatar customization component includes a Form User Interface 902, Form configuration 904 and a Video User Interface 906. Form configuration 904 can be defined in any suitable computer-readable format.
[0037] When the digital avatar is loaded for the first time, the Configuration Service may set default configuration values. When an update trigger is received, at step 1 the Configuration Service 908requests the latest configuration from the Form Configuration 904. The Form Configuration updates the Form User Interface to reflect the Configuration at step 953. At step 954, the Form Configuration Module 904 requests a preview deployment to the Configuration service 906. The Configuration service 906 returns connection information to the Form Configuration 904. The Configuration Service uploads the preview deployment to the Object Storage Service 910. The Object Storage Service 910 Sends preview configuration to the Session Server 912. The Session Server 912 Establishes a Video Stream to the Video User Interface 906. A user may modify the Form User Interface 902. An update may be triggered again, in any suitable manner. When the digital avatar is updated, at step 954, the Form Configuration 904 requests the latest configuration from the Configuration Service 906. The Configuration Service 906 returns the latest Configuration at step 952. Following this, steps 954 to 959 are repeated as described above.Advantages
[0038] The current methods for creating and editing digital avatars on electronic devices are often cumbersome and inefficient, with complex and time-consuming user interfaces. These existing methods waste user time and device energy, which is particularly problematic for battery-operated devices. This technique presents faster and more efficient methods and interfaces for creating and editing digital avatars on electronic devices. These methods and interfaces can be used to complement or replace other existing methods and can reduce the cognitive load on the user while also being more efficient. Additionally, for battery-operated computing devices, these methods and interfaces conserve power and increase the time between battery charges. The method described includes displaying an avatar navigation user interface on the electronic device's display, detecting a gesture directed to the user interface, and in response to the gesture, displaying an avatar of a first or second type in the avatar navigation user interface. The avatar may be deployed and reconfigured whilst users are interacting with the avatar in other sessions.SUMMARY OF INVENTIONThe invention relates to a computer-implemented method for customizing animation of an interactive digital avatar. The method provides an avatar authoring user interface with a customization interface for modifying animation parameters and a digital avatar animation display showing real-time animations based on these parameters.Users initiate the customization process by sending an update signal, indicating a request to modify one or more animation parameters of the digital avatar. This action triggers the deployment of an avatar session from an avatar cloud server, and a live video stream is received from the server.In real-time, the live video stream is displayed on the digital avatar animation display, updating the digital avatar within the display using the modified animation parameters. The update signal may be received via a graphical user interface or through conversational interaction between the user and the digital avatar.The user interface may also include an input field for content, allowing the digital avatar animation display to deliver content based on the specified delivery field. Additionally, the method supports speech input and output, where the user, through a microphone, can provide a customization request, and the digital avatar, through a speaker, responds with speech output while simultaneously animating on the display.In certain instances, the customization request may include a query for customization options, and the customization response may consist of at least one customization option. The available customization options may depend on the state of the current customization session. The method includes determining whether the customization request meets predefined constraints before customizing the digital avatar. The digital avatar is customized only if these constraints are met.INTERPRETATION
[0039] The methods and systems described may be utilised on any suitable electronic computing system. According to the embodiments described below, an electronic computing system utilises the methodology of the invention using various modules and engines. The electronic computing system may include at least one processor, one or more memory devices or an interface for connection to one or more memory devices, input and output interfaces for connection to external devices in order to enable the system to receive and operate upon instructions from one or more users or external systems, a data bus for internal and external communications between the various components, and a suitable power supply. Further, the electronic computing system may include one or more communication devices (wired or wireless) for communicating with external and internal devices, and one or more input / output devices, such as a display, pointing device, keyboard or printing device. The processor is arranged to perform the steps of a program stored as program instructions within the memory device. The program instructions enable the various methods of performing the invention as described herein to be performed. The program instructions, may be developed or implemented using any suitable software programming language and toolkit, such as, for example, a C-based language and compiler. Further, the program instructions may be stored in any suitable manner such that theycan be transferred to the memory device or read by the processor, such as, for example, being stored on a computer readable medium. The computer readable medium may be any suitable medium for tangibly storing the program instructions, such as, for example, solid state memory, magnetic tape, a compact disc (CD-ROM or CD-R / W), memory card, flash memory, optical disc, magnetic disc or any other suitable computer readable medium. The electronic computing system is arranged to be in communication with data storage systems or devices (for example, external data storage systems or devices) in order to retrieve the relevant data. It will be understood that the system herein described includes one or more elements that are arranged to perform the various functions and methods as described herein. The embodiments herein described are aimed at providing the reader with examples of how various modules and / or engines that make up the elements of the system may be interconnected to enable the functions to be implemented. Further, the embodiments of the description explain, in system related detail, how the steps of the herein described method may be performed. The conceptual diagrams are provided to indicate to the reader how the various data elements are processed at different stages by the various different modules and / or engines. It will be understood that the arrangement and construction of the modules or engines may be adapted accordingly depending on system and user requirements so that various functions may be performed by different modules or engines to those described herein, and that certain modules or engines may be combined into single modules or engines. It will be understood that the modules and / or engines described may be implemented and provided with instructions using any suitable form of technology. For example, the modules or engines may be implemented or created using any suitable software code written in any suitable language, where the code is then compiled to produce an executable program that may be run on any suitable computing system. Alternatively, or in conjunction with the executable program, the modules or engines may be implemented using, any suitable mixture of hardware, firmware and software. For example, portions of the modules may be implemented using an application specific integrated circuit (ASIC), a system-on-a-chip (SoC), field programmable gate arrays (FPGA) or any other suitable adaptable or programmable processing device. The methods described herein may be implemented using a general-purpose computing system specifically programmed to perform the described steps. Alternatively, the methods described herein may be implemented using a specific electronic computer system such as a data sorting and visualisation computer, a database query computer, a graphical analysis computer, a data analysis computer, a manufacturing data analysis computer, a business intelligence computer, an artificial intelligence computer system etc., where the computer has been specifically adapted to perform the described steps on specific data captured from an environment associated with a particular field.
Claims
CLAIMS1. A computer implemented method of customizing animation of an interactive digital avatar, comprising: a) providing an avatar authoring user interface, wherein the avatar authoring user interface includes: i) a customization interface for customizing one or more animation parameters of the digital avatar; ii) a digital avatar animation display for displaying real-time animation of the avatar animated configured using one or more animation parameters; b) receiving an update signal corresponding to a request to update one more animation parameters of the digital avatar with modified animation parameters; c) triggering deployment of an avatar session from an avatar cloud server; d) receiving a live video stream of the avatar cloud server; e) displaying the lived video stream of the digital avatar on the digital avatar animation display to in real time, update the digital avatar within the digital avatar animation display using the modified animation parameters.
2. The method of claim 1 wherein the update signal is received via a graphical user interface.
3. The method of claim 1 wherein the update signal is received via a conversational interaction between the user and the digital avatar.
4. The method of claim 1 wherein the update signal is received from a computer algorithm configured to detect one or more features of the user or the user’s environment.
5. The method of claim 1 wherein the user interface further comprises an input field for content, wherein the digital avatar animation display is configured to deliver content according to the delivery field.
6. The method of claim 3, further comprising: a) receiving, from the user via a microphone of the electronic device, speech input indicating a customization request; andb) providing, by the digital avatar via a speaker of the electronic device, speech output indicating a customization response and simultaneously animating the digital avatar on the display of the electronic device consistently with the speech output.
7. The method of claim 6 wherein the customization request comprises a query for customization options; and wherein the customization response comprises at least one customization option.
8. The method of claim 7, wherein the at least one customization option depends on a state of a current customization session.
9. The method of any one of the preceding claims 6 to 8 further comprising: a) determining whether the customization request meets one or more customization constraints; and b) customizing the digital avatar in accordance with the customization request if, preferably only if, the one or more customization constraints are met.
10. A data processing apparatus or system comprising means for carrying out the method of any one of claims 1-9.
11. A computer program or a computer-readable medium having stored thereon the computer program, the computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of claims 1-10.