Method for constructing robot expression function of vem-token world model
By using the VEM-Token world model and segmenting information lexical units using music and dialogue beats, the problem that NLP-Token cannot understand high-dimensional information modal emotions and expressions is solved, enabling robots to more accurately recognize and communicate social emotions and expressions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD
- Filing Date
- 2025-10-15
- Publication Date
- 2026-04-28
AI Technical Summary
Existing NLP-Token models cannot effectively understand and process emotions and expressions in high-dimensional information modalities, especially the tone of social dialogue, voice, singing, etc., which leads to errors and inaccuracies in AI's emotion and expression recognition and understanding.
Using the VEM-Token world model, information units are segmented through music and dialogue beats to establish vectors, enabling AI to directly understand the operational rules of music and dialogue emotions. Physical or virtual robots are used to recognize and display facial expressions in social situations.
This technology enables robots to more accurately understand and simulate human social emotions and expressions, improving their ability to recognize and communicate emotions and expressions, and is applicable to various intelligent agent applications.
Smart Images

Figure CN120951102B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, specifically to robotics and the subfields of sociology and psychology oriented towards world models, and particularly to robots' ability to understand social emotions and learn social expressions during social interactions. This invention differs from traditional NLP-token systems, which rely on internet text interpretation of facial emotions, verbal dialogue, music, singing, and vocal emotions. This invention proposes a new system—the VEM-fe expression function model—based on the VEM-Token vocal emotion multimodal model for understanding, generating, recognizing, and interacting with robot facial expressions and dialogue expressions, thereby reconstructing the analytical model system. Background Technology
[0002] In the field of artificial intelligence, today's AI is actually unable to understand high-dimensional information such as emotions and expressions. Future AI will need a "world model" to understand the operating rules of the physical world and the operating rules of the social and emotional world, and learn to use emotions and expressions to engage in social interactions.
[0003] As is well known, AI's "understanding" of information to date is still based on the natural language lexical segmentation method using NLP-Tokens (Natural-Language-Processing Tokens). NLP-Tokens have a natural advantage in analyzing and applying low-dimensional textual information modalities. Large models have already learned and memorized all the text-based information from books and web pages. However, for high-dimensional non-textual information modalities, such as the tone and intonation of social conversations, singing, voice, and recitation, and further for even higher-dimensional information modalities such as emotions and facial expressions, today's large models can only resort to the old methods of searching for possible textual descriptions in their memory. This approach of using low-dimensional modalities to process high-dimensional or even higher-dimensional information is clumsy and ridiculous, failing not only in achieving equivalent communication but also in the ability to understand.
[0004] For NLP-Tokens, an information modality based on word units, we cannot know where the large model obtains its corpus from, nor can we predict the accuracy of that corpus. Therefore, the "illusion" and "deviation" of the large model are unavoidable. Furthermore, many information modalities in nature cannot be accurately described directly through language. For example, human emotions, and even animal emotions, according to psychological classifications, include at least the following categories: joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, affection, and enmity. A closer examination of these emotions reveals that:
[0005] (1) The boundaries of emotion classification are not accurately described in words. For example, it is difficult to define the boundary between "joy" and "love" accurately. How do we define this ambiguity and convertibility in words?
[0006] (2) Emotions are spatially interconnected, which is difficult to describe in words. Even if we accept the above psychological classifications, such as "fear," "disgust," and "surprise" being different categories, how do they influence and relate to each other? Although NLP-Tokens can measure the angle and magnitude of vectors in a high-dimensional coordinate system based on the dimension of emotion type, what if there are problems with this coordinate system itself?
[0007] (3) Emotions are continuous analog functions that change over time. Textual descriptions can only be limited slices of this function and cannot describe it in real time. For example, when a song describes "sadness" and "hatred," the singer's emotions will fluctuate as the beat goes by. Moreover, this fluctuation will vary depending on the singer, the environment, and even the number of times the song has been sung. For another example, when Xianglin's wife repeatedly tells the story of Amao being eaten by wolves, sometimes she cries sadly, while at other times she remains calm and does not shed a single tear...
[0008] Following this logic, it's easy to see that even with today's large models learning vast amounts of knowledge, they still cannot track emotions in real time and output accurate "explanations" and "understandings." The root cause of this situation is that the NLP-Token model itself passively allows AI to receive knowledge (such as emotions and expressions) rather than actively understand emotions themselves.
[0009] The invention team has proposed a new set of designs based on the concept of a world model. This design does not adopt the traditional NLP-token model at all, but instead allows AI to directly understand the emotions in music and conversational voices. This series of inventions includes one authorized Chinese invention patent, "VEM-Token Vocal Emotion Multimodal Tokenization Deep Learning Method for Singing and Accompaniment (CN120126506)," one published invention patent, "VEM-Token Beat Capture and Alignment Model Construction Method (CN120748450A)," two invention patents that have passed the preliminary examination, "VEM-Token Vocal Emotion Multimodal Modified Model Construction Method (202511340091.8)" and "VEM-Token Emotion Synchronization Function Layered Fusion Method (202511428558.4)," plus this invention application, which constitute a patent pool of 5 invention patents. This patent pool successfully uses music beats and dialogue beats to segment information tokens, namely VEM-Token (Vocal-Emotion-Multimodal Token), to segment music and dialogue, establish vectors, and thus enable AI to understand music emotions, language and dialogue emotions. This becomes a branch of the world model that enables AI to understand the operating rules of the world of music emotions and dialogue emotions.
[0010] In reality, traditional NLP-token segmentation methods cannot directly process this type of information. They can only understand it by searching the model library for textual descriptions of similar content from web pages and books. Our team's invention, however, is based on a world model approach, enabling AI to understand the workings of the physical world. This completely bypasses NLP-tokens based on text in natural language processing, allowing AI to understand the emotional world of music and dialogue, thus directly comprehending this type of information.
[0011] Furthermore, through careful and in-depth research into real-world social scenarios, the invention team discovered the following characteristics in human facial expressions during social interactions:
[0012] (1) Consistency. The expression of a human face is relatively uniform. That is to say, it is rare for expressions to conflict in both direction, such as crying and laughing. Unless it is a highly skilled actor or the plot is extremely complex, the direction of the expression is generally assumed to be consistent.
[0013] (2) Simplification. In ordinary social situations, such as at a hotel front desk, expressions do not need to be absolutely delicate and rich. A few typical expressions can be listed to meet the needs of the job.
[0014] (3) Ambiguity. In work and social situations, facial expressions do not need to be absolutely precise. As long as the margin of error is not too out of line, it can still meet the needs of the job.
[0015] (4) Compensatory. In social situations, according to psychology, when a person's face is highly attractive or particularly beautiful or handsome, a beautiful / handsome and friendly face is better than a variety of expressions. Therefore, the aesthetic appeal of a robot's face is necessary.
[0016] Based on the above characteristics, in this embodiment, both the micro-expression simulation mode and the real-time expression mode can be designed according to the above four properties. For example, in the micro-expression simulation mode, these four properties can be used to simplify the expression difficulty and reduce computing costs, and the same applies to the real-time expression mode.
[0017] It is important to emphasize that in this patent application, the concept of "space robot" is the same as that of space intelligent robot, embodied robot, and embodied intelligent robot. Unless otherwise stated, it will not be distinguished in this patent application and will be uniformly referred to as "space robot".
[0018] This invention can be applied to the research and application of current large-scale models, such as OpenAI, DeepSeek, GoogleGemini, Kimi, Doubao large-scale model, and Wenxin Yiyan, etc. These can be used for end-to-end applications, or they can be used to form various intelligent agents after connecting to front-end and back-end applications. This invention accesses these large-scale models and achieves bidirectional communication with them to form music / vocal-based artificial intelligence applications, and even music AI intelligent agents, providing strong innovation and support for broadening the application of AI.
[0019] Insufficiency of existing technical methods
[0020] (1) Existing NLP-token-based technologies cannot directly identify and understand the emotions and expressions in music and voice dialogues. They can only use text tokens and existing text in the knowledge base to output annotations in order to indirectly describe emotions.
[0021] (2) Real-world emotions and expressions are not only continuous analog quantities and concepts with fuzzy boundaries, but also time functions that change over time. NLP tokens, through discretized text, cannot accurately describe these analog functions of emotions and expressions that change over time. Summary of the Invention
[0022] Based on the shortcomings of existing technologies, this invention proposes a novel method for constructing VEM-Token world model robot facial expression functions to enable AI to understand the operational rules of the emotional world of music and dialogue. Instead of the traditional NLP-token model, this invention uses physical spatial robots or virtual digital robots in real social situations to understand the emotions and expressions of human social objects through speech, language, and music, and to realize the recognition, display, and communication of simulated human facial and vocal expressions.
[0023] The purpose and intent of this invention are achieved through the following technical solutions and working steps:
[0024] 1. Method for constructing facial expression functions for VEM-Token world model robots
[0025] This invention provides a method for constructing facial expression functions for VEM-Token world model robots, including but not limited to the following steps:
[0026] S1000: Construct a VEM-fe expression function model, including but not limited to: synthesizing one or more expression vectors representing emotional input into a synthetic expression output, aligning expression vectors in the current beat, eliminating expression conflicts between two or more expression vectors, and transitioning from the synthetic expression of the current beat to the synthetic expression of the next beat, enabling the robot to understand emotions and learn expressions.
[0027] S2000: The robot uses a physical expression generator or an animated expression generator to generate expressions. The expression driving model is connected by the VEM-fe expression function model to transmit one or more synthetic expressions and display facial expressions.
[0028] 2. Basic Steps of the VEM-fe Expression Function Model
[0029] Building upon the aforementioned basic scheme, this invention, in terms of the VEM-fe expression function model, includes, but is not limited to, one or more combinations of the following steps or methods:
[0030] S1210: Based on the settings, the VEM-fe expression function model includes, but is not limited to, an input expression vector labeled VEMn(t) and an output synthesized expression labeled FEm(t), where n is the sequence number of the expression vector, m is the sequence number of the synthesized expression, and t is the time sequence value.
[0031] S1220: Define the facial expression transition, including but not limited to the process of the VEM-fe facial expression function model from VEMn(t) to VEMn(t+1), during which the transition time is fet, and generate a facial expression deviation vector labeled FEDT from FEm(t) to FEm(t+1), such that FEDT is less than or equal to the maximum acceptable FED-MAX, which is considered an acceptable transition.
[0032] S1230: Based on the settings, the VEM-fe expression function model includes, but is not limited to, the steps of beat capture, to capture beats starting from time t with a transition length of fet, including beats with a fixed length of fet and beats with a variable length of fet.
[0033] S1240: Expression alignment includes two or more expression vectors or composite expressions, with start and end points aligned on the beat.
[0034] S1250: Expression conflict includes, but is not limited to: In the same beat, when two or more VEMn(t) and VEMp(t) generate FEm(t) and FEq(t), FEDC is marked as an expression conflict vector, such that FEDC is greater than the minimum acceptable conflict threshold FED-MIN, which is considered an expression conflict, where p < n and q < m.
[0035] 3. Extended steps of the VEM-fe expression function model
[0036] Based on the aforementioned solution, the present invention includes, but is not limited to, one or more of the following steps or methods in terms of the extension steps of the VEM-fe expression function model:
[0037] S1310: Based on the settings, the synthesized expressions include, but are not limited to, micro-expression simulation mode and real-time expression mode, specifically:
[0038] The micro-expression simulation mode includes micro-expression components, including but not limited to one or a combination of facial component movement micro-expressions, head movement micro-expressions, and vocal micro-expressions, recorded as MFk, where k is the component or combination number. The vector marking the micro-expression includes but is not limited to direction values and motion values, with direction values ranging from 0 to 360 degrees and motion values ranging from 0 to 1, where the static value during sleep is 0.
[0039] Real-time facial expression patterns, including but not limited to one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hate, affection, enmity, and voice, are recorded as FEj, where j is the number of the corresponding emotion or combination. The degree of facial expression marking ranges from 0 to 1, with 0 being the static sleep scale.
[0040] S1320: Expression library, including but not limited to photos or videos of typical expressions of models of different genders and ages, which are labeled by human experts and trained using supervised learning and reinforcement learning, and also including but not limited to deep learning based on existing expression libraries to train and statistically update expression libraries.
[0041] S1330: 3D mesh model. Collect photos of the front and side views of the head of one or more typical models or users, convert them into a 3D mesh model, and add it to the emoji library.
[0042] S1340: Stability detection. The VEM-fe expression function model employs external and internal stability detection. Specifically, the VEM-fe expression function model uses initial, system-specified, and user-specified periods. For all expression vectors, the values are forced to the maximum value to detect whether the synthesized expression exhibits oscillations. If oscillations are found, the stability detection is deemed to have failed. Alternatively, for all expression vectors, the values are forced to 0 to detect whether the synthesized expression exhibits oscillations. If oscillations are found, the stability detection is deemed to have failed, and a stability detection error message is output.
[0043] S1350: Robustness detection, performed by the VEM-fe expression function model, including but not limited to user input error detection and conflicts arising between two or more input expression vectors. The VEM-fe expression function model outputs a robustness detection error message.
[0044] 4. Emotions in the VEM-fe Expression Function Model
[0045] Based on the aforementioned scheme, the present invention includes, but is not limited to, one or more combinations of the following steps or methods in the VEM-fe expression function model:
[0046] S1410: Employs emotion classifications including but not limited to happiness, sadness, anger, disgust, fear, surprise, and contempt to generate emotion values, which serve as the expression vector input to the VEM-fe expression function model. Specifically, this includes, but is not limited to:
[0047] S1411: Obtain the expression vector. For the vocal, music, recitation, command and dialogue modalities in the emotion classification, use an audio-visual sensor to record audio-visual signals and use the VEM-Token emotion synchronization function to generate one or more data units. The data unit includes, but is not limited to, emotion name, emotion value and emotion weight, as the expression vector.
[0048] S1412: Set a unit social task, including but not limited to a sequence of more than one synthesized emoticons, to complete the emoticon output of a social interaction, including but not limited to the synthesized emoticons of the entire beat of a dialogue or emoticon exchange.
[0049] S1413: Set up continuous social tasks, including but not limited to synthesized expressions for multiple consecutive unit social tasks completed by the robot based on a complete social activity.
[0050] S1414: An emotion generator is used to generate emotion values, which are used as the expression vectors input to the VEM-fe expression function model. The emotion generator includes an editor for emotion names and emotion values provided by a software interface, and a manually operated editor for emotion names and emotion values made of hardware, including a remote control.
[0051] 5. Dialogue emoticons using the VEM-fe emoticon function model
[0052] Based on the aforementioned solution, the present invention includes, but is not limited to, one or more combinations of the following steps or methods in the process of generating dialogue expressions using the VEM-fe expression function model:
[0053] S1510: Uses a microphone to collect speech and employs speech recognition to identify and distinguish more than one social object.
[0054] S1520: Semantic recognition is used to identify the dialogue content of a single unit social task of social interaction, and the VEM-Token emotion synchronization function is used to identify the emotions in the dialogue. The VEM-fe expression function model is then used to identify the synthetic expressions in the dialogue.
[0055] S1530: When the robot engages in one-on-one voice dialogue with a social object, it designs dialogue content based on the dialogue initiated by the social object, establishes a one-to-one pre-dialogue expression spectrum including dialogue content and dialogue expressions, adopts the VEM-fe expression function model, and processes the pre-dialogue expression spectrum including but not limited to synthesizing expressions, expression transitions, beat capture, expression alignment, and expression conflicts to generate the output dialogue expression spectrum. The robot outputs voice with expressions and dialogue content through a sound player, which is a unit social task.
[0056] S1540: When the robot engages in one-to-many voice dialogue with social objects, it designs dialogue content one by one based on the dialogue initiated by the social objects, establishes a one-to-many pre-dialogue expression spectrum including dialogue content and dialogue expressions, adopts the VEM-fe expression function model, and processes the pre-dialogue expression spectrum including but not limited to synthesizing expressions, expression transitions, beat capture, expression alignment, and expression conflicts to generate the output dialogue expression spectrum. The robot outputs voice with expressions and dialogue content through a sound player in multiple unit social tasks.
[0057] S1550: When a robot engages in many-to-many voice dialogue with a social object, multiple robots establish communication, assign a social object to a robot, and complete multiple social tasks by adopting a one-to-many voice dialogue process.
[0058] 6. Real-time expression mode of the entity expression generator
[0059] Based on the aforementioned solution, the present invention includes, but is not limited to, one or more combinations of the following steps or methods in the real-time expression mode of the entity expression generator:
[0060] S2610: Real-time expression mode includes, but is not limited to, a controllable bionic muscle layer and an epidermal layer installed on the robot's head. The controllable bionic muscle layer consists of a control mode and movable bionic muscles, so that facial expressions can be generated in a micro-expression simulation mode or a real-time expression mode under the control of the control mode. The epidermal layer is attached to the surface of the controllable bionic muscle layer and is used to imitate the facial organs of real people to produce expressions and skin color. It also includes, but is not limited to, tongue, teeth, hair, headdress, jewelry and clothing.
[0061] S2620: Control modes include, but are not limited to, electrical signals, force signals, chemical signals and neural signals. It connects to the VEM-fe expression function model through a communication interface to obtain signals for synthesized expressions, driving a controllable bionic muscle layer to display expressions through the epidermis.
[0062] S2630: Real-time expression mode also includes, but is not limited to, bionic eyes with pupils and eyelids, bionic ears, and bionic mouths that can move to generate visual and auditory input and sound output.
[0063] S2640: Bionic eyes include, but are not limited to, dual cameras, dual pupils, and dual eyelids, capable of tracking targets and automatically adjusting the video footage; bionic ears include, but are not limited to, microphones, with binaural sound localization capabilities; bionic mouth includes, but is not limited to, a built-in sound player and mouth movement mechanism; vision, hearing, and sound communicate with and are controlled by the VEM-fe facial expression function model via a communication interface; control steps include, but are not limited to, the motion functions and training results of the facial expression vectors of the bionic eyes, bionic ears, and bionic mouth in the facial expression library.
[0064] S2650: The VEM-fe expression function model extracts control and communication content from the expression library based on the needs of on-site social interaction, and completes the real-time expression mode operation.
[0065] S2660: During the dialogue, based on the dialogue content, continuously acquire facial expression vectors to decompose and assemble them into a sequence of facial expression vectors, obtain synthetic facial expressions of the control modality of real-time facial expression mode, continuously drive the controllable bionic muscle layer, and continuously complete facial expressions and dialogue expressions for unit social tasks and continuous social tasks to complete the dialogue process.
[0066] S2670: Based on user needs, adjust the emotion weight function and VEM-sync emotion vector in the VEM-Token emotion synchronization function to intensify or reduce the emotion in the specified direction, so that the output of the controllable bionic muscle layer is driven to synchronously present the intensification or reduction of the emotion direction.
[0067] S2680: Complete user-specified unit social tasks and continuous social tasks.
[0068] 7. Micro-expression simulation mode of the physical expression generator
[0069] Based on the aforementioned solution, the present invention includes, but is not limited to, one or more combinations of the following model steps or methods in the micro-expression simulation mode of the entity expression generator:
[0070] S2700: The local simulation mode includes decomposing rich human facial expressions into more than one micro-expression that can be realized by micro-expression components, in order to mimic real-time facial expression patterns in typical social situations, including the following:
[0071] S2710: Micro-expressions of the eyes, including but not limited to the movement of components that drive the left and right eyeballs, eyebrows, and eyelids, specifically including but not limited to: tracking eye contact steps where the pupils track the eyes or specific parts of the face of a social object; gazing steps where the pupils track the hands or objects carried by a social object; searching for a sound source where the pupils move to find the location of a sound source; awkward and tense steps where the pupils move away from the social object to find another location; inattentive steps where the pupils track non-social objects or moving objects in the environment; blinking steps where the eyes blink; furrowed brows indicating contemplation and unhappiness; raised brows on one side indicating suspicion; raised brows on both sides indicating surprise and joy; wide-eyed steps with raised eyelids; drooping eyelids or closed eyes; and user-defined steps where the pupil and eyebrow movement trajectories are set according to the user's needs.
[0072] S2720: Mouth micro-expressions, including but not limited to the movement of components that drive the upper and lower lips, tongue, teeth and jaw, specifically including but not limited to: natural pursing of lips, slightly open mouth, open mouth in surprise, blowing, pouting, pursing lips, tightly closed lips, upturned corners of mouth, downturned corners of mouth, biting lips, licking lips, and also includes movements learned and trained from real people's mouths when speaking to drive the generation of mouth micro-expressions.
[0073] S2730: Head micro-expressions, including but not limited to head shaking, nodding, tilting forward, tilting backward, tilting left, tilting right, turning the head left, turning the head right, and tilting the head to the side, as well as movements learned and trained from real people's heads when they speak, which drive the generation of head micro-expressions.
[0074] S2740: Dialogue micro-expressions, including but not limited to, during the dialogue process, continuously acquiring expression vectors based on the dialogue content, decomposing and assembling them into an expression vector sequence to obtain more than one micro-expression simulation mode, continuously driving eye micro-expressions, mouth micro-expressions and head micro-expressions, continuously completing facial expressions and dialogue expressions for unit social tasks and continuous social tasks to complete the dialogue process.
[0075] S2750: Based on user needs, including but not limited to adjusting the emotion weight function and VEM-sync emotion vector in the VEM-Token emotion synchronization function, the emotion can be intensified or reduced in the specified emotion direction, so that the output of eye micro-expressions, mouth micro-expressions, head micro-expressions and dialogue micro-expressions are presented in a synchronized manner with intensification or reduction.
[0076] S2760: Complete user-specified unit social tasks and continuous social tasks.
[0077] 8. Micro-expression simulation mode of the animated expression generator
[0078] Based on the aforementioned solution, the present invention, in the steps of the micro-expression simulation mode of the animated expression generator, further includes, but is not limited to, one or more of the following combined steps or methods:
[0079] S2810: Including but not limited to, based on the Facial Action Coding System (FACS) and the Facial Action Unit (AU) model, taking each AU unit as a micro-expression unit, decomposing the emotion value of the expression vector into the driving parameters of the micro-expression unit of more than one AU unit, and adjusting the driving parameters according to the VEM-fe expression function model and the micro-expression simulation mode.
[0080] S2820: The steps for setting the driving parameters of the micro-expression unit to drive the micro-expression simulation mode to generate two-dimensional and three-dimensional digital human video animations. The driving parameters include, but are not limited to, the steps for modifying the modeling, lighting settings, skin layer rendering and playback of the animation, including but not limited to real-time automatic completion and time-sharing manual adjustment.
[0081] S2830: Drives the micro-expression simulation mode to generate animations of micro-expressions, and combines animations to display facial expressions.
[0082] S2840: Dialogue micro-expressions. During the dialogue process, based on the dialogue expressions in the VEM-fe expression function model and the dialogue content, expression vectors are continuously acquired to decompose and assemble into expression vector sequences to obtain more than one micro-expression simulation mode. The micro-expression simulation mode is continuously driven to generate unit micro-expression animations, and the animations are combined to display facial expressions. The facial expressions and dialogue expressions of unit social tasks and continuous social tasks are continuously completed to complete the dialogue process.
[0083] S2850: Complete user-specified unit social tasks and continuous social tasks.
[0084] 9. Real-time emoji mode of the animated emoji generator
[0085] Based on the aforementioned solution, the present invention further includes, but is not limited to, one or more combinations of the following steps or methods in the real-time expression mode of the animated expression generator:
[0086] S2910: For more than one target industry, collect typical facial expression photos or videos, use the VEM-fe facial expression function model, perform the steps of facial expression library collection, train and generate an industry facial expression library with the characteristics of the target industry, and incorporate it into the facial expression library.
[0087] S2920: Decompose the expressions in the industry expression library into combined micro-expressions one by one according to the proximity principle, and establish an approximate correspondence. The approximate correspondence includes using the expression conflict detection step to calculate the sum of FEDC of all expression sequences in the unit social task, and take the expression sequence with the smallest sum to form the sequence of the closest synthetic expression.
[0088] S2930: Outputs the sequence of synthesized facial expressions to the real-time facial expression mode, becoming combined micro-expressions.
[0089] S2940: Dialogue micro-expressions. During the dialogue, based on the dialogue expressions in the VEM-fe expression function model and the dialogue content, expression vectors are continuously acquired to decompose and assemble into expression vector sequences to obtain more than one micro-expression simulation mode. The micro-expression simulation mode is continuously driven to generate unit micro-expression animations, and the animations are combined to display facial expressions. The facial expressions and dialogue expressions of unit social tasks and continuous social tasks are continuously completed to complete the dialogue process.
[0090] S2950: Facial expressions and dialogue expressions to complete user-specified unit social tasks and continuous social tasks.
[0091] 10. Applications of Space Robots and Digital Robots
[0092] Based on the aforementioned solutions, the present invention specifically includes one or more of the following steps or methods in its application to space robots and digital robots:
[0093] S3010: Connect to a large model, set target industries, including but not limited to hotel front desk, restaurant front desk, bank lobby staff, shopping mall sales assistants, and luxury brand store sales assistants, and distill to obtain semantic text of industry dialogues and social dialogues.
[0094] S3020: Train industry-specific emoji packs and social emoji packs. Input industry text for a specified target industry, and incorporate the trained emoji packs into the local offline emoji library.
[0095] S3030: Designed to support end-to-end lightweight computing hardware, installed in space robot systems or digital robot systems, operating in interactive social micro-expression simulation mode or real-time expression mode.
[0096] 11. Purpose and Intent of the Invention
[0097] The purpose and intent of this invention, which describes a method for constructing facial expression functions for a VEM-Token world model robot, is as follows:
[0098] (1) Transcending the traditional NLP-Token, AI can truly understand the operating rules of the world of music and social dialogue emotions, and realize the basic concept of the world model.
[0099] (2) Enable robots to understand human language and emotions, and to conduct language dialogues with emotions and expressions.
[0100] (3) Enable robots to understand emotions and learn facial expressions, and realize innovative design of auxiliary robots for target industries.
[0101] 12. Beneficial effects of the invention
[0102] (1) Create new ones. From the forward-looking perspective of the world model, enable robots to understand emotions and learn to express expressions, so as to realize the emotional understanding, expression display and synchronous language social interaction of space robots and digital robots.
[0103] (2) Novelty: A brand-new industry robot was designed and implemented through the original VEM-fe expression function model.
[0104] (3) Practicality: For the target industry, a series of industrially practical designs are provided to achieve the purpose of the invention.
[0105] (4) Specifically, the design of the expression function model, the design of the real-time expression of the space robot and the design of the micro-expression mechanism, and the design of the real-time expression of the digital robot and the design of the micro-expression mechanism were implemented.
[0106] (5) Implementation examples of industry robot design are provided. Attached Figure Description
[0107] List of attached images:
[0108] Figure 1 Model system schematic diagram
[0109] Figure 2 Facial Expression Vector Flowchart
[0110] Figure 3 Dialogue Expression Chart
[0111] Figure 4Facial expression-driven diagram
[0112] Figure 5 Singing Expression Chart Diagram
[0113] Detailed description of the attached diagram:
[0114] For detailed descriptions of each figure, please refer to the corresponding figure number analysis in the "Specific Embodiments" section. Detailed Implementation
[0115] This invention application involves a patent pool of 5 related patents, and the invention application itself also possesses independent inventiveness, novelty, and usability. The patent pool includes one authorized Chinese invention patent, "VEM-Token Vocal Emotion Multimodal Tokenization Deep Learning Method for Singing and Accompaniment, CN120126506", one published invention patent application, "VEM-Token Beat Capture and Alignment Model Construction Method, CN120748450A", and two invention patents that have passed preliminary examination, "Construction Method of VEM-Token Vocal Emotion Multimodal Modified Model, 202511340091.8" and "VEM-Token Emotion Synchronization Function Layered Fusion Method, 202511428558.4".
[0116] The objectives and intentions of this invention can be achieved through the following specific embodiments. It should be noted that each specific embodiment has its own specific use and industrial applicability. Therefore, the following embodiments do not encompass all features and steps of this invention, nor do they constitute a limitation thereof. The description in the claims is the summary of the invention.
[0117] Specific embodiments of the present invention are illustrated below:
[0118] An innovative method for constructing facial expression functions for VEM-Token world model robots—a hotel front desk robot that understands emotions and learns facial expressions.
[0119] Diagram Explanation
[0120] This embodiment mainly includes, but is not limited to, the following main schematic diagrams: Figures 1 to 5 .
[0121] Implementation steps
[0122] The method steps in this embodiment mainly include steps 1 to 10. Each of these 10 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not entirely necessary, and their order is not required, but rather selected by the patent implementer based on specific task requirements.
[0123] The specific work steps are explained below:
[0124] 1. Method for constructing facial expression functions for VEM-Token world model robots
[0125] This invention provides a method for constructing facial expression functions for VEM-Token world model robots, including but not limited to the following steps:
[0126] Constructing a VEM-fe expression function model includes, but is not limited to: synthesizing one or more expression vectors representing emotional input into a synthetic expression output; aligning expression vectors in the current beat; eliminating expression conflicts between two or more expression vectors; and transitioning from the synthetic expression of the current beat to the synthetic expression of the next beat, enabling the robot to understand emotions and learn expressions.
[0127] The robot uses a physical expression generator or an animated expression generator to generate expressions. The expression driving model is connected to the VEM-fe expression function model to transmit one or more synthetic expressions and display facial expressions.
[0128] As a typical world model, the independent and core innovative goal of this invention is to enable robots to understand emotions, learn facial expressions, and excel at social interaction, allowing AI to understand the operational rules of the social and emotional world. As a patent pool result of a series of innovative world models, the following models are continued and referenced:
[0129] VEM-Token Vocal Emotion Multimodal Model
[0130] This refers to the model in "VEM-Token Vocal Emotion Multimodal Tokenization Deep Learning Method for Singing and Accompaniment, CN120126506", which specifically includes the following steps 1-4:
[0131] (1) Use one or more modalities to record emotions, label vocal emotion multimodalities as VEM, construct VEM classification, VEM coordinate system, VEM function and VEM library. Vocal emotions include one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hate, affection and enmity. Multimodalities include one or a combination of lyrics, singing, accompaniment, vocal style, music, emotional basis, accompanying instruments, video and image. The VEM coordinate system includes a coordinate axis system established based on independent emotions, opposite emotion pairs and related opposite emotion groups.
[0132] (2) Collect vocal samples according to VEM classification, and have human vocal experts evaluate the vocal samples in terms of emotion in singing and accompaniment. Use supervised learning and deep learning to train the VEM function to obtain VEM parameters and add them to the VEM library.
[0133] (3) The VEM processor is used to calibrate the rhythm of the vocal file and separate the vocal stream and the accompaniment stream. The vocal file is split into VEM-Tokens according to the rhythm, the vocal stream is converted into a VEM-Token1 sequence, the accompaniment stream is converted into a VEM-Token2 sequence, and added to the preprocessing library.
[0134] (4) Using deep learning, the dialogue expression chart, lyrics chart, VEM-Token song chart, VEM-Token accompaniment chart, and VEM-Token music score are generated respectively.
[0135] VEM-Token beat capture and beat alignment model
[0136] This refers to the model in "A Method for Constructing a VEM-Token Beat Capture and Alignment Model, 202511249168.0", which specifically includes the following steps 5-7:
[0137] (5) For vocal files, based on the VEM-Token vocal emotion multimodal model, a beat model including beat capture and beat alignment is set to capture the beat of the vocal file. Based on the beat, the vocal file is divided into VEM-Token sequences, and the starting point and ending point of the beat in each VEM-Token are marked.
[0138] (6) Set the starting point alignment model, including:
[0139] The vocal files, including the sample files and the user files generated by the user imitating the sample files, are divided into VEM-Token1 sequences and VEM-Token2 sequences, respectively. Based on the starting point of each VEM-Token1, a starting point fine-tuning step is used to adjust the starting point of the corresponding VEM-Token2 one by one, so that it is aligned with the starting point of the corresponding VEM-Token1.
[0140] For the dialogue segments and rhythms included in the voice dialogue files in social interactions, taking the first segment as a reference, starting from the second segment, a starting point fine-tuning step is adopted to adjust the starting point of each VEM-Token in each segment one by one, so as to align with the starting point of the VEM-Token at the corresponding position in the first segment, until all loop segments are completed.
[0141] (7) Set the endpoint alignment model, including:
[0142] Based on the endpoint of each VEM-Token1, a fine-tuning step is used to adjust the endpoint of the corresponding VEM-Token2 one by one, so that it is aligned with the endpoint of the corresponding VEM-Token1.
[0143] For each segment of the dialogue, taking the first segment as a reference, starting from the second segment, a fine-tuning step is used to adjust the endpoint of each VEM-Token in each segment one by one, so that it is aligned with the endpoint of the VEM-Token at the corresponding position in the first segment, until all loop segments are completed.
[0144] VEM-Token Vocal Emotion Multimodal Modified Model
[0145] This refers to the model in "Construction Method of VEM-Token Vocal Emotion Multimodal Modification Model, 202511340091.8", which specifically includes the following steps 8 and 9:
[0146] (8) Collect sample files and user files. Based on the VEM-Token model, use beat capture and VEM-Token segmentation to obtain the VEM-Token1 sequence of the sample files and the VEM-Token2 sequence of the user files respectively. Based on the VEM-Token1 sequence, use beat alignment for all VEM-Token2 sequences to generate the VEM-Token2 sequence.
[0147] (9) Based on the VEM parameters included in the VEM-Token model, identify the VEM parameters of the VEM-Token1 sequence, determine the modification scheme by the user, process the VEM parameters of the VEM-Token2 sequence, and modify and generate the VEM-Token2 sequence that matches the style of the VEM-Token1 sequence.
[0148] A Layered Fusion Method for VEM-Token Sentiment Synchronization Function
[0149] This refers to the model in "A Method for Layered Fusion of VEM-Token Sentiment Synchronization Function, 202511428558.4", which specifically includes the following steps 10, 11, and 12:
[0150] (10) The VEM-Token vocal emotion multimodal model is adopted. The song file is divided into VEM-Token sequences in units of beats. The emotion synchronization function is set as VEM-sync. VEM-sync includes a sequence of VEM-sync vectors that correspond to and are aligned with the VEM-Token sequence. The VEM-sync vector includes synchronization content and a synchronization pointer pointing to the corresponding VEM-Token beat.
[0151] (11) Based on the VEM-Token sequence, calculate and obtain the synchronization content of the VEM-sync vector sequence. The synchronization content includes beat attributes and emotion attributes, where the emotion attributes include more than one emotion name, emotion value, and emotion weight.
[0152] (12) The emotion weight function includes: using hierarchical processing of song files to obtain the corresponding VEM-sync vector sequence, and using one or a combination of multi-layer weighted scanning, recurrent neural network, long short-term memory network, self-attention mechanism and retrieval enhancement to generate forward propagation, backward propagation and omnidirectional propagation to obtain the emotion synchronization function with the song file.
[0153] While this invention application references and draws inspiration from the ideas of these four patents, it is also an independent invention. As a typical world model, the independent and core innovative goal of this invention is to enable robots to understand emotions, learn facial expressions, and be adept at social interaction, allowing AI to understand the operational rules of the social and emotional world.
[0154] Figure 1 This is a schematic diagram of the system model of the present invention, and a diagram of the VEM-fe expression function model, explained as follows:
[0155] exist Figure 1 On the left are several facial expression vectors input to the facial expression vector synthesis module. It's important to note that this synthesis is based on classification. For example, regarding the eye movement modality, it might involve facial expression vectors related to nearby sounds, the arrival of guests at a hotel reception, or the speaking of guests. Therefore, these various facial expression vectors need to be considered together, i.e., "synthesized." Furthermore, these facial expression vectors may also become inputs for other modalities, such as head rotation. Therefore, the relationship between the input facial expression vectors and the output synthesized vectors is not a one-to-one correspondence, but a many-to-many mapping. It's worth noting that within each module, the beat division of the facial expression vectors can be set, and beat markers can be added. For details, refer to the model generation of beat markers in "A Method for Constructing a VEM-Token Beat Capture and Alignment Model, 202511249168.0". Users of this patent can also design their own beat division methods.
[0156] From synthesizing expression vectors to aligning expression vectors, from aligning expression vectors to eliminating expression conflicts, and from eliminating expression conflicts to transitioning to synthesized expressions, refer to steps 2, 3, 4, and 5 below to complete the corresponding step formulas. After generating the synthesized expression, submit it to either the physical expression generator in spatial robot mode or the animated expression generator in digital robot mode, using micro-expression simulation or real-time expressions, and output it to the spatial robot or real-time expression generator respectively to display the expression. Here, the spatial robot is the head and face of a physical robot used to display the expression; the area below the neck can be connected by the spatial robot. The real-time expression is displayed as an animation on a monitor.
[0157] 2. Basic Steps of the VEM-fe Expression Function Model
[0158] Building upon the aforementioned basic scheme, this invention, in terms of the VEM-fe expression function model, includes, but is not limited to, one or more combinations of the following steps or methods:
[0159] Furthermore, based on the settings, the VEM-fe expression function model includes, but is not limited to, an input expression vector labeled VEMn(t) and an output synthesized expression labeled FEm(t), where n is the sequence number of the expression vector, m is the sequence number of the synthesized expression, and t is the time sequence value.
[0160] Preferably, the expression transition is set to include, but is not limited to, the process of the VEM-fe expression function model from VEMn(t) to VEMn(t+1) with a transition time of fet, generating an expression deviation vector labeled FEDT from FEm(t) to FEm(t+1), such that FEDT is less than or equal to the maximum acceptable FED-MAX, thus becoming an acceptable transition.
[0161] The transition of facial expressions here is because the transition from one expression to the next requires a psychologically consistent process, such as from smiling to crying. This is defined as an expression deviation vector. If this vector exceeds the normal range in actual social interactions, it does not conform to normal social dynamics. Furthermore, this transition must have a natural grace period (fet), which cannot be completed in a very short time. In this invention, this expression deviation vector and natural grace period are initially determined through supervised learning and reinforcement learning, under the guidance of human experts.
[0162] Furthermore, based on the settings, the VEM-fe expression function model includes, but is not limited to, the step of beat capture, to capture beats starting from time t with a transition length of fet, including beats with a fixed length of fet and beats with a variable length of fet.
[0163] The rhythm here refers to the conversation or social actions between one party and another in a social process, which can be decomposed. This invention adopts the concept of rhythm decomposition.
[0164] Preferably, the expression alignment includes two or more expression vectors or synthesized expressions, with the start point and end point aligned on the beat.
[0165] Preferably, the expression conflict includes, but is not limited to: in the same beat, when two or more VEMn(t) and VEMp(t) generate FEm(t) and FEq(t), FEDC is marked as the expression conflict vector, such that FEDC is greater than the minimum acceptable conflict threshold FED-MIN, which is considered as an expression conflict, where p < n and q < m.
[0166] The facial expression alignment here involves factors such as the uncertainty of model computation time due to different input facial expression vectors, and the uncertainty of psychological and facial responses from a social psychology perspective. For example, when a front desk robot is talking to a customer, facial expression vector 1 is "a customer" and facial expression vector 2 is "inconvenient to move". The optimal result for these two facial expression vectors might be "the robot comes out to help the customer and says: Please slow down, let me help you", rather than "say: Hello". Since the model output produces two conflicting results, it is necessary to compare these two results and then make a decision on which facial expression to execute. Therefore, both facial expression alignment and facial expression conflict decision-making are essential.
[0167] In this invention, the facial expression vector that serves as the input to the VEM-fe facial expression function model is actually the output vector of the previous model. For example, when connecting to a robot remote control, this facial expression vector is a block of instructions generated by the remote control. When connecting to the aforementioned VEM model, this facial expression vector is an emotion vector, including vocal emotion, dialogue emotion, emotion synchronization function, etc.
[0168] 3. Extended steps of the VEM-fe expression function model
[0169] Based on the aforementioned solution, the present invention includes, but is not limited to, one or more of the following steps or methods in terms of the extension steps of the VEM-fe expression function model:
[0170] Furthermore, based on the settings, the synthesized expressions include, but are not limited to, micro-expression simulation mode and real-time expression mode, specifically:
[0171] Preferably, the micro-expression simulation mode includes micro-expression components, including but not limited to one or a combination of facial component movement micro-expressions, head movement micro-expressions, and vocal micro-expressions, recorded as MFk, where k is the component or combination number, and the vector marking the micro-expression includes but is not limited to direction values and motion values, with direction values ranging from 0 to 360 degrees and motion values ranging from 0 to 1, wherein the static value during sleep is 0.
[0172] Preferably, the real-time facial expression mode includes, but is not limited to, one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hate, affection, enmity, and voice, which are recorded as FEj, where j is the number of the corresponding emotion or combination, and the degree of facial expression marking ranges from 0 to 1, with 0 being the static sleep scale.
[0173] The invention team's research revealed the following characteristics of facial expressions:
[0174] (1) Consistency. The expression of a human face is relatively uniform. That is to say, it is rare for a person to both cry and laugh. Unless it is a highly skilled actor or the plot is extremely complex, the expression is generally considered to be consistent.
[0175] (2) Simplification. In ordinary social situations, such as at a hotel front desk, expressions do not need to be absolutely delicate and rich. A few typical expressions can be listed to meet the needs of the job.
[0176] (3) Ambiguity. In work and social situations, facial expressions do not need to be absolutely precise. As long as the margin of error is not too large, it can meet the needs of the job.
[0177] (4) Compensatory. In social situations, according to psychology, when a person's face is highly attractive or particularly beautiful / handsome, a beautiful / handsome and friendly face is better than a variety of expressions. Therefore, the aesthetic appeal of a robot's face is necessary.
[0178] Based on the above characteristics, in this embodiment, both the micro-expression simulation mode and the real-time expression mode can be designed according to actual applications. For example, in the micro-expression simulation mode, these four characteristics can be used to simplify the expression difficulty and reduce costs, and the same applies to the real-time expression mode.
[0179] Furthermore, the expression library includes, but is not limited to, collecting typical expression photos or videos of models of different genders and ages, having human experts label the expressions, and training the expression library using supervised learning and reinforcement learning. It also includes, but is not limited to, using deep learning based on existing expression libraries to train and statistically update the expression library.
[0180] In the initial stages, the facial expression database was acquired beforehand through supervised learning and reinforcement learning, under the guidance of human experts. Subsequently, it can be upgraded synchronously during model runtime.
[0181] Preferably, a three-dimensional mesh model is created by collecting photos of the front and side views of the head of one or more typical models or users, converting them into a three-dimensional mesh model, and adding it to the emoji library.
[0182] This is especially useful in the design of animated emoticons, where clients can swap faces.
[0183] Preferably, stability detection is performed using the VEM-fe expression function model, employing both external and internal stability checks. Specifically, the VEM-fe expression function model uses checks for the initial stage, system-specified periods, and user-specified periods. For all expression vectors, the values are forced to the maximum value, and the synthesized expression is checked for oscillations. If oscillations are found, the stability detection is deemed to have failed. Alternatively, for all expression vectors, the values are forced to 0, and the synthesized expression is checked for oscillations. If oscillations are found, the stability detection is deemed to have failed, and a stability detection error message is output.
[0184] Preferably, robustness detection is performed by the VEM-fe expression function model, including but not limited to user input error detection and conflicts arising between two or more input expression vectors. The VEM-fe expression function model outputs a robustness detection error message.
[0185] Stability and robustness testing are particularly important for models in both closed-loop and open-loop control, as they can reduce "illusions" similar to those in large-scale model dialogues, and minimize facial expression drift and distortion.
[0186] 4. Emotions in the VEM-fe Expression Function Model
[0187] Based on the aforementioned solutions, this invention, in terms of emotion classification, emotion vectors, and unit social tasks of model input, includes, but is not limited to, one or more combinations of the following steps or methods:
[0188] Furthermore, emotion classifications, including but not limited to happiness, sadness, anger, disgust, fear, surprise, and contempt, are used to generate emotion values, which serve as the expression vectors input to the VEM-fe expression function model. Specifically, these include, but are not limited to:
[0189] Preferably, an expression vector is obtained. For the vocal, music, recitation, command, and dialogue modalities in the emotion classification, audio-visual signals are recorded using an audio-visual sensor, and one or more data units are generated using the VEM-Token emotion synchronization function. These data units include, but are not limited to, emotion name, emotion value, and emotion weight, and are used as expression vectors.
[0190] In this embodiment, the audio and video sensors are used to collect video, images and sound using the robot's eyes and ears, respectively. Furthermore, the VEM-Token emotion synchronization function software generates emotion names, emotion values and emotion weights as expression vectors.
[0191] Preferably, a unit social task is set, including but not limited to a sequence of one or more synthetic emoticons, to complete the emoticon output of a social interaction, including but not limited to the synthetic emoticons of all beats of a dialogue or an emoticon exchange.
[0192] Preferably, continuous social tasks are set, including but not limited to synthesized emoticons that are completed by the robot on multiple consecutive units of social tasks based on a complete social activity.
[0193] Preferably, an emotion generator is used to generate emotion values, which are used as the expression vectors input to the VEM-fe expression function model. The emotion generator consists of an editor for emotion names and emotion values provided by a software interface, and a manually operated editor for emotion names and emotion values made of hardware, including a remote control.
[0194] In this embodiment, which is limited to a robot engaged in the service industry at the hotel front desk, the emotion classification needs to weaken or even eliminate the emotion classifications of sadness, anger, disgust, fear, surprise, and contempt, and focus on strengthening the emotion of happiness and the sub-emotion classifications.
[0195] Figure 2 This is a flowchart illustrating the expression vector process, explained below:
[0196] Figure 2 From left to right, the diagram illustrates the process by which the robot acquires audio and video information from the scene. First, the robot's camera captures a video image stream. The facial expression recognition module then obtains the sequence of facial images of the social object (in this example, a potential booking customer), and outputs it to the VEM-Token emotion synchronization function. This function captures facial expressions, obtains facial expression vectors, and outputs them to the VEM-fe expression function model. Simultaneously, the robot's ear (microphone) captures an audio stream. The speech and facial expression recognition module obtains the speech sequence of the social object (in this example, a potential booking customer), and outputs it to the VEM-Token emotion synchronization function. This function captures speech and facial expressions, obtains speech and facial expression vectors, and outputs them to the VEM-fe expression function model. In the VEM-fe expression function model, facial expression vectors and speech and facial expression vectors are synthesized. These vectors undergo expression beat segmentation, expression beat alignment, expression conflict elimination, and transitional synthesis to produce the final synthesized expression output.
[0197] 5. Dialogue emoticons using the VEM-fe emoticon function model
[0198] Based on the aforementioned solutions, the present invention, in terms of emoticons for social interaction dialogues, includes, but is not limited to, one or more combinations of the following steps or methods:
[0199] Furthermore, a microphone is used to collect speech, and speech recognition is employed to identify and differentiate more than one social object.
[0200] Preferably, semantic recognition is used to identify the dialogue content of a single unit social task of social object dialogue, the VEM-Token emotion synchronization function is used to identify the dialogue emotion, and the VEM-fe expression function model is input to identify the synthetic expression of the dialogue.
[0201] Preferably, when the robot engages in one-on-one voice dialogue with a social object, the dialogue content is designed based on the dialogue initiated by the social object, and a one-to-one pre-dialogue expression spectrum including dialogue content and dialogue expressions is established. The VEM-fe expression function model is adopted to process the pre-dialogue expression spectrum, including but not limited to synthesizing expressions, expression transitions, beat capture, expression alignment, and expression conflicts, to generate the output dialogue expression spectrum. The robot outputs voice with expressions and dialogue content through a sound player, which is a single unit social task.
[0202] Preferably, when the robot engages in one-to-many voice dialogue with social objects, dialogue content is designed one by one based on the dialogue initiated by the social objects, and a one-to-many pre-dialogue expression spectrum including dialogue content and dialogue expressions is established. The VEM-fe expression function model is adopted to process the pre-dialogue expression spectrum, including but not limited to synthesizing expressions, expression transitions, beat capture, expression alignment, and expression conflicts, to generate the output dialogue expression spectrum. The robot outputs voice with expressions and dialogue content through a sound player in multiple unit social tasks.
[0203] Preferably, when a robot engages in a many-to-many voice dialogue with a social object, multiple robots establish communication, assign a social object to a robot, and complete multiple social tasks by adopting a one-to-many voice dialogue process.
[0204] In this embodiment, the front-end is set as a robot whose work is social, requiring voice dialogue and facial expression display, and needs one-to-one and one-to-many modes.
[0205] Figure 3 It is a dialogue emoticon chart, and also the musical score proposed in this invention application. Figure 3 In the diagram, the vertical blue dashed lines represent beat lines, with one beat between two beat lines representing a single beat in the dialogue. The black horizontal lines from top to bottom are modal segmentation lines. The first modal segmentation line represents "I," the robot itself, and its dialogue content is displayed from left to right, highlighted in red. The second modal segmentation line represents the guest, specifically the guest in the hotel front desk scene of this example, and their dialogue content is displayed from left to right, highlighted in green. Subsequent modal segmentation lines represent: cue words, tone of voice, eyebrows, eyes, mouth, and head, with red text on each line indicating the modal annotation. All of this content represents instructions generated by the VEM-fe expression function model based on artificial intelligence steps, executed by the robot to produce facial expressions and dialogue expressions.
[0206] 6. The robot's real-time facial expressions
[0207] Based on the aforementioned solution, this invention employs a real-time expression mode of a physical expression generator, including but not limited to one or more of the following steps or methods:
[0208] Furthermore, the real-time expression mode includes, but is not limited to, a controllable bionic muscle layer and an epidermal layer installed on the robot's head. The controllable bionic muscle layer consists of a control mode and movable bionic muscles, so that facial expressions can be generated in a micro-expression simulation mode or a real-time expression mode under the control of the control mode. The epidermal layer is attached to the surface of the controllable bionic muscle layer and is used to imitate the facial organs of real people to produce expressions and skin color. It also includes, but is not limited to, tongue, teeth, hair, headdress, jewelry and clothing.
[0209] Preferably, the control modality includes, but is not limited to, electrical signals, force signals, chemical signals, and neural signals. The signal for synthesizing facial expressions is obtained by connecting to the VEM-fe facial expression function model through a communication interface, which drives the controllable bionic muscle layer to display facial expressions through the epidermis.
[0210] Preferably, the real-time expression mode also includes, but is not limited to, bionic eyes with pupils and eyelids, bionic ears, and bionic mouths that can move to generate visual and auditory input and sound output.
[0211] Preferably, the bionic eyes include, but are not limited to, dual cameras, dual pupils, and dual eyelids, capable of tracking targets and automatically adjusting the video footage. The bionic ears include, but are not limited to, microphones, with binaural sound localization capabilities. The bionic mouth includes, but is not limited to, a built-in sound player and mouth movement mechanism. Vision, hearing, and sound communicate with and are controlled by the VEM-fe expression function model via a communication interface. The control steps include, but are not limited to, the motion functions and training results of the expression vectors of the bionic eyes, ears, and mouth in the expression library.
[0212] The automatic adjustments here include adjusting the focal length, aperture, searching for and focusing on social objects. The bionic eye also includes cameras with different functions, such as one for recording video and one for ranging, and can also be designed for binocular parallax ranging.
[0213] Preferably, the VEM-fe expression function model extracts control and communication content from the expression library according to the needs of on-site social interaction, and completes the real-time expression mode operation.
[0214] During the dialogue, based on the dialogue content, facial expression vectors are continuously acquired, decomposed and assembled into a sequence of facial expression vectors, and synthetic facial expressions of the control modality of real-time facial expression mode are obtained. This continuously drives the controllable bionic muscle layer to continuously complete facial expressions and dialogue expressions for unit social tasks and continuous social tasks, thereby completing the dialogue process.
[0215] Based on user needs, the emotion weight function and VEM-sync emotion vector in the VEM-Token emotion synchronization function are adjusted to intensify or reduce the emotion in the specified direction, thereby driving the output of the controllable bionic muscle layer to synchronously present the intensification or reduction of the emotion in the specified direction.
[0216] Complete user-specified unit social tasks and continuous social tasks.
[0217] This invention uses a space robot to perform tasks at the hotel front desk. However, it should be emphasized that the design focuses on the robot's head, not its entire body.
[0218] For the space robot's head, a 3-axis rotation is recommended, using the spine connecting the head and shoulders as the base point. The three axes are head rotation, left-right swaying, and forward / backward nodding. The eyes, including both eyeballs, are recommended to use a 2-axis design, including left-right and up-down rotation. The eyebrows are recommended to move up and down. The eyelids are recommended to move in tandem with the eyeballs. The mouth is recommended to extend up and down and left and right. The chin is recommended to use a single-axis design with up-down movement. The face is recommended to be able to move up and down and left and right. The movement of all the head's organs should be supported by a trained facial expression database. In particular, for the robot's speech, the corresponding organs need to move synchronously.
[0219] It's important to note that for the VEM-fe expression function model, the expression database contains previously annotated content by human vocal experts and expression videos categorized for typical emotions, which are then used for training. Furthermore, for speech playback, the mouth movement patterns during the pronunciation of typical words, trained on this method, are incorporated into the expression database.
[0220] Figure 4 This is a schematic diagram of the facial expression driving mechanism, as shown in the diagram when the facial expressions are displayed on the terminal according to this invention. The output of the VEM-fe facial expression function model is a synthesized facial expression, which is connected to a space robot or digital robot via driving circuitry to drive the robot to display the expression. It is important to note that the driving circuitry varies depending on the type of robot. For example, if the space robot is a solid structure robot, such as an electric, hydraulic, or pneumatic robot, then the driving circuitry consists of electrical circuits, hydraulic circuits, and pneumatic circuits. If it is a digital robot, then the driving circuitry consists of the driving parameters of the animation software.
[0221] 7. Micro-expression simulation of robots
[0222] Based on the aforementioned solution, the present invention includes, but is not limited to, one or more combinations of the following model steps or methods in the micro-expression simulation mode of the entity expression generator:
[0223] Furthermore, the local simulation mode includes decomposing the rich human facial expressions into one or more micro-expressions that can be realized by micro-expression components, in order to mimic real-time facial expression patterns in typical social situations, including the following:
[0224] Preferably, the micro-expressions of the eyes include, but are not limited to, the movement of components that drive the movement of the left and right eyeballs, eyebrows, and eyelids. Specifically, this includes, but is not limited to: a tracking eye contact step where the pupils track the eyes or specific parts of the face of a social subject; a gazing step where the pupils track the hands or objects carried by a social subject; a sound-finding step where the pupils move to find the location of a sound source; an awkward or tense step where the pupils move away from the social subject and find another location; a nonchalant step where the pupils track non-social subjects or moving objects in the environment; a blinking step where the eyes blink; a furrowed brow step indicating contemplation or unhappiness; a raised eyebrow step indicating suspicion; a raised eyebrow step indicating surprise or joy; a wide-eyed step with raised eyelids; a drooping eyelid or closed eye step; and user-defined steps where the pupil and eyebrow movement trajectories are set according to the user's needs.
[0225] Preferably, the micro-expressions of the mouth include, but are not limited to, the movement of components that drive the upper and lower lips, tongue, teeth and jaw, specifically including but not limited to: natural pursing of the lips, slightly opening the mouth, opening the mouth in surprise, blowing air, pouting, pursing the lips, tightly closing the lips, raising the corners of the mouth, lowering the corners of the mouth, biting the lips, licking the lips, and also includes movements learned and trained from the mouth when speaking to drive the generation of micro-expressions of the mouth.
[0226] Preferably, head micro-expressions include, but are not limited to, head shaking, nodding, tilting forward, tilting backward, tilting left, tilting right, turning the head left, turning the head right, and tilting the head to the side, as well as movements learned and trained from real people's heads while speaking to drive the generation of head micro-expressions.
[0227] Preferably, the dialogue micro-expressions include, but are not limited to, continuously acquiring expression vectors based on the dialogue content during the dialogue process, decomposing and assembling them into an expression vector sequence to obtain more than one micro-expression simulation mode, continuously driving eye micro-expressions, mouth micro-expressions and head micro-expressions, continuously completing facial expressions and dialogue expressions for unit social tasks and continuous social tasks, so as to complete the dialogue process.
[0228] Furthermore, based on user needs, including but not limited to adjusting the emotion weight function and VEM-sync emotion vector in the VEM-Token emotion synchronization function, the intensity can be increased or decreased in the specified emotion direction, so that the output of eye micro-expressions, mouth micro-expressions, head micro-expressions and dialogue micro-expressions are simultaneously intensified or decreased.
[0229] Preferably, the user-specified unit social tasks and continuous social tasks are completed.
[0230] It's important to note that the micro-expression simulation mode here refers to statistical analysis and summarization of facial expressions used in industry-specific (e.g., front desk staff) interactions. This aims to refine and extract several typical social expression patterns that existing micro-expression technology can mimic, rather than a complete simulation of real-life expressions. Through this refinement, it's possible to achieve a near-perfect imitation of industry-specific social interactions within the existing micro-expression capabilities. In actual social activities, by simulating typical industry-specific social expressions and combining them with social rhythms, it can achieve an effect similar to real-life expressions.
[0231] In addition, micro-expressions can be categorized at the business level of industry-specific social networking, so that they can be categorized and invoked when used.
[0232] 8. Micro-expression simulation in animation
[0233] Based on the aforementioned solution, the present invention further includes, but is not limited to, one or more combinations of the following steps or methods in the micro-expression simulation mode of the animated expression generator:
[0234] Furthermore, including but not limited to, based on the Facial Action Coding System (FACS) and the Facial Action Unit (AU) model, each AU unit is considered as a micro-expression unit, and the emotion value of the decomposed expression vector is used as the driving parameter of the micro-expression unit of one or more AU units. Based on the VEM-fe expression function model and the micro-expression simulation mode, the driving parameters are adjusted.
[0235] Preferably, the steps of setting the driving parameters of the micro-expression unit to drive the micro-expression simulation mode to generate two-dimensional and three-dimensional digital human video animations include, but are not limited to, modifying the modeling, lighting settings, skin layer rendering and playback of the animation, including but not limited to real-time automatic completion and time-sharing manual adjustment.
[0236] Preferably, the micro-expression simulation mode generates animations of micro-expressions in the unit, and the animations are combined to display facial expressions.
[0237] Preferably, in the dialogue micro-expression, during the dialogue process, based on the dialogue expressions in the VEM-fe expression function model and the dialogue content, expression vectors are continuously acquired to decompose and assemble into expression vector sequences to obtain more than one micro-expression simulation mode. The micro-expression simulation mode is continuously driven to generate animations of unit micro-expressions, and the animations are combined to display facial expressions. The facial expressions and dialogue expressions of unit social tasks and continuous social tasks are continuously completed to complete the dialogue process.
[0238] Preferably, the user-specified unit social tasks and continuous social tasks are completed.
[0239] It's important to note that since facial expressions are achieved using animation, such as 3D or 2D animation, the difficulty of achieving these expressions differs from that of physical spatial robots. Therefore, the introduction of micro-expressions is not for simplification or refinement; the users of this patent aim to make the expressions as delicate and realistic as possible. However, the industry-specific social characteristics of facial expressions still exist, and the micro-expression simulation here refers to the simulation of these industry-specific social characteristics, in order to achieve an effect as similar as possible to real human expressions.
[0240] 9. Real-time animated emoticons
[0241] Based on the aforementioned solution, the present invention further includes, but is not limited to, one or more combinations of the following steps or methods in the real-time expression mode of the animated expression generator:
[0242] Furthermore, for more than one target industry, typical facial expression photos or videos are collected statistically. Using the VEM-fe facial expression function model, the steps of facial expression library collection are performed to train and generate an industry facial expression library with the characteristics of the target industry, which is then incorporated into the facial expression library.
[0243] Preferably, the expressions in the industry expression library are decomposed into combined micro-expressions one by one according to the proximity principle to establish an approximate correspondence. The approximate correspondence includes using an expression conflict detection step to calculate the sum of the FEDC of all expression sequences in the unit social task, and taking the expression sequence with the smallest sum to form the sequence of the closest synthetic expression.
[0244] Preferably, the sequence of synthesized facial expressions is output to the real-time facial expression mode to become a combined micro-expression.
[0245] Preferably, in the dialogue micro-expression, during the dialogue process, based on the dialogue expressions in the VEM-fe expression function model and the dialogue content, expression vectors are continuously acquired to decompose and assemble into expression vector sequences to obtain more than one micro-expression simulation mode. The micro-expression simulation mode is continuously driven to generate animations of unit micro-expressions, and the animations are combined to display facial expressions. The facial expressions and dialogue expressions of unit social tasks and continuous social tasks are continuously completed to complete the dialogue process.
[0246] Preferably, facial expressions and dialogue expressions are used to complete user-specified unit social tasks and continuous social tasks.
[0247] It is important to note that the micro-expressions in dialogue here include not only the emotional expressions conveyed through the voice during the conversation, but also the simultaneous facial expressions.
[0248] 10. Applications of Space Robots and Digital Robots
[0249] Based on the aforementioned solutions, the present invention specifically includes one or more of the following steps or methods in its application to space robots and digital robots:
[0250] By accessing a large model and setting target industries, including but not limited to hotel front desk staff, restaurant front desk staff, bank lobby staff, shopping mall sales assistants, and luxury brand store sales assistants, semantic text of industry dialogues and social dialogues can be obtained through distillation.
[0251] Preferably, industry-specific emojis and social emojis are trained by inputting industry-specific text for a specified target industry, and then incorporated into a local offline emoji library.
[0252] Preferably, the design supports end-to-end lightweight computing hardware, which is installed in a space robot system or a digital robot system and operates in an interactive social micro-expression simulation mode or a real-time expression mode.
[0253] Figure 5 This is a schematic diagram of singing expression patterns. As a specific example of this embodiment, the robot of this invention, while performing social dialogue, can also complete singing expressions, dialogue expressions, and facial expressions in the context of music and singing. Furthermore, it can perform multimodal recognition for guests' music, singing, and other multimodal expressions, in a similar manner.
Claims
1. A method for constructing facial expression functions for a VEM-Token world model robot, characterized in that, include: S1000: Constructing the VEM-fe expression function model, including: synthesizing one or more expression vectors representing emotional input into the output synthetic expression, aligning expression vectors in the current beat, eliminating expression conflicts between two or more expression vectors, and transitioning from the synthetic expression of the current beat to the synthetic expression of the next beat, so that the robot can understand emotions and learn expressions; S2000: The robot uses a physical expression generator or an animated expression generator to generate expressions. The expression driving model is connected by the VEM-fe expression function model to transmit one or more synthetic expressions and display facial expressions.
2. The method according to claim 1, characterized in that, The VEM-fe expression function model includes: S1210: Based on the settings, the VEM-fe expression function model includes those labeled as VEM. n(t) The input expression vector and the label FE m(t) The output is the synthesized facial expression, where n is the sequence number of the facial expression vector, m is the sequence number of the synthesized facial expression, and t is the time sequence value; S1220: Setting facial expression transitions includes the VEM-fe facial expression function model derived from VEM. n(t) To VEM n(t+1) During the process, the transition time is fet, and the process is generated by FE. m(t) To FE m(t+1) The expression deviation vector labeled FEDT is such that FEDT is less than or equal to the maximum acceptable FED-MAX, thus becoming an acceptable transition; S1230: Based on the settings, the VEM-fe expression function model includes a beat capture step to capture beats starting from time t with a transition time of length fet. The beats include beats with a fixed length fet and beats with a variable length fet. S1240: Expression alignment includes the start and end points of two or more expression vectors or composite expressions, aligned on the beat. S1250: Expression conflict includes: when two or more VEMs are in the same beat. n(t) VEM p(t) FE generation m(t) and FE q(t) Then, FEDC is labeled as the expression conflict vector. If FEDC is greater than the minimum acceptable conflict threshold FED-MIN, it is considered an expression conflict, where p < n and q < m.
3. The method according to claim 2, characterized in that, The VEM-fe expression function model also includes: S1310: Based on the settings, the synthesized expressions include a micro-expression simulation mode and a real-time expression mode, specifically: The micro-expression simulation mode includes one or a combination of facial component movement micro-expressions, head movement micro-expressions, and vocal micro-expressions, recorded as MFk, where k is the number of the component or combination. The vector marking the micro-expression includes direction value and motion value, with the direction value ranging from 0 to 360 degrees and the motion value scale from 0 to 1, where the static state during sleep is 0. Real-time facial expression patterns include one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hate, affection, and enmity, which are recorded as FEj, where j is the number of the corresponding emotion or combination. The degree of facial expression marking ranges from 0 to 1, with 0 being the static sleep scale. S1320: Expression library, which includes photos or videos of typical expressions of models of different genders and ages, which are labeled by human experts and trained using supervised learning and reinforcement learning. It also includes a statistically updated expression library trained by deep learning based on the existing expression library. S1330: A 3D mesh model is created by collecting frontal and side-view photos of the head of one or more typical models or users, converting them into a 3D mesh model, and adding it to the emoji library; and / or, S1340: Stability detection, which uses external and internal stability detection by the VEM-fe expression function model. Specifically, the VEM-fe expression function model uses the initial stage, system-specified stage, and user-specified stage. For all expression vectors, the value is forced to the maximum value to detect whether the synthesized expression oscillates. If there is oscillation, the stability detection is determined to be wrong. For all expression vectors, the value is forced to 0 to detect whether the synthesized expression oscillates. If there is oscillation, the stability detection is determined to be wrong, and a stability detection error message is output. S1350: Robustness detection, performed by the VEM-fe expression function model, includes user input error detection and conflict between two or more input expression vectors. The VEM-fe expression function model outputs a robustness detection error message.
4. The method according to claim 3, characterized in that, The VEM-fe expression function model also includes: S1410: Employs emotion classifications including happiness, sadness, anger, disgust, fear, surprise, and contempt to generate emotion values, which serve as the expression vector input to the VEM-fe expression function model. Specifically, this includes: S1411: Obtain the expression vector. For the vocal, music, recitation, command and dialogue modalities in the emotion classification, use an audio-visual sensor to record audio-visual signals and use the VEM-Token emotion synchronization function to generate one or more data units. The data unit includes emotion name, emotion value and emotion weight, which are used as expression vectors. S1412: Set a unit social task, including a sequence of more than one synthesized emoticons to complete the emoticon output of a social interaction, including the synthesized emoticons of the entire beat of a dialogue and an emoticon exchange; S1413: Set up continuous social tasks, including synthesized expressions for multiple consecutive unit social tasks completed by the robot based on a complete social activity; and / or, S1414: An emotion generator is used to generate emotion values as the expression vectors input to the VEM-fe expression function model. The emotion generator includes an editor for emotion names and emotion values provided by a software interface, and / or a manually operated editor for emotion names and emotion values made of hardware, including a remote control.
5. The method according to claim 4, characterized in that, The VEM-fe expression function model also includes dialogue expressions, specifically: S1510: Uses a microphone to collect speech and employs speech recognition to identify and distinguish more than one social object; S1520: Uses semantic recognition to identify the dialogue content of a single unit social task of social object dialogue, uses VEM-Token emotion synchronization function to identify dialogue emotion, and inputs VEM-fe expression function model to identify the synthetic expression of dialogue. S1530: When the robot engages in one-on-one voice dialogue with a social object, based on the dialogue initiated by the social object, the robot designs dialogue content and establishes a one-to-one pre-dialogue expression spectrum including dialogue content and facial expressions. Using the VEM-fe expression function model, the robot processes the pre-dialogue expression spectrum, including expression synthesis, expression transition, beat capture, expression alignment, and expression conflict resolution, to generate the output dialogue expression spectrum. The robot then outputs voice with expressions and dialogue content through a sound player, completing a single-unit social task; or... S1540: When the robot engages in one-to-many voice dialogues with social objects, it designs dialogue content for each dialogue initiated by the social objects, establishes a one-to-many pre-dialogue expression spectrum including dialogue content and expressions, and uses the VEM-fe expression function model to process the pre-dialogue expression spectrum, including expression synthesis, expression transition, beat capture, expression alignment, and expression conflict resolution, to generate the output dialogue expression spectrum. The robot then outputs voice with expressions and dialogue content through a sound player in multiple unit social tasks; or, S1550: When a robot engages in many-to-many voice dialogue with a social object, multiple robots establish communication, assign a social object to a robot, and complete multiple social tasks by adopting a one-to-many voice dialogue process.
6. The method according to claim 5, characterized in that, The entity-based facial expression generator includes a real-time facial expression mode: S2610: The real-time expression mode includes a controllable bionic muscle layer and an epidermal layer installed on the robot's head. The controllable bionic muscle layer consists of a control mode and movable bionic muscles, so that facial expressions can be generated in a micro-expression simulation mode or a real-time expression mode under the control of the control mode. The epidermal layer is attached to the surface of the controllable bionic muscle layer and is used to imitate the facial organs of real people to produce expressions and facial skin color. It also includes a tongue, teeth, hair, headdress, jewelry and clothing. S2620: The control modality includes electrical signals, force signals, chemical signals and neural signals. It connects to the VEM-fe facial expression function model through the communication interface to obtain signals for synthesizing facial expressions, and drives the controllable bionic muscle layer to display facial expressions through the epidermal layer. and / or, S2630: The real-time expression mode also includes movable bionic eyes with pupils and eyelids, bionic ears, and a bionic mouth to generate visual and auditory input and sound output; S2640: The bionic eyes include dual cameras, dual pupils, and dual eyelids, which can track targets and automatically adjust the video shooting image. The bionic ears include microphones and have sound binaural localization function. The bionic mouth includes a built-in sound player and mouth movement mechanism. Vision, hearing, and sound communicate with and are controlled by the VEM-fe expression function model through a communication interface. The control steps include the motion functions of the expression vectors of the bionic eyes, bionic ears, and bionic mouth in the expression library and the training results. S2650: The VEM-fe expression function model extracts control and communication content from the expression library based on the needs of on-site social interaction, and completes the real-time expression mode operation. S2660: During the dialogue, based on the dialogue content, continuously acquire facial expression vectors to decompose and assemble them into a sequence of facial expression vectors, obtain synthetic facial expressions of the control mode of real-time facial expression mode, continuously drive the controllable bionic muscle layer, and continuously complete facial expressions and dialogue expressions for unit social tasks and continuous social tasks to complete the dialogue process. S2670: Based on user needs, adjust the emotion weight function and VEM-sync emotion vector in the VEM-Token emotion synchronization function to intensify or reduce the emotion in the specified direction, so that the output of the controllable bionic muscle layer is driven to synchronously present the intensification or reduction of the emotion in the specified direction. S2680: Complete user-specified unit social tasks and continuous social tasks.
7. The method according to claim 5, characterized in that, The entity facial expression generator also includes a micro-expression simulation mode: S2700: The local simulation mode includes decomposing rich human facial expressions into more than one micro-expression that can be realized by micro-expression components, in order to mimic real-time facial expression patterns in typical social situations, including the following: S2710: Micro-expressions of the eyes, driving the movement of the eyeballs, eyebrows, and eyelids of both eyes, specifically including: tracking eye contact steps where the pupils track the eyes or specific parts of the face of a social object facing them; gazing at a hand or a hand-held object where the pupils track the hand or a hand-held object facing them; searching for a sound source where the pupils move to find the sound source; awkward and tense steps where the pupils move away from the social object to find another location; inattentive steps where the pupils track a non-social object or a moving object in the environment; blinking steps; furrowed brows indicating contemplation and unhappiness; raised brows on one side indicating suspicion; raised brows on both sides indicating surprise and joy; wide-eyed steps with raised eyelids; drooping eyelids or closed eyes; and user-defined steps where the pupil and eyebrow movement trajectories are set according to the user's needs. S2720: Mouth micro-expressions, driving the movement of the upper and lower lips, tongue, teeth and jaw components, specifically including: natural pursing of the lips, slightly open mouth, open mouth in surprise, blowing, pouting, pursing lips, tightly closed lips, upturned corners of the mouth, downturned corners of the mouth, biting lips, licking lips, and also includes movements learned and trained from real people's mouths when speaking to drive the generation of mouth micro-expressions; S2730: Head micro-expressions, including head shaking, nodding, tilting forward, tilting backward, tilting left, tilting right, turning the head left, turning the head right, and tilting the head to the side. It also includes movements learned and trained from real people's head when they speak to drive the generation of head micro-expressions. S2740: Dialogue micro-expressions. During the dialogue, based on the dialogue content, expression vectors are continuously acquired, decomposed and assembled into expression vector sequences to obtain more than one micro-expression simulation mode. This continuously drives eye micro-expressions, mouth micro-expressions and head micro-expressions to continuously complete the facial expressions and dialogue expressions of unit social tasks and continuous social tasks to complete the dialogue process. S2750: Based on user needs, adjust the emotion weight function and VEM-sync emotion vector in the VEM-Token emotion synchronization function to intensify or reduce the emotion in the specified direction, so that the output of eye micro-expressions, mouth micro-expressions, head micro-expressions and dialogue micro-expressions are simultaneously intensified or reduced. S2760: Complete user-specified unit social tasks and continuous social tasks.
8. The method according to claim 5, characterized in that, The animated facial expression generator includes a micro-expression simulation mode step, specifically including: S2810: Based on the Facial Action Coding System (FACS) and the Facial Action Unit (AU) model, each AU unit is considered as a micro-expression unit. The emotion value of the decomposed expression vector is the driving parameter of the micro-expression unit of more than one AU unit. Based on the VEM-fe expression function model and the micro-expression simulation mode, the driving parameters are adjusted. S2820: The steps to set the driving parameters of the micro-expression unit to drive the micro-expression simulation mode to generate two-dimensional and three-dimensional digital human video animations. The driving parameters include the steps of modifying the modeling of the animation, lighting settings, animation skin layer rendering and playback, including real-time automatic completion and time-sharing manual adjustment. S2830: Drives the micro-expression simulation mode to generate animations of micro-expressions in the unit, and combines the animations to display facial expressions; S2840: Dialogue micro-expressions. During the dialogue process, based on the dialogue expressions in the VEM-fe expression function model and the dialogue content, expression vectors are continuously acquired to decompose and assemble into expression vector sequences to obtain more than one micro-expression simulation mode. The micro-expression simulation mode is continuously driven to generate unit micro-expression animations, and the animations are combined to display facial expressions. The facial expressions and dialogue expressions of unit social tasks and continuous social tasks are continuously completed to complete the dialogue process. S2850: Complete user-specified unit social tasks and continuous social tasks.
9. The method according to claim 5, characterized in that, The animated emoji generator also includes a real-time emoji mode step, specifically including: S2910: For more than one target industry, collect typical facial expression photos or videos, use the VEM-fe facial expression function model, perform the facial expression library collection steps, train and generate an industry facial expression library with the characteristics of the target industry, and incorporate it into the facial expression library; S2920: Decompose the expressions in the industry expression library into combined micro-expressions one by one according to the proximity principle, and establish an approximate correspondence. The approximate correspondence includes using the expression conflict detection step to calculate the sum of FEDC of all expression sequences in the unit social task, and take the expression sequence with the smallest sum to form the closest synthetic expression sequence. S2930: Outputs the sequence of synthesized facial expressions to the real-time facial expression mode, becoming combined micro-expressions; S2940: Dialogue micro-expressions. During the dialogue, based on the dialogue expressions in the VEM-fe expression function model and the dialogue content, expression vectors are continuously acquired to decompose and assemble into expression vector sequences to obtain more than one micro-expression simulation mode. The micro-expression simulation mode is continuously driven to generate unit micro-expression animations, and the animations are combined to display facial expressions. The facial expressions and dialogue expressions of unit social tasks and continuous social tasks are continuously completed to complete the dialogue process. S2950: Facial expressions and dialogue expressions to complete user-specified unit social tasks and continuous social tasks.
10. The method according to any one of claims 6, 7, 8, and 9, characterized in that, include: S3010: Connect to a large model, set target industries, including hotel front desk, restaurant front desk, bank lobby staff, shopping mall sales assistants, and luxury brand store sales assistants, and distill to obtain semantic text of industry dialogues and social dialogues; S3020: Train industry-specific and social emoji packs by inputting industry-specific text for a target industry, and then incorporate the trained packs into the local offline emoji library. S3030: Designed to support end-to-end lightweight computing hardware, installed in space robot systems or digital robot systems, operating in interactive social micro-expression simulation mode or real-time expression mode.
Citation Information
Patent Citations
VEM-Token beat capture and alignment model construction method
CN120748450A
Construction method of vem-token vocal emotion multimodal magic modification model
CN120853611B
VEM-Token emotion synchronization function hierarchical fusion method
CN120913602A