VEM-robot emotional right brain model construction method

By constructing the VEM-robot emotional right brain model, the problems of rigid emotional expression and synchronization in robots were solved, enabling emotional communication and personalized responses between robots and real people, and the VEM-ROS operating system was used for coordination.

CN122165432APending Publication Date: 2026-06-09GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD
Filing Date
2026-05-07
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing robots lack independent emotional state maintenance, their emotional expression is stiff and out of context, they cannot synchronize facial expressions and movements, and they cannot respond to users' emotional expression styles in a personalized way.

Method used

A VEM-robot emotional right brain model was constructed. Through the VEM-Token model and beat capture technology, emotional and logical perceptions were decomposed and synthesized to achieve decoupling and synergy between emotion and logic. The VEM-ROS operating system was used for coordination.

Benefits of technology

It enables emotional communication between robots and real people, simulating natural rhythms to synchronize facial expressions and body movements, and can respond to users' emotional expressions in a personalized way.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122165432A_ABST
    Figure CN122165432A_ABST
Patent Text Reader

Abstract

The VEM-robot emotional right-brain model construction method decomposes the robot's brain into an emotional right brain, an intellectual left brain, and a motor cerebellum, forming the embodied or humanoid robot operating system VEM-ROS. This enables emotional communication between the robot and humans or other robots. Human output is used as the perceptual spectrum, and robot output as the deductive spectrum. Both the perceptual and deductive spectra are segmented, synthesized, and aligned using the emotional rhythm of the multimodal VEM-Token. LLM-Token decomposition of the large language model calculates lexicalized logical perception, and VEM-Token decomposition calculates the micro-expressions of multimodal emotional components. VEM-ROS also includes priority and interruption mechanisms, VEM memory mechanisms, dialogue relationships, craniofacial and limb sensors and actuators and interfaces, robot cloning, and operating system encapsulation. The right-brain model supports emotional micro-expression communication between the robot and humans, similar to the Turing test, and is expected to distinguish whether the emotional test subject is a robot or a human.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, specifically to the subfield of emotional computing and implementation in embody robots and humanoid robots, and also to a brain-based operating system for robots. This invention proposes to decompose the brain operating system into three parts: the sensory left brain, the emotional right brain, and the motor cerebellum. Furthermore, for the emotional right brain, this invention proposes an innovative method for model construction. Background Technology

[0002] The robotics industry is developing rapidly. The hardware design and basic motion control of humanoid and embodied robots have moved from the remote-controlled toy stage to a stage of limited autonomous control. In the remote-controlled toy stage, robot motion control was based on motion hardware such as motors, reducers, and lead screws, along with underlying control programs designed around these components. After adjusting the parameters, semi-automatic or automatic operation of certain movements could be achieved. However, if the motors, reducers, or lead screws were replaced, the corresponding programs and their parameters had to be changed. In the limited autonomous control stage, some manufacturers have designed more universal programs for hardware such as motors, reducers, and lead screws. When replacing these hardware components, only the parameters need to be changed, while the program itself remains unchanged. This provides the benefit of partial generalization in subsequent robot design.

[0003] Looking at the history of the computer industry, these "mini-programs" can be seen as a prototype of a "robot operating system." Some scholars and inventors have referred to these "mini-programs" as the "cerebellum" of a robot. Here, we provide a more systematic classification as follows: 1. Simulation robot operating system: It includes the emotional right brain, the intellectual left brain, and the motor cerebellum, among which: Emotional right brain: Its core functions are multimodal emotion perception, emotion state maintenance, and emotion expression generation. Technical support includes VEM-Token, VEM-sync, VEM-fe, and modified models, among other VEM innovations.

[0004] Intelligent Left Brain: The core functions are logical reasoning, language comprehension, knowledge-based question answering, and task planning. Technical support is provided by existing LLMs (such as DeepSeek and ChatGPT).

[0005] Dynamic cerebellum: Its core functions are balance control, motion coordination, and reflexive movement. The technological support comes from the automated control system for the robot's limb movements.

[0006] 2. Detailed Explanation of the Emotional Right Brain: Core Innovations: For the first time, it is proposed to use VEM-ROS (Vocal-Emotion-Multimodal - Robot Operating System) as an independent "emotional right brain" module of the robot, which works in parallel with the logical left brain and the kinematic cerebellum to achieve decoupling and synergy between emotion and logic.

[0007] The technical problem to be solved: 2.1 Emotion and Logic Overlap: Existing robots treat emotional expression as an "additional function" of the Large Language Model (LLM), resulting in rigid emotional responses that are out of touch with the context.

[0008] 2.2 Lack of independent emotional state maintenance: The robot does not have "emotional memory" and cannot maintain a consistent emotional tendency during long-term interactions.

[0009] 2.3 Emotional expression is out of sync with beat / movement: Unable to accompany speech with natural, beat-synchronized facial expressions and body movements like humans do.

[0010] 2.4 Inability to mimic emotions and personalize: The robot is unable to learn the emotional expression style of a specific user and respond to the user in a personalized way.

[0011] 3. Search Results According to the invention team's search, there are relatively few research results on robot brain models to date. The similar ones found are: The following patents are relevant to the field of robot motion cerebellum: Peking University's "A Cerebellum Training System and Method for Embodied Intelligent Robots Based on Simulation and Reality Fusion (Publication No. CN121122100A); Senard Digital Technology (Jiaxing) Co., Ltd.'s "A Robot Motion Control Method and System Based on Cerebellum Reinforcement Learning (Publication No. CN121756368A); and Suzhou Apache Robotics Technology Co., Ltd.'s "A Real-Time and Deterministic Communication Monitoring System and Method for Robot Cerebellum and Cerebellum (Publication No. CN121864645A)". Among them, "A Real-Time and Deterministic Communication Monitoring System and Method for Robot Cerebellum and Cerebellum (Publication No. CN121864645A)" falls under the field of cerebellum and cerebellum communication monitoring technology.

[0012] 4. VEM-Token Model Technology This invention team has for the first time proposed a set of novel and innovative VEM (Vocal-Emotion-Multimodal) technologies, comprising a patent pool of the following eight invention patents: 4.1 VEM-Token Vocal Emotion Multimodal Model This refers to the model in "VEM-Token Vocal Emotion Multimodal Tokenization Deep Learning Method for Singing and Accompaniment, CN120126506B", which specifically includes the following steps and methods 1-4: (1) Use one or more modalities to record emotions, label vocal emotion multimodalities as VEM, construct VEM classification, VEM coordinate system, VEM function and VEM library. Vocal emotions include one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, affection and enmity. Multimodalities include one or a combination of lyrics, singing, accompaniment, vocal style, music, emotional basis, accompanying instruments, video and image. The VEM coordinate system includes a coordinate axis system established based on independent emotions, opposite emotion pairs and related opposite emotion groups.

[0013] (2) Collect vocal samples according to VEM classification, and have human vocal experts evaluate the vocal samples in terms of emotion in singing and accompaniment. Use supervised learning and deep learning to train the VEM function to obtain VEM parameters and add them to the VEM library.

[0014] (3) The VEM processor is used to calibrate the rhythm of the vocal file and separate the vocal stream and the accompaniment stream. The vocal file is split into VEM-Tokens according to the rhythm, the vocal stream is converted into a VEM-Token1 sequence, the accompaniment stream is converted into a VEM-Token2 sequence, and added to the preprocessing library.

[0015] (4) Using deep learning, the dialogue expression chart, lyrics chart, VEM-Token song chart, VEM-Token accompaniment chart, and VEM-Token music score are generated respectively.

[0016] 4.2 Model for VEM-Token Beat Capture and Beat Alignment This refers to the model in "Method for Constructing VEM-Token Beat Capture and Alignment Model, CN120748450B", which specifically includes the following steps and methods (5-7): (5) For vocal files, based on the VEM-Token vocal emotion multimodal model, a beat model including beat capture and beat alignment is set to capture the beat of the vocal file. Based on the beat, the vocal file is divided into VEM-Token sequences, and the starting point and ending point of the beat in each VEM-Token are marked.

[0017] (6) Set the starting point alignment model, including: The vocal files, including the sample files and the user files generated by the user imitating the sample files, are divided into VEM-Token1 sequences and VEM-Token2 sequences, respectively. Based on the starting point of each VEM-Token1, a starting point fine-tuning step is used to adjust the starting point of the corresponding VEM-Token2 one by one, so that it is aligned with the starting point of the corresponding VEM-Token1.

[0018] For the dialogue segments and rhythms included in the voice dialogue files in social interactions, taking the first segment as a reference, starting from the second segment, a starting point fine-tuning step is adopted to adjust the starting point of each VEM-Token in each segment one by one, so as to align with the starting point of the VEM-Token at the corresponding position in the first segment, until all loop segments are completed.

[0019] (7) Set the endpoint alignment model, including: Based on the endpoint of each VEM-Token1, an endpoint fine-tuning step is used to adjust the endpoint of the corresponding VEM-Token2 one by one, so that it is aligned with the endpoint of the corresponding VEM-Token1.

[0020] For each segment of the dialogue, taking the first segment as a reference, starting from the second segment, a fine-tuning step is used to adjust the endpoint of each VEM-Token in each segment one by one, so that it is aligned with the endpoint of the VEM-Token at the corresponding position in the first segment, until all loop segments are completed.

[0021] 4.3 VEM-Token Vocal Emotion Multimodal Modified Model This refers to the model in "Construction Method of VEM-Token Vocal Emotion Multimodal Modification Model, CN120853611B", which specifically includes the following 8 and 9 steps and methods: (8) Collect sample files and user files. Based on the VEM-Token model, use beat capture and VEM-Token segmentation to obtain the VEM-Token1 sequence of the sample files and the VEM-Token2 sequence of the user files respectively. Based on the VEM-Token1 sequence, use beat alignment for all VEM-Token2 sequences to generate the VEM-Token2 sequence.

[0022] (9) Based on the VEM parameters included in the VEM-Token model, identify the VEM parameters of the VEM-Token1 sequence, determine the modification scheme by the user, process the VEM parameters of the VEM-Token2 sequence, and modify and generate the VEM-Token2 sequence that matches the style of the VEM-Token1 sequence.

[0023] 4.4 A Layered Fusion Method for VEM-Token Sentiment Synchronization Functions This refers to the model in "A Method for Layered Fusion of VEM-Token Sentiment Synchronization Function, CN120913602B", which specifically includes the following steps and methods: 10, 11, and 12. (10) The VEM-Token vocal emotion multimodal model is adopted. The song file is divided into VEM-Token sequences in units of beats. The emotion synchronization function is set as VEM-sync. VEM-sync includes a sequence of VEM-sync vectors that correspond to and are aligned with the VEM-Token sequence. The VEM-sync vector includes synchronization content and a synchronization pointer pointing to the corresponding VEM-Token beat.

[0024] (11) Based on the VEM-Token sequence, calculate and obtain the synchronization content of the VEM-sync vector sequence. The synchronization content includes beat attributes and emotion attributes, where the emotion attributes include more than one emotion name, emotion value, and emotion weight.

[0025] (12) The emotion weight function includes: using hierarchical processing of song files to obtain the corresponding VEM-sync vector sequence, and using one or a combination of multi-layer weighted scanning, recurrent neural network, long short-term memory network, self-attention mechanism and retrieval enhancement generation to generate forward propagation, backward propagation and omnidirectional propagation to obtain the emotion synchronization function with the song file.

[0026] 4.5 VEM-Token World Model Robot Expression Functions This refers to the "Method for Constructing Expression Functions of VEM-Token World Model Robots, CN120951102B", which specifically includes the following steps and methods in 13 and 14: (13) Constructing the VEM-fe expression function model, including: synthesizing one or more expression vectors representing emotional input into the output synthetic expression, aligning the expression vectors in the current beat, eliminating expression conflicts between two or more expression vectors, and transitioning from the synthetic expression of the current beat to the synthetic expression of the next beat, so that the robot can understand emotions and learn expressions.

[0027] (14) The robot uses a physical expression generator or an animated expression generator to form an expression display. The expression driving model is connected by the VEM-fe expression function model to transmit more than one synthetic expression and display facial expressions.

[0028] 4.6 Model Construction Method for VEM-4D Dynamic Skin Structure and Dynamic Expressions This refers to the "VEM-4D Dynamic Skin Structure and Dynamic Expression Model Construction Method, CN120951102B", which specifically includes the following 15 to 18 steps and methods: (15) The VEM-4D dynamic skin includes a 4D dynamic skin and a driving module and a 4D expression model, wherein: (16) The 4D dynamic skin includes two or more electromagnets. The electromagnets include parallel electromagnets composed of soft magnetic materials and excitation coils. The parallel electromagnets are arranged side by side at interval A. The front of the 4D dynamic skin includes a skin that is adhered to the magnetic poles on the front of the parallel electromagnets. The skin is made of an elastic and stretchable membrane.

[0029] (17) The drive module applies an excitation current to the excitation coil, which generates an attractive or repulsive force between the magnetic poles at both ends of the electromagnet, thereby driving the skin to produce expansion and contraction deformation.

[0030] (18) The VEM-4D dynamic skin also includes a 4D expression model, which provides static and dynamic expression instructions to the driving module to drive the skin to stretch and deform.

[0031] 4.7 VEM - Multimodal Emotional Skin System This refers to the "VEM-Multimodal Emotional Skin System, CN121541787B", which specifically includes the following steps and methods (19 to 23): (19) VEM - Multimodal Emotional Skin System, including electromagnets, soft magnetic colloids and multimodal mutual inductance system.

[0032] (20) An electromagnet includes a soft magnetic core and an excitation coil. The excitation current passes through the excitation coil to generate magnetic poles in the electromagnet. Two or more electromagnets are arranged at intervals to form an electromagnet array. The intervals are filled with soft magnetic colloid to form a mutual inductance magnetic circuit, which constitutes the emotional skin.

[0033] (21) Soft magnetic colloid is a mixture of soft magnetic material particles and elastic colloid, and has the resilience to restore its original shape when static.

[0034] (22) The multimodal mutual inductance system receives and executes the expression command, provides excitation current to the electromagnet, and generates magneto-induced peristalsis on the epidermis of the emotional skin to simulate the expression.

[0035] (23) The multimodal mutual inductance system also has a touch sensing function. When an external object applies a touch action on the surface, the touch action is converted and calculated to form a touch signal.

[0036] 4.8 VEM-Robotic Craniofacial End-to-End Emotional Bionic System This refers to the "VEM-Robot Craniofacial End-to-End Emotional Bionic System, CN121706840A", which specifically includes the following 24 to 26 steps and methods.

[0037] (24) Includes a robotic craniofacial structure and an end-to-end emotional bionic system, wherein: (25) The robotic craniofacial structure includes bionic eyeballs, bionic skin, bionic mouth, bionic ears and bionic neck mounted on the craniofacial support.

[0038] (26) The end-to-end emotional bionic system captures visual, auditory, lip tactile and skin tactile input environmental perception information from bionic eyeballs, bionic ears, bionic mouth and bionic skin, generates emotional bionic expression commands, and outputs them from bionic eyeballs, bionic skin, bionic mouth, bionic ears and bionic neck, producing eyeball movement, eyelid movement, facial skin movement, sound output, jaw movement, mouth movement and neck movement, realizing end-to-end bionic expression output from input end to output end.

[0039] Insufficiency of existing technical methods 1. Current technology does not offer a comprehensive solution for robots at the operating system level.

[0040] 2. A holistic model of the robot in terms of emotions has not yet been built.

[0041] 3. There is no mechanism for establishing emotional dialogue and communication between robots or between humans. Summary of the Invention

[0042] In response to the shortcomings of existing technologies, the invention team has proposed a novel “construction method of VEM-robot emotional right brain model” for the first time. This method aims to realize emotional communication between robots and real people, simulating real people. It not only achieves the anthropomorphic realization of voice, tone and facial micro-expressions, but also realizes the micro-expression movements of the robot’s craniofacial and limb movements.

[0043] The purpose and intent of this invention are achieved by the following structure and method: 1. Construction Method of VEM-Robot Emotional Right Brain Model This invention, as a method for constructing the VEM-robot emotional right brain model, includes, but is not limited to, the following methods: This robot, designed to mimic a human, engages in emotional communication with one or more interlocutors, including but not limited to real people or robots. An emotional feature database records and distinguishes the emotional characteristics of the mimic and the interlocutor. The robot includes, but is not limited to, the emotional right brain. The Emotional Right Brain Model uses a beat segmentation and alignment method to decompose the perceptual spectrum emitted by the communicator and perceived by the robot through sensors into emotional perception and logical perception, including logical expression, according to the emotional feature method.

[0044] The right-brain model of emotion, based on emotion perception and an emotion feature database, uses emotion computation to generate emotion expression, which is then synchronized with logical expression through rhythm synthesis to produce a deductive spectrum.

[0045] Preferably, the deductive spectrum is output to the robot's cranial actuators to form emotional communication between facial expressions and vocal expressions.

[0046] Furthermore, the actuator output of the deductive spectrum is used to communicate emotions through the robot's limb movements.

[0047] 2. The right-brain model of emotion also includes: Based on the aforementioned basic solution, this invention also includes, but is not limited to, the following methods in terms of the emotional right-brain model: The VEM-Token vocal emotion multimodal method is used to construct an emotion feature library and VEM-Token lexical units. The emotion feature library includes, but is not limited to, one or a combination of emotion, music, rhythm, imagination, creativity and art.

[0048] We employ a method for constructing a VEM-Token beat capture and alignment model to complete the calibration and alignment of VEM-Token lexical beats, and use a layered fusion method of VEM-Token sentiment synchronization function to obtain the sentiment function.

[0049] The VEM coordinate system is used to manage multimodal sentiment components, where: Emotional components include, but are not limited to, one or a combination of the following multimodalities: joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, affection, and enmity.

[0050] The emotional components also include, but are not limited to, the weighting functions of each component, which are determined by setting or learning based on the emotional characteristics of the simulated human.

[0051] The sentiment function includes, but is not limited to, sentiment components and weighting functions. The sentiment function is a dependent variable with sentiment components and weighting functions as independent variables, and sentiment components and sentiment functions are incorporated into the sentiment feature library.

[0052] The generation of the sentiment feature database also includes, but is not limited to, methods such as initial definition, supervised learning, learning of the emotional communication process, and one or a combination of long-term attention mechanisms.

[0053] 3. Includes beat model: Based on the aforementioned solution, the present invention also includes, but is not limited to, a beat model: The beat model is divided using a VEM-Token beat capture and alignment model construction method, including but not limited to: The music beat model, based on the VEM-Token model and beat theory in music theory, marks the start and end points of the music beat in the perceptual spectrum and the deductive spectrum according to the start and end points of the music beat.

[0054] The punctuation rhythm model, based on the LLM-Token model and the grammatical conventions of text sentences, marks the start and end points of the punctuation rhythm according to the start and end points of the punctuation.

[0055] The paragraph beat model, based on a complete dialogue logic paragraph and dialogue habits, marks the start and end points of the paragraph beat according to the start and end points of the paragraph's perceptual spectrum and deductive spectrum.

[0056] Beat conflict fine-tuning involves fine-tuning the start and end points of the perceptual spectrum when using beat segmentation alignment for emotional perception and logical perception, and when using beat synthesis alignment for emotional expression and logical expression for deductive spectrum, to achieve beat alignment.

[0057] Beat capture methods also include, but are not limited to, converting the perceptual spectrum and deductive spectrum into a spectrum format file, then decomposing it into a harmonic component spectrum and an impulse component spectrum, detecting abrupt changes in spectral energy, and finally determining the start point of the beat.

[0058] 4. Deductive spectrum and perceptual spectrum include: Based on the aforementioned scheme, the present invention further includes, in addition to the deductive spectrum and the perceptual spectrum: S41: For a given punctuation beat, the text included in its logical expression and the emotional components and weighting functions included in the emotional expression of the corresponding musical beat are aligned through beat synthesis to generate a punctuation beat's interpretation spectrum.

[0059] S42: For a dialogue segment, repeat S41 to complete the deductive spectrum of the dialogue segment.

[0060] S43: Output the deduction spectrum of a dialogue segment and perform sentiment function smoothing verification. Specifically, within a dialogue segment, compare the fluctuation of the sentiment function in the consecutive punctuation beats before and after. If the fluctuation residual is less than the preset value, the verification is deemed qualified, and the deduction spectrum is output. If the fluctuation residual is greater than the preset value, the verification is deemed unqualified, and the sentiment function is smoothed. After smoothing, the deduction spectrum is output.

[0061] Preferably, or further, S44: For a dialogue segment, the logical expression remains unchanged, but the emotional components and weighting functions in the emotional expression are unconditionally changed to generate a modified deductive spectrum.

[0062] S45: The deductive spectrum is output to the robot's craniofacial region or the robot's craniofacial region and limbs in units of a dialogue segment to complete the emotional communication of a dialogue segment.

[0063] S46: In the emotional perception and logical perception decomposed by the perceptual spectrum, the beat capture method is used to determine the start and end points of the musical beat and the punctuation beat, and the beat conflict fine-tuning method is used to align the start and end points of the musical beat and the punctuation beat.

[0064] S47: Deductive spectrum and perceptual spectrum also include, but are not limited to, question-and-answer mode. When the communicator sends out a question dialogue segment, the robot's answer dialogue segment is given the highest priority. When the robot sends out a question dialogue segment, if the communicator does not answer within a set time, the question is suspended and listed as a pending issue.

[0065] S48: The perception spectrum is generated by sensors of the robot, including but not limited to eyes, ears, skin and spatial sensing. The deductive spectrum is output by actuators of the robot, including but not limited to speaker sound output, craniofacial skin expression dynamic output, eye movement, eyelid movement and mouth movement.

[0066] 5. The robot includes an operating system: Based on the aforementioned solution, the present invention also includes an operating system on the robot: The robot also includes, but is not limited to, an intelligent left brain and a kinematic cerebellum. The intelligent left brain is used for logical reasoning and analysis, language and text comprehension, knowledge question answering, task planning and training memory. The intelligent left brain also includes, but is not limited to, working with a large language model using LLM-Token. The kinematic cerebellum is used for limb balance control, movement coordination, reflexive movement and training memory.

[0067] The robot's brain consists of an emotional right brain, an intelligent left brain, and a motor cerebellum. These three components are integrated to form the robot's brain operating system. The emotional right brain model communicates bidirectionally with the intelligent left brain through an emotion-logic interface, and the emotional right brain model communicates bidirectionally with the motor cerebellum through an emotion-action interface. The emotional right brain model, along with the intelligent left brain and the motor cerebellum, works together to achieve end-to-end emotional simulation and emotional communication with the communicator.

[0068] For the robot's emotional right brain, intellectual left brain, and motor cerebellum, there are options including but not limited to using their own independent processors and operating system software, as well as using a unified processor and unified operating system software for the emotional right brain, intellectual left brain, and motor cerebellum.

[0069] 6. Includes priority and interruption: Based on the aforementioned solution, this invention includes priority and interruption: The brain's working priority includes at least the following: the cerebellum as the first priority, the right brain as the emotional brain, and the left brain as the intellectual brain as the second priority, with the first priority being higher than the second priority.

[0070] Brain-based work priorities include, but are not limited to: high response speed priority, short response time priority, high importance priority, and custom priority.

[0071] Interruption priorities include, but are not limited to: when the robot is handling concurrent conflicting tasks, according to the priority order, allowing high-priority tasks to interrupt low-priority tasks, and allowing custom tasks to interrupt other tasks based on their priority.

[0072] Interrupt task recovery includes, but is not limited to: before the original task is interrupted by the high-level interrupt task, the processor saves the interrupt context parameters of the original task, enters the execution of the high-level interrupt task, and after the high-level interrupt task is completed, restores the interrupt context and resumes the original interrupt task.

[0073] 7. Includes VEM memory mechanism: Based on the aforementioned solution, this invention also includes a VEM memory mechanism: VEM memory mechanisms include, but are not limited to, one or a combination of person recognition mechanisms, event recording mechanisms, time memory mechanisms, multi-headed hotspot attention mechanisms, excitation-inhibition mechanisms, dormancy-awakening mechanisms, selective memory-forgetting mechanisms, and are stored in an emotional feature database.

[0074] The recognition mechanism includes, but is not limited to, a one-to-many relationship memory when a simulated human engages in emotional communication with multiple communicators. The memory content includes, but is not limited to, the craniofacial features, voice features, relationship features between the communicator and the simulated human, and communicator ID.

[0075] The recording mechanism includes, but is not limited to, extracting the content of the emotional exchange between the simulated human and the communicator, setting hot topics and hot function values, and recording the timestamp of the exchange, provided that the communicator has been identified.

[0076] The time memory mechanism includes, but is not limited to, setting a decay function based on the length of time since the communication timestamp, calculating the decay residual for hot content and hot function values, and marking the memory as valid when the residual is greater than the set value; otherwise, marking the memory as forgotten.

[0077] The multi-head hotspot attention mechanism includes, but is not limited to, setting two or more hotspot contents and hotspot function values. When two or more hotspot contents appear consecutively in a single search, the attenuation function weight is modified, and the attenuation residual is calculated. When the residual is greater than the set value, the memory is marked as valid; otherwise, the memory is marked as forgotten.

[0078] The excitation-inhibition mechanism includes, but is not limited to, setting excitation and inhibition values ​​for event time, event frequency, and event type, calculating the excitation-inhibition function, calculating the function residual, and marking the memory as valid when the residual is greater than the set value, otherwise marking the memory as forgotten.

[0079] The hibernation-wake mechanism includes, but is not limited to, setting hibernation and wake-up values ​​based on event time and event frequency, and calculating the hibernation-wake function.

[0080] Selective memory-forgetting mechanisms include, but are not limited to, setting memory events and forgetting events based on the personality traits of simulated humans, retrieving control events, and labeling memory and forgetting values.

[0081] 8. Including the dialogue relationship between the robot and the communicator: Based on the aforementioned solution, the present invention further includes the following regarding the dialogue relationship between the robot and the communicator: The robot establishes a one-to-one relationship with the communicator, using a human recognition mechanism to identify a communicator. It establishes a logical relationship and spatial visual positioning between the robot and the communicator based on the perceptual spectrum, deductive spectrum, and questions and answers. Based on the questions in the perceptual spectrum, it calculates the answers in the deductive spectrum and outputs them to the craniofacial and limb areas. Alternatively, based on the questions in the deductive spectrum and the answers in the perceptual spectrum, it determines the logic and completeness of the dialogue and the spatial visual positioning of the robot when it looks at the communicator during the dialogue.

[0082] The robot and the communicator form a one-to-many relationship. Based on the logical relationship between questions and answers in the perceptual spectrum and deductive spectrum, the robot records and buffers the questions or answers of the communicator, and outputs the robot's perceptual spectrum or deductive spectrum one by one.

[0083] The relationship between robots and communicators is a many-to-one relationship. Based on the dialogue relationship between multiple robots and one communicator, dialogue agreements are established for multiple robots, including question-based dialogues, agreed-upon robot dialogues, and custom robot dialogues.

[0084] The relationship between the robot and the communicator is many-to-many. The dialogue relationship between the robot and the communicator is agreed upon in advance, and is decomposed, determined and executed as a one-to-one or one-to-many dialogue relationship.

[0085] The switching of the relationship between the robot and the communicator is determined by setting or negotiation between the robot and the communicator, and the relationship switching of the next dialogue segment is determined on a segment-by-segment basis, including switching the simulator and switching the communicator.

[0086] 9. Includes craniofacial region and limbs: Based on the aforementioned solutions, the present invention further includes, on the craniofacial region and limbs: Craniofacial systems include, but are not limited to, the VEM-robot craniofacial end-to-end emotional bionic system; The craniofacial structure also includes, but is not limited to, actuators mounted on a craniofacial support that conform to the shape and movement habits of a simulated human, such as simulated eyeballs, simulated skin, simulated mouth, simulated ears, and simulated neck, used to mimic the output process of dynamic facial expressions and vocal output of a simulated human. It also includes, but is not limited to, the spatial visual positioning of a simulated human visual communicator, and is used to capture video input signals within the field of view and record audio input signals within the auditory range.

[0087] The craniofacial region also includes, but is not limited to, spatial coordinate sensors, motion sensors, navigation sensors, time sensors, and communication interfaces. Through the communication interface, craniofacial signals are transmitted as a perceptual spectrum to the emotional right brain model.

[0088] Limbs include, but are not limited to, actuators and limb interfaces of movable parts such as the robot's limbs, arms, legs, neck, and waist. The limb interface is connected to the emotional output, performs the communication of the emotional output, and generates body expressions.

[0089] The deductive spectrum connects the actuators of the craniofacial region and the limbs through a communication interface to realize the output of emotional communication of a simulated human. It also captures the emotional communication of the communicator through craniofacial sensors and converts the input into a perceptual spectrum.

[0090] 10. Includes clone mode: Based on the aforementioned scheme, the present invention includes a cloning mode: The hardware and operation of the emotional right brain model are located locally on the robot.

[0091] The emotional feature database is stored in the robot's local hardware.

[0092] In the intelligent left-brain model, the large language model is either placed locally or connected to the cloud center via the external network.

[0093] Cloning mode includes, but is not limited to, replicating one humanoid into two or more in real time, allowing two or more humanoids to engage in emotional communication with the communicator simultaneously.

[0094] In the monitoring mode, which is in clone mode, a simulated human is ensured to receive only the perceptual spectrum and not generate the deductive spectrum.

[0095] 11. The emotional right-brain model is encapsulated using an operating system. Based on the aforementioned solution, this invention includes an operating system encapsulation in the emotional right-brain model: Based on the EOS operating system for emotional robots, the emotional right-brain model is rewritten and encapsulated into the emotional multimodal VEM-ROS and merged into the EOS operating system, forming a part of EOS, specifically including but not limited to: Rewrite and encapsulate the physical drivers of sensors and actuators into interrupt or process modes.

[0096] Based on network communication protocols, rewrite and encapsulate emotion expressions to conform to network communication protocols, including but not limited to ISO-OSI protocol, AI agent protocol, model and tool protocol, agent-to-agent collaboration protocol, agent-to-user interface protocol.

[0097] Based on priority and interruption, rewrite and encapsulate the definition of emotional right brain functions.

[0098] 12. Purpose and Intent of the Invention The purpose and intent of the method for constructing the VEM-robot emotional right brain model of this invention are as follows: (1) It solves the shortcomings of existing technologies and proposes the core innovation for the first time - the construction method of VEM-robot emotional right brain model.

[0099] (2) VEM-ROS is used as an independent “emotional right brain” module of the robot, working in parallel with the logical left brain and the dynamic cerebellum to achieve decoupling and synergy between emotion and logic.

[0100] (3) Solve the following technical problems: 3.1 Conflict between emotion and logic: Existing robots treat emotional expression as an "additional function" of LLM, resulting in rigid emotional responses that are out of touch with the context.

[0101] 3.2 Lack of independent emotional state maintenance: The robot does not have "emotional memory" and cannot maintain a consistent emotional tendency during long-term interactions.

[0102] 3.3 Emotional expression is out of sync with beat / movement: Unable to accompany speech with natural, beat-synchronized facial expressions and body movements like humans do.

[0103] 3.4 Inability to mimic emotions and personalize: The robot is unable to learn the emotional expression style of a specific user and respond to the user in a personalized way.

[0104] 13. Beneficial effects of the invention (1) Construct a model of the brain operating system of the simulation robot and the embodied robot.

[0105] (2) To enable robots to simulate the emotions of real people and to achieve emotional communication with real people.

[0106] (3) Realize the generalized encapsulation of the VEM-robot emotional right brain model for operating system. Attached Figure Description

[0107] List of attached images: Figure 1 VEM Robot Operating System Structure and Flowchart Figure 2 Schematic diagram of perceptual and deductive spectra based on beat Figure 3 : Diagram of Emotional Communication Figure 4 Music demonstration illustration Detailed description of the attached diagram: For detailed descriptions of each accompanying drawing, please refer to the description of each drawing corresponding to the embodiments. It should be emphasized that these drawings are merely one illustration of the innovative concept of this invention and are simply schematic diagrams. The system and structural parts in the drawings are not specifically labeled; users of this invention can provide detailed design annotations based on specific applications. Therefore, the accompanying drawings are not intended to limit the innovative concept of this invention. Users of this invention can certainly draw other types of drawings and their resulting understanding and explanation based on industry-standard knowledge. Detailed Implementation

[0108] The objectives and intentions of this invention are achieved through the following specific embodiments. It should be noted that each specific embodiment has its own specific use and industrial applicability. Therefore, the following embodiments do not encompass all system features and implementation methods of this invention, and the accompanying drawings are merely one implementation method and do not constitute a limitation of this invention. The description in the claims of this invention is the core summary of this invention.

[0109] Specific embodiments of the present invention are illustrated below: VEM - A Method for Constructing a Robotic Emotional Right Brain Model - Design Examples Diagram Explanation This embodiment mainly includes, but is not limited to, the following schematic diagrams, which are described in detail in the accompanying drawings. Figures 1 to 4 .

[0110] System Implementation Instructions This embodiment mainly includes method implementations 1 to 11. A portion of these is the basic system of this application, while the other portion consists of optional combinations based on actual applications. This is explained below: In this specification, a logical line refers to a paragraph. The word "furthermore" in a logical line indicates that the logical line is a required option for the next step of the previous logical line, while "preferably" indicates that the logical line is not a required option but an optional option for optimization in some applications.

[0111] Unless otherwise specified, these systems are not entirely necessary.

[0112] Unless otherwise specified, their order is not required.

[0113] Unless otherwise specified, the selection of materials and design parameters is not mandatory.

[0114] Based on the specific task requirements, the patent implementer makes the preferred and further selections.

[0115] 1. Construction Method of VEM-Robot Emotional Right Brain Model This invention, as a method for constructing the VEM-robot emotional right brain model, includes, but is not limited to, the following methods: This robot, designed to mimic a human, engages in emotional communication with one or more interlocutors, including but not limited to real people or robots. An emotional feature database records and distinguishes the emotional characteristics of the mimic and the interlocutor. The robot includes, but is not limited to, the emotional right brain. The Emotional Right Brain Model uses a beat segmentation and alignment method to decompose the perceptual spectrum emitted by the communicator and perceived by the robot through sensors into emotional perception and logical perception, including logical expression, according to the emotional feature method.

[0116] The right-brain model of emotion, based on emotion perception and an emotion feature database, uses emotion computation to generate emotion expression, which is then synchronized with logical expression through rhythm synthesis to produce a deductive spectrum.

[0117] Preferably, the deductive spectrum is output to the robot's cranial actuators to form emotional communication between facial expressions and vocal expressions.

[0118] Furthermore, the deductive spectrum is output to the actuators of the robot's limbs to form emotional communication of actions.

[0119] Here, the robot is used to mimic a human, engaging in emotional communication with a human, who can be a real person or another robot. This emotional communication differs from current mechanical human-computer interaction, or the monotonous imitation of human voices. The reason is that current human-computer interaction is almost entirely based on large language models using LLM-Tokens. In these models, both input and output are textual units, unable to understand or process emotional variables. Even if some voice output resembles a certain actor's voice, it is merely mechanical imitation, not genuine emotional dialogue.

[0120] This invention is the first to propose the VEM-Robot Emotional Right Brain Model, a novel and original innovation. First, it establishes the brain concept of the robot operating system—VEM-ROS (Vocal-Emotion-Multimodal Robot Operating System)—which handles the basic management of robot hardware in a standardized, unified, and operating system-like manner. Second, it decomposes VEM-ROS into three parts: the emotional right brain, the intellectual left brain, and the motor cerebellum, enabling coordinated operation. Third, it deepens the research and design of the emotional right brain model, completing beat segmentation alignment and beat synthesis alignment, and achieving emotional communication across the perceptual and deductive spectra.

[0121] Figure 1This is a schematic diagram of the VEM robot operating system structure of the present invention. The pink functional blocks represent the content of the emotional right brain, including beat synthesis and alignment of the deductive spectrum, emotional expression, VEM-Token, emotional computation, emotional feature library, and emotional perception. The light blue functional blocks represent the content of the intellectual left brain, including beat segmentation and alignment of the perceptual spectrum, logical perception, logical computation, LLM-Token, and logical expression. The light green functional blocks represent the content of the kinematic cerebellum, including balance control and motion decomposition. Figure 1 The right side is a lifelike human figure.

[0122] Figure 1 The workflow is as follows: the simulated human engages in emotional exchange with the communicator. Assume the information sequence and flow of this emotional exchange are as follows: S1. The communicator first initiates an emotional exchange question.

[0123] S2. The simulated human obtains the perception spectrum from the sensor.

[0124] S3. The perceptual spectrum is divided and aligned by the beat, decomposing into emotional perception and logical perception, which are sent to the emotional right brain and the intellectual left brain respectively. The logical perception may also need to be sent to the motor cerebellum.

[0125] S4.1 The emotion perception is sent to VEM-Token to parse out the lexical units of the vocal emotion multimodal, obtain the emotion function and emotion components, and further decompose them by emotion operation and store them in the emotion feature library.

[0126] S4.2 Preferably or further, the emotional function and emotional component are sent to the kinematic cerebellum and transmitted to the robot's limbs through motion decomposition and balance control to form emotional communication of movements.

[0127] S4.3. Logical perception is delivered to the LLM-Token to parse out the lexical units of the large language model, obtain the text in the logical perception, and, with the support of logical operations, parse out the logical meaning of the communicator in this question.

[0128] S4.4 Preferably or further, the corresponding kinematic cerebellum components, such as nodding and shaking, are calculated by logical operations and transmitted to the robot's limb actuators through motion decomposition and balance control to form emotional communication of actions.

[0129] S5, the intelligent left brain generates a logical expression of response based on logical perception, which is sent to the beat synthesis alignment and emotion calculation respectively. In the emotion calculation, an emotional expression is generated based on VEM-Token and the emotion feature library. In the beat synthesis alignment, the emotional expression and the logical expression are synthesized together to form a deductive spectrum.

[0130] S6. Preferably or further, the deductive spectrum is output to the robot's craniofacial actuator to form emotional communication of facial expressions and vocal expressions. For some needs, the deductive spectrum is also output to the robot's limbs to form emotional communication of movements.

[0131] Regarding humanoid robots, these are robots that imitate humans. Humanoid robots include those whose facial features are defined by their skull shape; for example, if the skull shape imitates the appearance of Zhang San, then the robot is a humanoid imitating Zhang San. Interactors, on the other hand, are people who engage in emotional communication with the robot. These include real people and other robots, and can include more than one. It's important to note that when there are more than one interactor, it's necessary to distinguish between different robot dialogues. For example, if the interactors are Zhang San, Li Si, and Wang Wu, the emotional dialogue output pattern must be strictly tailored to the interactor group.

[0132] 2. The right-brain model of emotion also includes: Based on the aforementioned basic solution, this invention also includes, but is not limited to, the following methods in terms of the emotional right-brain model: The VEM-Token vocal emotion multimodal method is used to construct an emotion feature library and VEM-Token lexical units. The emotion feature library includes, but is not limited to, one or a combination of emotion, music, rhythm, imagination, creativity and art.

[0133] We employ a method for constructing a VEM-Token beat capture and alignment model to complete the calibration and alignment of VEM-Token lexical beats, and use a layered fusion method of VEM-Token sentiment synchronization function to obtain the sentiment function.

[0134] The VEM coordinate system is used to manage multimodal sentiment components, where: Preferably, the emotional components include, but are not limited to, one or a combination of multiple modalities such as joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, affection, and enmity.

[0135] Preferably, the emotional components also include, but are not limited to, weighting functions for each component, which are determined by setting or learning based on the emotional characteristics of the simulated human.

[0136] Preferably, the sentiment function includes, but is not limited to, sentiment components and weighting functions. The sentiment function is a dependent variable with sentiment components and weighting functions as independent variables, and the sentiment components and sentiment functions are incorporated into the sentiment feature library.

[0137] Preferably, the generation of the emotional feature library also includes, but is not limited to, methods such as initial definition, supervised learning, learning of emotional communication processes, and long-term attention mechanisms or combinations thereof.

[0138] Among them, VEM-Token includes methods supported by the basic principles of vocal emotion multimodality, including at least the patented methods of VEM-Token vocal emotion multimodality method, VEM-Token beat capture and alignment model construction method, and VEM-Token emotion synchronization function hierarchical fusion method, as well as the obtained emotion function and emotion component.

[0139] Regarding the emotional feature database, the emotional components and functions are specifically configured and categorized based on the application of the right brain in each specific situation. For example, in music applications, configuration and categorization are needed based on music theory and music genres; in emotional support applications, configuration and categorization are needed based on psychology and the cultural background of the person being cared for. In particular, this configuration and categorization also includes supervised learning and deep learning for discovery and creation. Based on these creations, configurations, and categorizations, the data is recorded and stored in the situational feature database.

[0140] 3. Includes a section model: Based on the aforementioned solution, the present invention also includes, but is not limited to, a beat model: The beat model is divided using a VEM-Token beat capture and alignment model construction method, including but not limited to: The music beat model, based on the VEM-Token model and beat theory in music theory, marks the start and end points of the music beat in the perceptual spectrum and the deductive spectrum according to the start and end points of the music beat.

[0141] The punctuation rhythm model, based on the LLM-Token model and the grammatical conventions of text sentences, marks the start and end points of the punctuation rhythm according to the start and end points of the punctuation.

[0142] The paragraph beat model, based on a complete dialogue logic paragraph and dialogue habits, marks the start and end points of the paragraph beat according to the start and end points of the paragraph's perceptual spectrum and deductive spectrum.

[0143] Beat conflict fine-tuning involves fine-tuning the start and end points of the perceptual spectrum when using beat segmentation alignment for emotional perception and logical perception, and when using beat synthesis alignment for emotional expression and logical expression for deductive spectrum, to achieve beat alignment.

[0144] Beat capture methods also include, but are not limited to, converting the perceptual spectrum and deductive spectrum into a spectrum format file, then decomposing it into a harmonic component spectrum and an impulse component spectrum, detecting abrupt changes in spectral energy, and finally determining the start point of the beat.

[0145] It's important to note that the VEM-Token model and the LLM-Token model are completely different in principle and mechanism. The VEM-Token model uses a beat capture and alignment method, setting a start-point alignment model and an end-point alignment model. It converts the perceptual spectrum into a spectral format file, and then, based on the music theory definition of strong and weak beats, completes beat capture and alignment. It also includes emotional expressions, such as emotional components, emotional functions, and emotional weights, such as scalars and vectors for joy, anger, sorrow, and happiness, thus generating a music beat model. The LLM-Token model, on the other hand, is based on a large language model, a model based on word units, such as Chinese characters and phrases, or English letters and words. Therefore, the length of a VEM-Token is much greater than that of an LLM-Token, and the content of a VEM-Token is far richer than that of an LLM-Token.

[0146] 4. Deductive spectrum and perceptual spectrum also include: Based on the aforementioned scheme, the present invention further includes, in addition to the deductive spectrum and the perceptual spectrum: S41: For a given punctuation beat, the text included in its logical expression and the emotional components and weighting functions included in the emotional expression of the corresponding musical beat are aligned through beat synthesis to generate a punctuation beat's interpretation spectrum.

[0147] S42: For a dialogue segment, repeat S41 to complete the deductive spectrum of the dialogue segment.

[0148] S43: Output the deduction spectrum of a dialogue segment and perform sentiment function smoothing verification. Specifically, within a dialogue segment, compare the fluctuation of the sentiment function in the consecutive punctuation beats before and after. If the fluctuation residual is less than the preset value, the verification is deemed qualified, and the deduction spectrum is output. If the fluctuation residual is greater than the preset value, the verification is deemed unqualified, and the sentiment function is smoothed. After smoothing, the deduction spectrum is output.

[0149] Preferably, or further, S44: For a dialogue segment, the logical expression remains unchanged, but the emotional components and weighting functions in the emotional expression are unconditionally changed to generate a modified deductive spectrum.

[0150] S45: The deductive spectrum is output to the robot's craniofacial region or the robot's craniofacial region and limbs in units of a dialogue segment to complete the emotional communication of a dialogue segment.

[0151] S46: In the emotional perception and logical perception decomposed by the perceptual spectrum, the beat capture method is used to determine the start and end points of the musical beat and the punctuation beat, and the beat conflict fine-tuning method is used to align the start and end points of the musical beat and the punctuation beat.

[0152] S47: Deductive spectrum and perceptual spectrum also include, but are not limited to, question-and-answer mode. When the communicator sends out a question dialogue segment, the robot's answer dialogue segment is given the highest priority. When the robot sends out a question dialogue segment, if the communicator does not answer within a set time, the question is suspended and listed as a pending issue.

[0153] S48: The perception spectrum is generated by sensors of the robot, including but not limited to eyes, ears, skin and spatial sensing. The deductive spectrum is output by actuators of the robot, including but not limited to speaker sound output, craniofacial skin expression dynamic output, eye movement, eyelid movement and mouth movement.

[0154] The deductive spectrum and perceptual spectrum record information using two or more dimensions. In two-dimensional recording, the horizontal axis records time, including beats. The vertical axis records components, including the textual components of LLM-Tokens and multiple emotional components of VEM-Tokens. On the horizontal axis, the components for each beat are aligned vertically. In the beat synthesis and alignment module, emotional and logical expressions are synthesized into a deductive spectrum. This deductive spectrum is sent to the robot's craniofacial and limb systems, generating emotional communication. When the communicator emits emotional communication, the robot collects the perceptual spectrum through sensors and decomposes it into emotional perception and logical perception in the beat segmentation and alignment module. Specifically, the VEM-Token model is used to decompose beats and emotional components, while the LLM-Token model is used to decompose textual units, punctuation, and dialogue segments. Based on the horizontal axis, beats are aligned on both the emotional perception and logical perception axes.

[0155] Figure 2 This is a schematic diagram of the perceptual and deductive spectra based on rhythm. From left to right, the color blocks in the table represent the rhythm content (white), the deductive spectrum (pink), and the perceptual spectrum (blue). Specifically: Within the beat content, there are LLM-Token rows representing the large language model and VEM-Token rows representing the sentiment model and its sentiment components. The LLM-Token rows include word units from the deductive and perceptual spectra, such as the robot saying, "Hello sir!" or the guest saying, "Hello, do you have a double room?". The VEM-Token sentiment component rows also include sentiment components such as cue words, tone, and speech rate, corresponding to the descriptions of sentiment components in the deductive and perceptual spectra, respectively.

[0156] In the deductive spectrum, the LLM-Token and VEM-Token are aligned according to the beat, and then the emotional components are obtained based on the emotion calculation. According to the parameters given in the emotional components, the corresponding emotional components are adjusted, and the components are aligned by beat synthesis and loaded onto the LLM-Token to form an emotional communication with facial expressions, body expressions and voice expressions. The output is sent to the robot's actuator to simulate and perform the emotional communication.

[0157] In the perception spectrum, this is the guest's response to the robot. The robot's sensors collect the guest's voice, video, and body movements, which are segmented into the LLM-Token "Hello, do you have a king-size bed?" and the VEM-Token emotional component, as shown in the table.

[0158] Thus, it can be seen that this invention, through the right-brain emotion model, completes the processing and acquisition of the perceptual spectrum and the processing and generation of the deductive spectrum, thereby forming emotional communication between the robot and the real person.

[0159] Figure 3 This is a schematic diagram of a complete business conversation involving emotional exchange. Here, you can see that the dialogue between the robot and the customer is arranged in a sentence-by-sentence rhythm. Below this, emotional components are arranged according to the aligned rhythm, such as tone of voice, micro-expressions of eyebrows, eyes, and mouth. Thus, the dialogue output not only includes vocal expressions but also facial expressions, forming completely human-like micro-expressions that perfectly simulate real human expressions.

[0160] Figure 4 This is a diagram illustrating the emotional expression of a piece of music. It shows the sheet music and interpretation of a song, from bottom to top: lyrics, simplified notation, standard notation, rhythm, and the intensity of subtle facial expressions such as chin, eyes, and eyebrows. In this example, emotional components are expressed using analog quantities. For instance, the eyelid opening gradually increases and then decreases from left to right, representing the singer's use of eyelid opening to convey the emotion of their performance.

[0161] 5. The robot includes an operating system: Based on the aforementioned solution, the present invention also includes an operating system on the robot: The robot also includes, but is not limited to, an intelligent left brain and a kinematic cerebellum. The intelligent left brain is used for logical reasoning and analysis, language and text comprehension, knowledge question answering, task planning and training memory. The intelligent left brain also includes, but is not limited to, working with a large language model using LLM-Token. The kinematic cerebellum is used for limb balance control, movement coordination, reflexive movement and training memory.

[0162] The robot's brain consists of an emotional right brain, an intelligent left brain, and a motor cerebellum. These three components are integrated to form the robot's brain operating system. The emotional right brain model communicates bidirectionally with the intelligent left brain through an emotion-logic interface, and the emotional right brain model communicates bidirectionally with the motor cerebellum through an emotion-action interface. The emotional right brain model, along with the intelligent left brain and the motor cerebellum, works together to achieve end-to-end emotional simulation and emotional communication with the communicator.

[0163] Furthermore, for the robot's emotional right brain, intellectual left brain, and motor cerebellum, there are options including but not limited to using their own independent processors and operating system software, as well as using a unified processor and unified operating system software for the emotional right brain, intellectual left brain, and motor cerebellum.

[0164] It's important to note here that the robot's brain specifically refers to the operating system VEM-ROS (Vocal-Emotion-Multimodal - Robot Operating System). This is a general, hierarchical definition, and within the operating system, the divisions include: The hardware abstraction layer encapsulates the hardware drivers for bionic eyes, skin, ears, neck, etc. The emotional core layer includes emotional state vector maintenance, beat synchronization, and emotional scheduling. The logical lexical layer decomposes emotional dialogue into textual lexical units and synthesizes them into speech; Balanced motion layer, motion decomposition, and balance control; The system service layer includes emotion perception service, emotion expression service, emotion memory service, and emotion learning service. The API interface layer provides standardized interfaces for calling emotional functions. The application framework layer supports the development of third-party emotion applications.

[0165] 6. Includes priority and interruption: Based on the aforementioned solution, this invention includes priority and interruption: The brain's working priority includes at least the following: the cerebellum as the first priority, the right brain as the emotional brain, and the left brain as the intellectual brain as the second priority, with the first priority being higher than the second priority.

[0166] Brain-based work priorities include, but are not limited to: high response speed priority, short response time priority, high importance priority, and custom priority.

[0167] Interruption priorities include, but are not limited to: when the robot is handling concurrent conflicting tasks, according to the priority order, allowing high-priority tasks to interrupt low-priority tasks, and allowing custom tasks to interrupt other tasks based on their priority.

[0168] Interrupt task recovery includes, but is not limited to: before the original task is interrupted by the high-level interrupt task, the processor saves the interrupt context parameters of the original task, enters the execution of the high-level interrupt task, and after the high-level interrupt task is completed, restores the interrupt context and resumes the original interrupt task.

[0169] Priorities and interruptions include tasks that need to be processed in a single processor and single thread. If each of the emotional right brain, the intellectual left brain, and the motor cerebellum has its own independent processor, then since there is no time conflict between their respective tasks, there is no need to set priorities and interruptions for tasks between brains.

[0170] 7. Includes VEM memory mechanism: Based on the aforementioned solution, this invention also includes a VEM memory mechanism: VEM memory mechanisms include, but are not limited to, one or a combination of person recognition mechanisms, event recording mechanisms, time memory mechanisms, multi-headed hotspot attention mechanisms, excitation-inhibition mechanisms, dormancy-awakening mechanisms, selective memory-forgetting mechanisms, and are stored in an emotional feature database.

[0171] The recognition mechanism includes, but is not limited to, a one-to-many relationship memory when a simulated human engages in emotional communication with multiple communicators. The memory content includes, but is not limited to, the craniofacial features, voice features, relationship features between the communicator and the simulated human, and communicator ID.

[0172] The recording mechanism includes, but is not limited to, extracting the content of the emotional exchange between the simulated human and the communicator, setting hot topics and hot function values, and recording the timestamp of the exchange, provided that the communicator has been identified.

[0173] The time memory mechanism includes, but is not limited to, setting a decay function based on the length of time since the communication timestamp, calculating the decay residual for hot content and hot function values, and marking the memory as valid when the residual is greater than the set value; otherwise, marking the memory as forgotten.

[0174] The multi-head hotspot attention mechanism includes, but is not limited to, setting two or more hotspot contents and hotspot function values. When two or more hotspot contents appear consecutively in a single search, the attenuation function weight is modified, and the attenuation residual is calculated. When the residual is greater than the set value, the memory is marked as valid; otherwise, the memory is marked as forgotten.

[0175] The excitation-inhibition mechanism includes, but is not limited to, setting excitation and inhibition values ​​for event time, event frequency, and event type, calculating the excitation-inhibition function, calculating the function residual, and marking the memory as valid when the residual is greater than the set value, otherwise marking the memory as forgotten.

[0176] The hibernation-wake mechanism includes, but is not limited to, setting hibernation and wake-up values ​​based on event time and event frequency, and calculating the hibernation-wake function.

[0177] Selective memory-forgetting mechanisms include, but are not limited to, setting memory events and forgetting events based on the personality traits of simulated humans, retrieving control events, and labeling memory and forgetting values.

[0178] As an emotional right-brain model, the VEM memory mechanism is crucial. It not only determines how well the robot "acts like a human," "recognizes people," and "remembers events" during interactions with real people, but it can even learn "human relationships," making the robot's interactions with real people more emotional and humane, and making communication more like human interaction.

[0179] In the VEM memory mechanism, learning is a commonly used method. Here, learning refers to various forms of memorization and selection. Learning methods include supervised learning, unsupervised learning, reinforcement learning, and deep learning.

[0180] 8. Including the dialogue relationship between the robot and the communicator: Based on the aforementioned solution, the present invention further includes the following regarding the dialogue relationship between the robot and the communicator: The robot establishes a one-to-one relationship with the communicator, using a human recognition mechanism to identify a communicator. It establishes a logical relationship and spatial visual positioning between the robot and the communicator based on the perceptual spectrum, deductive spectrum, and questions and answers. Based on the questions in the perceptual spectrum, it calculates the answers in the deductive spectrum and outputs them to the craniofacial and limb areas. Alternatively, based on the questions in the deductive spectrum and the answers in the perceptual spectrum, it determines the logic and completeness of the dialogue and the spatial visual positioning of the robot when it looks at the communicator during the dialogue.

[0181] The robot and the communicator form a one-to-many relationship. Based on the logical relationship between questions and answers in the perceptual spectrum and deductive spectrum, the robot records and buffers the questions or answers of the communicator, and outputs the robot's perceptual spectrum or deductive spectrum one by one.

[0182] The relationship between robots and communicators is a many-to-one relationship. Based on the dialogue relationship between multiple robots and one communicator, dialogue agreements are established for multiple robots, including question-based dialogues, agreed-upon robot dialogues, and custom robot dialogues.

[0183] The relationship between the robot and the communicator is many-to-many. The dialogue relationship between the robot and the communicator is agreed upon in advance, and is decomposed, determined and executed as a one-to-one or one-to-many dialogue relationship.

[0184] The switching of the relationship between the robot and the communicator is determined by setting or negotiation between the robot and the communicator, and the relationship switching of the next dialogue segment is determined on a segment-by-segment basis, including switching the simulator and switching the communicator.

[0185] The dialogue relationship between the robot and the communicator includes mimicking human-to-human dialogue. It is important to emphasize here that: 1. It is necessary to always be clear about "Who am I?", that is, "Who is the simulated person?". It is recommended to check this point when starting each dialogue segment.

[0186] 2. Always be clear about "Who am I talking to?", that is, who is the communicator? Remember that different communicators have different dialogue logic, content, and emotional scale. It is recommended to review this when starting each dialogue segment.

[0187] 3. It is necessary to remember the content before the current dialogue segment time point and calculate and predict the content after the current dialogue segment time point, and then review and verify it to avoid errors in logical content and emotional scale.

[0188] 9. Includes craniofacial region and limbs: Based on the aforementioned solutions, the present invention further includes, on the craniofacial region and limbs: Craniofacial systems include, but are not limited to, the VEM-robot craniofacial end-to-end emotional bionic system; The craniofacial structure also includes, but is not limited to, actuators mounted on a craniofacial support that conform to the shape and movement habits of a simulated human, such as simulated eyeballs, simulated skin, simulated mouth, simulated ears, and simulated neck, used to mimic the output process of dynamic facial expressions and vocal output of a simulated human. It also includes, but is not limited to, the spatial visual positioning of a simulated human visual communicator, and is used to capture video input signals within the field of view and record audio input signals within the auditory range.

[0189] The craniofacial region also includes, but is not limited to, spatial coordinate sensors, motion sensors, navigation sensors, time sensors, and communication interfaces. Through the communication interface, craniofacial signals are transmitted as a perceptual spectrum to the emotional right brain model.

[0190] Limbs include, but are not limited to, actuators and limb interfaces of movable parts such as the robot's limbs, arms, legs, neck, and waist. The limb interface is connected to the emotional output, performs the communication of the emotional output, and generates body expressions.

[0191] Preferably, the deductive spectrum is connected to the craniofacial actuator and the limb actuator through a communication interface to realize the output of the simulated human's emotional communication, and the emotional communication of the communicator is captured by the craniofacial sensor and the input is converted into a perceptual spectrum.

[0192] As previously mentioned, in the VEM-robot craniofacial end-to-end emotional bionic system, the robotic craniofacial system includes actuators mounted on a craniofacial support, such as bionic eyeballs, bionic skin, bionic mouth, bionic ears, and bionic neck.

[0193] The end-to-end emotional bionic system captures visual, auditory, lip-touch, and skin-touch input environmental perception information from bionic eyes, bionic ears, bionic mouth, and bionic skin, generates emotional bionic facial expression commands, and outputs them from bionic eyes, bionic skin, bionic mouth, bionic ears, and bionic neck, producing eye movements, eyelid movements, facial skin movements, sound output, jaw movements, mouth movements, and neck movements, realizing end-to-end bionic facial expression output from input to output.

[0194] The 3D movements of the robot's limbs include the neck, waist, arms, hands, feet (for humanoid robots), and lower limbs (for wheeled robots). These movements are essential for the robot's emotional communication, making it more lifelike and enhancing its approachability.

[0195] 10. Includes clone mode: Based on the aforementioned scheme, the present invention includes a cloning mode: The hardware and operation of the emotional right brain model are located locally on the robot.

[0196] The emotional feature database is stored in the robot's local hardware.

[0197] In the intelligent left-brain model, the large language model is either placed locally or connected to the cloud center via the external network.

[0198] Cloning mode includes, but is not limited to, replicating one humanoid into two or more in real time, allowing two or more humanoids to engage in emotional communication with the communicator simultaneously.

[0199] Preferably, in the monitoring mode, in the clone mode, a simulated human is ensured to receive only the perceptual spectrum and not generate the deductive spectrum.

[0200] As a special case of industrial products, the robot's brain and operating system, VEM-ROS, are located within the robot's body, while its clone resides in a cloud center. As another special case, based on machine training, machine learning, and backup, this type of clone can be configured as a fully background clone that does not produce a true deductive spectrum output. This allows the robot to be inherited and reactivated via a completely new hardware clone located in the cloud center after physical damage.

[0201] 11. The emotional right-brain model is encapsulated using an operating system. Based on the aforementioned solution, this invention includes an operating system encapsulation in the emotional right-brain model: Based on the EOS operating system for emotional robots, the emotional right-brain model is rewritten and encapsulated into the emotional multimodal VEM-ROS and merged into the EOS operating system, forming a part of EOS, specifically including but not limited to: Rewrite and encapsulate the physical drivers of sensors and actuators into interrupt or process modes.

[0202] Based on network communication protocols, rewrite and encapsulate emotion expressions to conform to network communication protocols, including but not limited to ISO-OSI protocol, AI agent protocol, model and tool protocol, agent-to-agent collaboration protocol, agent-to-user interface protocol.

[0203] Based on priority and interruption, rewrite and encapsulate the definition of emotional right brain functions.

[0204] As an operating system for emotion-based robots, EOS considers the emotion model VEM-ROS as only a local part or a subset of the set within EOS. In the future, EOS will also include a world model based on spatial images, as well as other plug-and-play models.

Claims

1. A method for constructing a VEM-robot emotional right brain model, characterized by: A robot designed to mimic a human being engages in emotional communication with one or more interlocutors, including real people or other robots. An emotional feature database records and distinguishes the emotional characteristics of the mimic and the interlocutor. The robot includes an emotional right brain, wherein: The Emotional Right Brain Model uses a beat segmentation and alignment method to decompose the perceptual spectrum emitted by the communicator and perceived by the robot into emotional perception and logical perception, including logical expression. The right-brain model of emotion, based on emotion perception and an emotion feature database, uses emotion processing to generate emotional expression, which is then synchronized with logical expression through rhythmic synthesis to produce a deductive spectrum. The deductive spectrum output to the robot's craniofacial region forms emotional communication through facial expressions and vocal expressions, and / or, the deductive spectrum output to the robot's limbs forms emotional communication through movements.

2. The method according to claim 1, characterized in that, The emotional right-brain model also includes: The VEM-Token vocal emotion multimodal method is used to construct an emotion feature library and VEM-Token lexical units. The emotion feature library includes one or a combination of emotion, music, rhythm, imagination, creativity and art. We adopted a method of constructing a VEM-Token beat capture and alignment model to complete the calibration and alignment of VEM-Token lexical beats, and used a method of hierarchical fusion of VEM-Token sentiment synchronization function to obtain sentiment function; The VEM coordinate system is used to manage multimodal sentiment components, where: Emotional components include one or a combination of multiple modalities such as joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, affection, and enmity; The emotional component also includes the weighting function of each component, which is determined by setting or learning based on the emotional characteristics of the simulated human. The sentiment function includes sentiment components and weighting functions. The sentiment function is a dependent variable with sentiment components and weighting functions as independent variables. The sentiment components and sentiment function are incorporated into the sentiment feature library. The generation of the sentiment feature database also includes methods such as initial definition, supervised learning, learning of the emotional communication process, and generation of one or a combination of long-term attention mechanisms.

3. The method according to claim 2, characterized in that, It also includes a beat model, specifically including: The beat model is divided using a VEM-Token beat capture and alignment model construction method, specifically including: The music beat model, based on the VEM-Token model and beat theory in music theory, marks the start and end points of the music beat in the perceptual spectrum and the deductive spectrum according to the start and end points of the music beat. The punctuation rhythm model, based on the LLM-Token model and the grammatical habits of text sentences, marks the start and end points of the punctuation rhythm according to the start and end points of the punctuation; The paragraph beat model, based on a complete dialogue logic paragraph and dialogue habits, marks the start and end points of the paragraph beat according to the start and end points of the paragraph's perceptual spectrum and deductive spectrum. Beat conflict fine-tuning: When the perceptual spectrum is aligned by beat segmentation to form emotional perception and logical perception, and when the emotional expression and logical expression are aligned by beat synthesis to form the deductive spectrum, the start and end points are fine-tuned to achieve beat alignment. The beat capture method also includes converting the perceptual spectrum and deductive spectrum into a spectrum format file, then decomposing them into harmonic component spectrum and impulse component spectrum, detecting abrupt changes in spectral energy, and finally determining the start point of the beat.

4. The method according to claim 3, characterized in that, Deductive spectrum and perceptual spectrum also include: S41: For a given punctuation beat, the text included in its logical expression and the emotional components and weighting functions included in the emotional expression of the corresponding musical beat are aligned through beat synthesis to generate a punctuation beat's interpretation spectrum. S42: For a dialogue segment, execute S41 repeatedly to complete the deductive spectrum of a dialogue segment; S43: Output the deduction spectrum of a dialogue segment, or perform sentiment function smoothing verification. Specifically, within a dialogue segment, compare the fluctuation of the sentiment function in consecutive punctuation beats. If the fluctuation residual is less than a preset value, the verification is considered successful, and the deduction spectrum is output. If the fluctuation residual is greater than the preset value, the verification is considered unsuccessful, and the sentiment function is smoothed. The deduction spectrum is then output. S44: For a dialogue segment, keep the logical expression unchanged, unconditionally change the emotional components and weighting functions in the emotional expression, and generate a modified deductive spectrum. S45: The deductive spectrum is output to the robot's craniofacial region or the robot's craniofacial region and limbs in units of a dialogue segment to complete the emotional communication of a dialogue segment. S46: In the emotional perception and logical perception decomposed by the perceptual spectrum, the beat capture method is used to determine the start and end points of the musical beat and the pause beat, and the beat conflict fine-tuning method is used to align the start and end points of the musical beat and the pause beat. S47: The deductive spectrum and perceptual spectrum also include the question-and-answer mode. When the communicator sends out a question, the robot's answer is given the highest priority. If the communicator does not answer within a set time after the robot sends out a question, the question is suspended and listed as a pending issue. S48: The perception spectrum is obtained by sensors including the robot's eyes, ears, skin, and spatial sensing. The deductive spectrum is output by actuators including the robot's speaker sound output, craniofacial skin expression dynamic output, eye movement, eyelid movement, and mouth movement.

5. The method according to claim 1, characterized in that, Robots also include operating systems: The robot also includes an intelligent left brain and a dynamic cerebellum. The intelligent left brain is used for logical reasoning and analysis, language and text comprehension, knowledge question answering, task planning and training memory. The intelligent left brain also includes working with a large language model using LLM-Token. The dynamic cerebellum is used for limb balance control, movement coordination, reflexive movement and training memory. The robot's brain consists of an emotional right brain, an intelligent left brain, and a motor cerebellum. These three components are integrated to form the robot's brain operating system. The emotional right brain model communicates bidirectionally with the intelligent left brain through an emotion-logic interface, and the emotional right brain model communicates bidirectionally with the motor cerebellum through an emotion-action interface. The emotional right brain model works in collaboration with the intelligent left brain and the motor cerebellum to achieve end-to-end emotional simulation and emotional communication with the communicator. For the robot's emotional right brain, intellectual left brain, and motor cerebellum, there are options to have them operate with their own independent processors and operating system software, or to have them operate with a unified processor and operating system software for the emotional right brain, intellectual left brain, and motor cerebellum.

6. The method according to claim 5, characterized in that, Including priority and interruption: The brain-brain working priority includes at least the following: the kinematic cerebellum is the first priority, the emotional right brain and / or the intellectual left brain is the second priority, and the first priority is higher than the second priority; Brain-based work priorities include: high response speed priority, short response time priority, high importance priority, and custom priority. Interruption priorities include: when the robot is handling concurrent conflicting tasks, according to the priority order, high-priority tasks are allowed to interrupt low-priority tasks, and custom tasks are allowed to interrupt other tasks based on their priority. Interrupt recovery includes: before the original task is interrupted by the high-level interrupt task, the processor saves the interrupt context parameters of the original task, enters the execution of the high-level interrupt task, and after the high-level interrupt task is completed, the interrupt context is restored and the original interrupt task is restored.

7. The method according to claim 4 or 6, characterized in that, It also includes the VEM memory mechanism: VEM memory mechanisms include one or a combination of person recognition mechanism, event recording mechanism, time memory mechanism, multi-head hot spot attention mechanism, excitation-inhibition mechanism, dormancy-awakening mechanism, selective memory-forgetting mechanism, and are stored in the emotional feature database; The person recognition mechanism includes a one-to-many relationship memory when a simulated human engages in emotional communication with multiple communicators. The memory content includes the craniofacial features of the communicators, voice features, the relationship features between the communicators and the simulated human, and the communicator's number. The recording mechanism includes extracting the content of the emotional exchange between the simulated human and the communicator, setting hot topics and hot function values, and recording the timestamp of the exchange, given that the communicator has been identified; The time memory mechanism includes setting a decay function based on the length of time since the communication timestamp, for hot content and hot function values, calculating the decay residual, and marking the memory as valid when the residual is greater than the set value, otherwise marking the memory as forgotten. The multi-head hotspot attention mechanism includes setting two or more hotspot contents and hotspot function values. When two or more hotspot contents appear consecutively in a single search, the attenuation function weight is modified and the attenuation residual is calculated. When the residual is greater than the set value, the memory is marked as valid; otherwise, the memory is marked as forgotten. The excitation-inhibition mechanism includes setting excitation and inhibition values ​​for event time, event frequency, and event type, calculating the excitation-inhibition function, calculating the function residual, and marking the memory as valid when the residual is greater than the set value; otherwise, the memory is marked as forgotten. The sleep-wake mechanism includes setting sleep and wake-up values ​​based on event time and event frequency, and calculating the sleep-wake function; Selective memory-forgetting mechanism includes setting memory events and forgetting events based on the personality characteristics of simulated humans, retrieving control events, and labeling memory and forgetting values.

8. The method according to claim 7, characterized in that, The dialogue relationship between robots and communicators also includes: The robot establishes a one-to-one relationship with the communicator, using a human recognition mechanism to identify a communicator. It establishes the logical relationship and spatial visual positioning between the robot and the communicator in terms of the perceptual spectrum, deductive spectrum, questions, and answers. Based on the questions in the perceptual spectrum, it calculates the answers in the deductive spectrum and outputs them to the craniofacial and limb regions. Alternatively, based on the questions in the deductive spectrum and the answers in the perceptual spectrum, it determines the logic and completeness of the dialogue and the spatial visual positioning of the robot when it looks at the communicator during the dialogue. The robot and the communicator have a one-to-many relationship. Based on the logical relationship between the questions and answers in the perceptual spectrum and deductive spectrum, the robot's questions or answers to the communicator are recorded and buffered, and the robot's perceptual spectrum or deductive spectrum is output one by one. The relationship between robots and communicators is a many-to-one relationship. Based on the dialogue relationship between multiple robots and one communicator, dialogue agreements are established for multiple robots, including question-based dialogue, agreed-upon robot dialogue, and custom robot dialogue. The many-to-many relationship between the robot and the communicator is defined in advance, and the dialogue relationship between the robot and the communicator is decomposed, determined and executed as a one-to-one or one-to-many dialogue relationship. The switching of the relationship between the robot and the communicator is determined by setting or negotiation between the robot and the communicator, and the relationship switching of the next dialogue segment is determined on a segment-by-segment basis, including switching the simulator and / or switching the communicator.

9. The method according to claim 8, characterized in that, The craniofacial region and limbs also include: The craniofacial system includes the VEM-robotic craniofacial end-to-end emotional bionic system; The craniofacial structure also includes actuators mounted on the craniofacial support that conform to the shape and movement habits of a simulated human. These actuators include simulated eyeballs, simulated skin, simulated mouth, simulated ears, and simulated neck, which are used to mimic the output process of dynamic facial expressions and vocal output of a simulated human. It also includes the spatial visual positioning of a simulated human visual communicator, as well as the ability to capture video input signals within the field of view and record audio input signals within the auditory range. The craniofacial region also includes spatial coordinate sensors, motion sensors, navigation sensors, time sensors, and communication interfaces. The craniofacial signals are transmitted as a perceptual spectrum to the emotional right brain model through the communication interface. Limbs include actuators and limb interfaces for movable parts of the robot's limbs, arms, legs, neck, and waist. The limb interface is connected to the emotional output, performs the communication of the emotional output, and generates body expressions. The deductive spectrum connects the actuators of the craniofacial region and the limbs through a communication interface to realize the output of emotional communication of a simulated human. It also captures the emotional communication of the communicator through craniofacial sensors and converts the input into a perceptual spectrum.

10. The method according to claim 9, characterized in that, Including clone mode: The hardware and operation of the emotional right brain model are located locally on the robot; The emotional feature database is stored in the robot's local hardware; In the intelligent left-brain model, the large language model is either placed locally or connected to the cloud center via an external network. The cloning mode involves replicating one android into two or more in real time, allowing two or more androids to engage in emotional communication with the communicator simultaneously. In the monitoring mode, which is in clone mode, a simulated human is ensured to receive only the perceptual spectrum and not generate the deductive spectrum.

11. The method according to claim 10, characterized in that, The emotional right-brain model includes operating system encapsulation; Based on the EOS operating system for emotional robots, the emotional right-brain model is rewritten and encapsulated into the emotional multimodal VEM-ROS, which is then merged into the EOS operating system to form a part of EOS. Specifically, this includes: Rewrite and encapsulate the physical drivers of sensors and actuators into interrupt or process mode; Based on network communication protocols, we rewrite and encapsulate emotion expressions to conform to network communication protocols, including ISO-OSI protocol, AI agent protocol, model and tool protocol, agent-to-agent collaboration protocol, and agent-to-user interface protocol. Based on priority and interruption, rewrite and encapsulate the definition of emotional right brain functions.

Citation Information

Patent Citations

  • CN121122100A

  • CN121756368A

  • CN121864645A