Facial robot expression and action accurate control system based on multi-modal perception

By integrating information from multiple perception modes in the facial robot system, the problem of traditional systems lacking understanding and response capabilities for complex environments and user commands is solved, and more accurate and natural expression and motion control is achieved, improving the level of intelligence and interactive experience.

CN120206513APending Publication Date: 2025-06-27ZHENGZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510320972.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The traditional facial robot expression motion control system only relies on a single perception mode, resulting in insufficient accurate understanding and response capabilities of complex environments and user instructions, limiting the robot's intelligence level and interactive experience.

Method used

A system based on multimodal perception is adopted to integrate information from multiple perception modes such as vision, sound, tactile and bioelectric signals. Through the multimodal fusion processing module and expression action control module, accurate and natural expression action control is achieved.

Benefits of technology

By integrating information from multiple perception modes, accurate control of facial robot expressions and movements is achieved, improving the robot's intelligence level and interactive experience, and being able to understand and respond to complex environments and user instructions more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120206513A_ABST
    Figure CN120206513A_ABST
Patent Text Reader

Abstract

The invention discloses a facial robot expression and action precise control system based on multi-mode perception, and the system comprises a multi-mode perception module which is used for integrating the information of various perception modes, such as vision, sound, touch and bio-electricity signals at the same time; the multi-mode sensing module comprises a visual sensing module, a sound sensing module, a touch sensing module and a bio-electricity signal sensing module. According to the method, information of various sensing modes such as vision, sound, touch and bio-electricity signals can be integrated at the same time, comprehensive analysis is carried out through an advanced artificial intelligence algorithm, and therefore accurate and natural control over facial expressions and actions of the facial robot is achieved. The intelligent level and interaction experience of the robot can be greatly improved, and the robot can better adapt to complex and changeable environments and user requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of facial robots, and specifically to a precise control system for facial robot expression actions based on multi-modal perception. Background Art

[0002] With the rapid development of robot technology, facial robots are increasingly widely used in multiple fields such as human-computer interaction, entertainment, education, and medical treatment. However, traditional facial robot expression action control systems often rely only on a single perception modality, such as vision or sound, which greatly limits the robot's accurate understanding and response to complex environments and user instructions. In order to improve the intelligence level and interaction experience of facial robots, there is an urgent need for a system that can integrate multiple perception modalities to achieve more precise and natural expression action control;

[0003] Therefore, we propose a precise control system for facial robot expression actions based on multi-modal perception. Summary of the Invention

[0004] The purpose of the present invention is to provide a precise control system for facial robot expression actions based on multi-modal perception, and solve the problems raised in the background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A precise control system for facial robot expression actions based on multi-modal perception, including a multi-modal perception module for simultaneously integrating information from multiple perception modalities such as vision, sound, touch, and bioelectrical signals;

[0006] The multi-modal perception module includes a visual perception module, a sound perception module, a touch perception module, and a bioelectrical signal perception module;

[0007] It also includes: a multi-modal fusion processing module and an expression action control module;

[0008] The visual perception module is used to accurately perceive the user's facial expressions, actions, and environmental background;

[0009] The sound perception module is used to collect environmental sounds omnidirectionally and with high sensitivity;

[0010] The touch perception module is used to achieve precise perception of facial touches;

[0011] The bioelectrical signal perception module is used to capture the electrical activities of the user's facial muscles and perform feature extraction and classification in combination with machine learning algorithms;

[0012] The multi-modal fusion processing module fuses and processes the perception information from the visual, sound, touch, and bioelectrical signal modules;

[0013] The expression and action control module is used to drive the expression and action mechanism of the facial robot for precise and natural action control, so as to achieve rich and diverse facial expressions and actions.

[0014] As a preferred embodiment of the present invention, the multimodal fusion processing module includes a feature extraction sub-module and a feature fusion sub-module. The feature extraction sub-module is used to extract multimodal features, and the feature fusion sub-module is used to fuse the extracted multimodal features.

[0015] As a preferred embodiment of the present invention, the feature fusion sub-module includes a fusion model. The specific operation of the fusion model is as follows:

[0016] Calculation of attention weights:

[0017] For N modalities, for the i-th modality, its attention weight α i can be calculated by the following formula:

[0018] e i = score(F i , Q)

[0019]

[0020] where F i is the feature vector of the i-th modality, Q is the query vector, and score is a function for calculating the score.

[0021] Extract n perceptual modalities, the feature vector of each modality is F i , and the weight is α i ;

[0022] The fused feature vector F fused is:

[0023]

[0024] As a preferred embodiment of the present invention, the expression and action control module includes a high-precision action driving model, and the high-precision action driving model includes:

[0025] Generator: used to generate realistic facial expression actions;

[0026] Discriminator: used to judge whether the generated facial expression actions are realistic;

[0027] During the model training process, the generator tries to generate realistic facial expression actions to deceive the discriminator, while the discriminator tries to distinguish the generated expression actions from the real expression actions. Through continuous adversarial training, the generator can gradually improve the authenticity of the generated expression actions.

[0028] As a preferred embodiment of the present invention, the goal of the generator is to generate realistic facial expression actions according to the input expression parameters, and the generator algorithm is as follows:

[0029] During the training process, the loss function of the generator can be expressed as:

[0030]

[0031] where z is the input expression parameter vector, G(z) is the facial expression action generated by the generator, D(G(z)) is the judgment probability of the discriminator for the generated expression action, y is the real facial expression action, and λ is the regularization parameter. The first term is the loss of the generator deceiving the discriminator, and the second term is the mean square error loss between the generated expression action and the real expression action.

[0032] As a preferred embodiment of the present invention, the goal of the discriminator is to distinguish between the generated facial expression actions and the real facial expression actions, and the discriminator algorithm is as follows:

[0033] During the training process, the loss function of the discriminator can be expressed as:

[0034] LD = E x~pdata(x) [log(D(x))] + E z~pz(z) [log(1 - D(G(z)))];

[0035] where x is the real facial expression action, D(x) is the judgment probability of the discriminator for the real expression action, z is the input expression parameter vector, G(z) is the facial expression action generated by the generator, the first term is the judgment loss of the discriminator for the real expression action, and the second term is the judgment loss of the discriminator for the generated expression action.

[0036] As a preferred embodiment of the present invention, the visual perception unit includes:

[0037] A binocular infrared camera for detecting human face key points in a dark environment;

[0038] A structured light depth sensor configured to reconstruct a three-dimensional facial model at a rate of 60 frames per second.

[0039] As a preferred embodiment of the present invention, the tactile perception unit includes a capacitive touch array, a piezoresistive force sensor, and a tactile feature encoder.

[0040] As a preferred embodiment of the present invention, the bioelectrical signal perception unit includes a flexible electromyogram sensor array for attaching to the inner side of the bionic skin of the robot to detect changes in the intensity of muscle electrical signals.

[0041] As a preferred embodiment of the present invention, the sound perception module is an array microphone.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] The present invention can simultaneously integrate information of multiple perception modalities such as vision, sound, touch, and bioelectric signals, and fuse multiple signals through a fusion model to achieve comprehensive analysis, so as to accurately control the facial robot's expression actions;

[0044] Through the high-precision motion driving model in the robot surface control module, the fused multi-modal information can be converted into high-precision control signals, enabling the robot to more accurately display precise facial expressions and actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0046] Figure 1 It is an operation diagram of the precise control system for facial robot expression actions based on multi-modal perception of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0047] In order to make the technical means, creative features, achieved purposes, and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0048] Embodiment

[0049] A precise control system for facial robot expression actions based on multi-modal perception includes a multi-modal perception module for simultaneously integrating information of multiple perception modalities such as vision, sound, touch, and bioelectric signals;

[0050] The multi-modal perception module includes a visual perception module, a sound perception module, a touch perception module, and a bioelectric signal perception module;

[0051] It further includes: a multi-modal fusion processing module and an expression action control module;

[0052] The visual perception module perceives the user's facial expressions, actions, and environmental background in real time and with high precision;

[0053] The sound perception module collects environmental sounds omnidirectionally and with high sensitivity;

[0054] The touch perception module is used to accurately perceive facial touches;

[0055] The bioelectric signal sensing module is used to capture the electrical activities of the user's facial muscles and perform feature extraction and classification in combination with machine learning algorithms;

[0056] The multimodal fusion processing module fuses and processes the perception information from the vision, sound, touch, and bioelectric signal modules;

[0057] The facial expression and action control module is used to drive the facial expression and action mechanism of the facial robot to perform precise and natural action control, realizing rich and diverse facial expressions and actions.

[0058] Among them, the visual perception unit includes: a binocular infrared camera for detecting facial key points in a dark environment; a structured light depth sensor configured to reconstruct a three-dimensional facial model at a rate of 60 frames per second. The tactile perception unit includes a capacitive touch array, a piezoresistive force sensor, and a tactile feature encoder. The bioelectric signal sensing unit includes a flexible electromyography sensor array for attaching to the inner side of the bionic skin of the robot to detect changes in muscle electrical signal intensity. The sound perception module is an array microphone.

[0059] Moreover, the multimodal fusion processing module includes a feature extraction sub-module and a feature fusion sub-module. The feature extraction sub-module is used to extract features of multiple modalities, and the feature fusion sub-module is used to fuse the extracted multimodal features. The feature fusion sub-module performs fusion on the fusion model. The specific operations of the fusion model are as follows:

[0060] Calculation of attention weights:

[0061] For N modalities, for the i-th modality, its attention weight α i can be calculated by the following formula:

[0062] e i = score(F i , Q)

[0063]

[0064] where F i is the feature vector of the i-th modality, Q is the query vector, and score is a function for calculating the score.

[0065] Extract n types of perception modalities, and the feature vector of each modality is F i , and the weight is α i ;

[0066] The fused feature vector F fused is:

[0067]

[0068] The facial expression and action control module includes a high-precision action driving model, and the high-precision action driving model includes:

[0069] A generator: used to generate realistic facial expression actions;

[0070] A discriminator: used to judge whether the generated facial expression actions are realistic;

[0071] During the model training process, the generator attempts to generate realistic facial expression actions to deceive the discriminator, while the discriminator tries to distinguish between the generated expression actions and real expression actions. Through continuous adversarial training, the generator can gradually improve the authenticity of the generated expression actions.

[0072] The goal of the generator is to generate realistic facial expression actions based on the input expression parameters. The generator algorithm is as follows:

[0073] During the training process, the loss function of the generator can be expressed as:

[0074]

[0075] where z is the input expression parameter vector, G(z) is the facial expression action generated by the generator, D(G(z)) is the judgment probability of the discriminator for the generated expression action, y is the real facial expression action, λ is the regularization parameter, the first term is the loss of the generator deceiving the discriminator, and the second term is the mean square error loss between the generated expression action and the real expression action.

[0076] The goal of the discriminator is to distinguish between the generated facial expression actions and real facial expression actions. The discriminator algorithm is as follows:

[0077] During the training process, the loss function of the discriminator can be expressed as:

[0078] LD = E x~pdata(x) [log(D(x))] + E z~pz(z) [log(1 - D(G(z)))];

[0079] where x is the real facial expression action, D(x) is the judgment probability of the discriminator for the real expression action, z is the input expression parameter vector, G(z) is the facial expression action generated by the generator, the first term is the judgment loss of the discriminator for the real expression action, and the second term is the judgment loss of the discriminator for the generated expression action.

[0080] The specific implementation of the system is as follows:

[0081] First, the binocular infrared camera captures the infrared image of the user's face and performs face key point detection. The structured light depth sensor reconstructs the three-dimensional model of the user's face, providing facial shape and motion data. The array microphone collects environmental sounds, enhances the target sound signal through beamforming technology, preprocesses the sound signal, and extracts the sound spectrum features. The capacitive touch array detects the slight touch on the face and outputs a touch signal. The piezoresistive force sensor measures the intensity of the touch and outputs an intensity signal. The tactile feature encoder encodes the touch signal and the intensity signal into digital signals. The flexible electromyogram sensor array detects the electrical activity of the user's facial muscles and outputs an electromyogram signal, and preprocesses the electromyogram signal to extract the features of muscle electrical activity;

[0082] Then, preprocess the raw data such as visual, sound, tactile, and bioelectrical signals, extract the feature vectors of each modality, calculate the attention weights of each modality using the attention mechanism, and weighted fuse the feature vectors of each modality according to the attention weights to obtain the fused feature vector F fused ,

[0083] Finally, the generator in the facial expression and motion control module generates realistic facial expression and motion according to the fused feature vector F fused The generator continuously optimizes the generated facial expression and motion to improve its authenticity and naturalness. The discriminator judges the generated facial expression and motion to distinguish whether it is realistic. According to the feedback of the discriminator, the generator adjusts the generation strategy to further improve the quality of the generated facial expression and motion. The facial expression and motion control module converts the generated facial expression and motion into control signals to drive the facial expression and motion mechanism of the facial robot to act. The facial robot makes accurate and natural facial expressions and motions according to the control signals.

[0084] The above shows and describes the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention.

[0085] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A precise control system for facial robot expression movements based on multimodal perception, characterized by: Includes a multimodal perception module for simultaneously integrating information from multiple perception modalities such as vision, sound, touch, and bioelectric signals; The multimodal perception module includes a visual perception module, a sound perception module, a tactile perception module, and a bioelectric signal perception module; It also includes: a multimodal fusion processing module and an expression action control module; The visual perception module is used to perceive the user's facial expressions, movements and environmental background with high precision; The sound sensing module is used to collect environmental sounds in an all-round and highly sensitive manner; The tactile sensing module is used to achieve accurate perception of facial touch; The bioelectric signal sensing module is used to capture the electrical activity of the user's facial muscles and perform feature extraction and classification in combination with a machine learning algorithm; The multimodal fusion processing module fuses and processes the perception information from the vision, sound, touch and bioelectric signal modules; The expression and motion control module is used to drive the expression and motion mechanism of the facial robot to perform precise and natural motion control, thereby realizing rich and diverse facial expressions and motions.

2. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The multimodal fusion processing module includes a feature extraction submodule and a feature fusion submodule. The feature extraction submodule is used to extract multimodal features, and the feature fusion submodule is used to fuse the extracted multimodal features.

3. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The feature fusion submodule includes a fusion model, and the fusion model fusion specific operations are as follows: Attention weight calculation: N modalities, for the i-th modality, its attention weight α i It can be calculated by the following formula: e i =score(F i ,Q) Among them, F i is the feature vector of the i-th mode, Q is the query vector, and score is the function for calculating the score. Extract n perceptual modalities, and the feature vector of each modality is F i , with weight α i ; The fused feature vector F fused for:

4. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The expression action control module includes a high-precision action drive model, and the high-precision action drive model includes: Generator: used to generate realistic facial expressions; Discriminator: used to judge whether the generated facial expressions are realistic; During the model training process, the generator tries to generate realistic facial expressions to deceive the discriminator, while the discriminator strives to distinguish between the generated expressions and real expressions. Through continuous adversarial training, the generator can gradually improve the authenticity of the generated expressions.

5. The facial robot expression and motion precision control system based on multimodal perception according to claim 4 is characterized by: The goal of the generator is to generate realistic facial expressions based on the input expression parameters. The generator algorithm is as follows: During the training process, the loss function of the generator can be expressed as: Where z is the input expression parameter vector, G(z) is the facial expression generated by the generator, D(G(z)) is the probability of the discriminator judging the generated expression, y is the real facial expression, λ is the regularization parameter, the first term is the loss of the generator deceiving the discriminator, and the second term is the mean square error loss between the generated expression and the real expression.

6. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The goal of the discriminator is to distinguish between generated facial expressions and real facial expressions. The discriminator algorithm is as follows: During the training process, the loss function of the discriminator can be expressed as: LD=E x~pdata(x) [log(D(x))]+E z~pz(z) [log(1-D(G(z)))]; Among them, x is the real facial expression, D(x) is the probability of the discriminator's judgment on the real expression, z is the input expression parameter vector, G(z) is the facial expression generated by the generator, the first term is the judgment loss of the discriminator on the real expression, and the second term is the judgment loss of the discriminator on the generated expression.

7. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The visual perception unit comprises: Binocular infrared camera, used for facial key point detection in dark environments; A structured light depth sensor configured to reconstruct a 3D model of a face at 60 frames per second.

8. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The tactile sensing unit includes a capacitive touch array, a piezoresistive force sensor and a tactile feature encoder.

9. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The bioelectric signal sensing unit includes a flexible myoelectric sensor array, which is used to be attached to the inner side of the robot's bionic skin to detect changes in muscle electrical signal intensity.

10. The facial robot expression and motion precision control system based on multimodal perception according to claim 1, characterized in that: The sound perception module is an array microphone.

Citation Information

Cited By

  • Bionic robot head system

    CN121200002A