Facial expression control method, device and equipment for anthropomorphic robot, medium and product

By inserting a silent delay into the robot's facial expression control, the rhythm of human cognitive processing is simulated, which solves the problem of the mechanical feel of robot facial expressions and movements, and achieves a more natural and friendly human-computer interaction.

CN121997973APending Publication Date: 2026-05-08SONGYAN POWER (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SONGYAN POWER (BEIJING) TECHNOLOGY CO LTD
Filing Date
2026-01-08
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing robot facial expression control technology lacks the rhythm of natural human reactions, resulting in mechanical and abrupt facial expressions and movements, and exacerbating the uncanny valley effect.

Method used

By inserting a silence delay into the facial expression sequence, the ratio of silence duration to execution duration is controlled within a preset range to simulate the rhythm of human cognitive processing.

Benefits of technology

It significantly improves the naturalness and lifelikeness of robot expressions, alleviates the uncanny valley effect, and enhances the warmth and harmony of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997973A_ABST
    Figure CN121997973A_ABST
Patent Text Reader

Abstract

The invention discloses an anthropomorphic robot facial expression control method, device and equipment, a medium and a product, and relates to the technical field of anthropomorphic humans. The method comprises the following steps: firstly, receiving an external expression trigger instruction, then planning a corresponding facial expression action sequence, specifically, when each facial expression action in the sequence is planned, actively inserting a silence delay before executing the action, the duration of the silence delay being the duration from instruction receiving to action execution starting, and the duration of the silence delay being the duration from instruction receiving to action execution starting. The ratio of the stimulation-reaction delay to the execution duration of the corresponding action is controlled within a preset interval, and finally the robot face execution mechanism is driven based on the sequence to complete the expression, so that the mechanical and abrupt feeling caused by transient response in the prior art is overcome by actively introducing and quantitatively controlling the proportion of the stimulation-reaction delay to the action duration, and the safety of the robot is improved. And the natural cognitive processing rhythm of human is simulated, so that the facial organ movement of the robot better conforms to the movement mode and reaction rule of human, and the expression naturalness and life feeling are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bionic human technology, specifically relating to a method, device, equipment, medium, and product for controlling facial expressions of an anthropomorphic robot. Background Technology

[0002] In recent years, with the advancement of artificial intelligence and robotics, humanoid robots (especially social robots, entertainment robots, and service robots) have developed rapidly. To achieve natural, harmonious, and emotionally engaging human-computer interaction, endowing robots with the ability to simulate human emotions and intentions has become crucial. Therefore, facial expression generation and control technology for robots has become a current research hotspot and a key area for patent strategy.

[0003] Currently, existing technologies for improving the naturalness of robot facial expressions mainly follow the following two optimization directions: (1) Structural and material optimization, namely, some existing technologies are dedicated to improving the naturalness of the visual and tactile aspects by improving the physical structure and surface materials of the robot; for example, using more and more precise micro servo motors or shape memory alloys to increase the number of facial degrees of freedom; or using materials that are closer to the texture and elasticity of human skin, such as flexible silicone or elastomers, to make the robot's face skin; although these hardware-level improvement schemes can alleviate the static stiffness of facial expressions to a certain extent, they are essentially in the category of mechanical design and are usually accompanied by problems such as high cost, complex system design and difficult maintenance; more importantly, such methods cannot fundamentally solve the problems of "mechanical feel" and "lifeless feel" of the facial movement patterns themselves; (2) Optimization of driving and control algorithms, another mainstream approach focuses on improving driving and execution algorithms from the software and control level. For example, more complex PID (Proportional Integral Derivative) control algorithms, fuzzy control or neural network models are used to accurately plan and control the motion trajectory of the motor, aiming to make the execution of single or multiple facial expressions smoother and more precise. However, these methods generally ignore a key factor affecting naturalness - the timing and rhythm of human facial expressions.

[0004] In-depth research and analysis have revealed a common shortcoming and bottleneck in existing technologies: upon receiving facial expression commands, robots' facial execution systems respond and execute almost instantaneously. This pursuit of technological efficiency and a "zero-delay" reaction mode runs counter to the true physiological and cognitive response patterns of humans. In natural interpersonal interactions, after receiving external stimuli (such as hearing a sentence or seeing a scene), the brain requires a brief cognitive process of information processing, emotion generation, and intention formation before the facial muscles are driven by nerves to display the corresponding expression. This subtle yet essential cognitive processing and physiological preparation time between "stimulus" and "overt response" is precisely the core rhythmic characteristic that reflects the "authenticity" and "naturalness" of a living organism's response.

[0005] The current robot facial expression control technology lacks precisely this "human-like" responsiveness. While its instantaneous and seamless facial expression switching and execution achieve high efficiency at the control system level, it appears abrupt, strange, and even uncomfortable to human observers. This is considered one of the key reasons for exacerbating the "uncanny valley effect" and generating a strong "mechanical" feeling. Extensive research and verification have revealed that no publicly available patents or existing technologies fundamentally solve the problem of naturalness in robot facial expressions by actively introducing and systematically controlling the "stimulus-response" delay time and its proportion to the total action duration.

[0006] Therefore, there is an urgent need in this field for a robot facial expression control scheme that can simulate the rhythm of natural human reactions, in order to overcome the above-mentioned defects of existing technologies and promote the development of human-computer interaction experience of anthropomorphic robots in a more natural and friendly direction. Summary of the Invention

[0007] The purpose of this invention is to provide a method, device, computer equipment, computer-readable storage medium, and computer program product for controlling facial expressions of humanoid robots, in order to solve the problems of mechanical, abrupt, and unnatural movements caused by the "zero-delay" instantaneous response of existing robot facial expression control technologies.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: Firstly, a method for controlling facial expressions in a humanoid robot is provided, including: Receive facial expression trigger commands from external sources; Based on the facial expression triggering command, at least one facial expression action that needs to be executed independently and the execution duration of each facial expression action in the at least one facial expression action are determined, and a facial expression action sequence composed of the at least one facial expression action is planned in the following manner: For each facial expression action, a silent delay with a corresponding silent duration is inserted before the corresponding action is executed, wherein the silent duration refers to the duration from the time of receiving the facial expression triggering command to the start time of the execution of the corresponding action, and the ratio of the silent duration to the execution duration of the corresponding action belongs to a preset range; Based on the planned sequence of facial expression movements, the anthropomorphic robot is driven to perform corresponding facial expression movements through its facial actuators, which are one-to-one with each of the facial expression movements.

[0009] Based on the above-mentioned invention, a robot facial expression control scheme that can simulate the rhythm of natural human reactions is provided. This involves first receiving an external facial expression trigger command, then planning a corresponding sequence of facial expression actions. Specifically, when planning each facial expression action in the sequence, a silent delay is actively inserted before the action is executed. The duration of this silent delay is the time from receiving the command to the start of the action, and its ratio to the execution time of the corresponding action is controlled within a preset range. Finally, the robot's facial actuators are driven to complete the expression based on this sequence. By actively introducing and quantifying the ratio of "stimulus-response" delay to action duration, this overcomes the mechanical and abrupt feeling caused by instantaneous response in existing technologies, simulates the natural cognitive processing rhythm of humans, and makes the robot's facial organ movements more consistent with human movement patterns and reaction rules. This significantly improves the naturalness and lifelikeness of the expressions, making the human-computer interaction experience more intimate and harmonious, and facilitating practical application and promotion.

[0010] In one possible design, the ratio corresponding to the first facial expression action for making eye movements is greater than the ratio corresponding to the second facial expression action for making eyelid movements.

[0011] In one possible design, the ratio is dynamically adjusted as follows: first, the emotion type is determined based on the facial expression trigger command; then, when the emotion type is a fast-response emotion type, the ratio is decreased, and when the emotion type is a slow-response emotion type, the ratio is increased.

[0012] In one possible design, the silence duration is dynamically generated as follows: Get the execution duration corresponding to the facial expression action. Minimum silence duration Maximum silence duration The range of values ​​for the ratio ; Randomly generate a value belonging to the interval Pure decimals The duration of silence was calculated. .

[0013] In one possible design, the minimum silence duration and the maximum silence duration are obtained as follows: Acquire multiple pre-stored real-person facial expression video data corresponding to the facial expression trigger command; For each video data in the multiple sets of real-person facial expression video data, firstly, analyze the corresponding video data to obtain at least one action intensity time sequence data that corresponds one-to-one with the at least one facial expression action. Then, determine the start time of the corresponding facial expression action based on the action intensity time sequence data, and calculate the relative time difference between the start time and the reference start time. Finally, obtain at least one relative time difference that corresponds one-to-one with the at least one facial expression action. For each action in the at least one facial expression action, a minimum silence duration and a maximum silence duration are determined based on multiple relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data. The minimum silence duration is determined based on the minimum value or lower percentile of the multiple relative time differences, and the maximum silence duration is determined based on the maximum value or upper percentile of the multiple relative time differences.

[0014] In one possible design, the preset range of the ratio is obtained as follows: Acquire multiple pre-stored real-person facial expression video data corresponding to the facial expression trigger command; For each video data in the multiple sets of real-person facial expression video data, firstly, analyze the corresponding video data to obtain at least one action intensity time sequence data corresponding to the at least one facial expression action. Then, determine the start time and completion time of the corresponding facial expression action based on the action intensity time sequence data, and calculate the first relative time difference between the start time and the reference start time, and calculate the second relative time difference between the completion time and the start time. Finally, obtain at least one first relative time difference and at least one second relative time difference corresponding to the at least one facial expression action. For each action in the at least one facial expression action, a minimum silence duration and a maximum silence duration are determined based on multiple first relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data. The multiple second relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data are then averaged to obtain a corresponding average execution duration. The result of dividing the minimum silence duration by the average execution duration is then used as the lower limit of the corresponding ratio, and the result of dividing the maximum silence duration by the average execution duration is used as the upper limit of the corresponding ratio. The minimum silence duration is determined based on the minimum or lower percentile of the multiple first relative time differences, and the maximum silence duration is determined based on the maximum or upper percentile of the multiple first relative time differences.

[0015] Secondly, a humanoid robot facial expression control device is provided, including a trigger instruction receiving unit, an action sequence planning unit, and an actuator driving unit that are sequentially connected in communication. The trigger instruction receiving unit is used to receive facial expression trigger instructions from external sources; The action sequence planning unit is used to determine, based on the facial expression triggering instruction, at least one facial expression action that needs to be executed independently and the execution duration of each facial expression action in the at least one facial expression action, and to plan a facial expression action sequence composed of the at least one facial expression action in the following manner: for each facial expression action, a silent delay with a corresponding silent duration is inserted before the corresponding action is executed, wherein the silent duration refers to the duration from the time the facial expression triggering instruction is received to the start time of the execution of the corresponding action, and the ratio of the silent duration to the execution duration of the corresponding action belongs to a preset range; The actuator driving unit is used to drive the humanoid robot and each facial actuator corresponding to each facial expression action to perform the corresponding facial expression action based on the planned facial expression action sequence.

[0016] Thirdly, the present invention provides a computer device comprising a storage module, a processing module, and a transceiver module connected in sequence for communication, wherein the storage module is used to store a computer program, the transceiver module is used to send and receive messages, and the processing module is used to read the computer program and execute the anthropomorphic robot facial expression control method as described in the first aspect or any possible design in the first aspect.

[0017] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the anthropomorphic robot facial expression control method as described in the first aspect or any possible design within the first aspect.

[0018] Fifthly, the present invention provides a computer program product, including a computer program or instructions, wherein the computer program or instructions, when executed by a computer, implement the anthropomorphic robot facial expression control method as described in the first aspect or any possible design in the first aspect.

[0019] The beneficial effects of the above scheme are: (1) This invention creatively provides a robot expression control scheme that can simulate the rhythm of human natural reactions. That is, firstly, it receives an external expression trigger command, and then plans the corresponding facial expression action sequence. Specifically, when planning each facial expression action in the sequence, a silent delay is actively inserted before the action is executed. The duration of the silent delay is the time from receiving the command to the start of the action, and the ratio of the silent delay to the execution time of the corresponding action is controlled within a preset range. Finally, the robot's facial actuator is driven to complete the expression based on the sequence. In this way, by actively introducing and quantifying the ratio of "stimulus-response" delay to action duration, the mechanical and abrupt feeling caused by instantaneous response in the prior art is overcome, and the natural cognitive processing rhythm of humans is simulated. This makes the robot's facial organ movement more in line with human movement patterns and reaction rules, significantly improving the naturalness and vitality of the expression, and making the human-computer interaction experience more intimate and harmonious. (2) It pioneered the “proportional delay” mechanism, laying the temporal foundation for natural interaction. That is, by introducing the control dimension of “the ratio of silence duration to action execution duration”, the optimization of the naturalness of facial expressions is upgraded from traditional hardware or trajectory optimization to bionic simulation of the temporal rhythm of “stimulus-response”. This makes the robot’s facial expression response break away from the rigid “zero delay” mode and have a cognitive logic similar to human “receive-process-response”, fundamentally solving the problem of abrupt and mechanical movements, and injecting the “life” foundation into human-computer interaction. (3) It achieves differentiated and emotional realistic expression. On the one hand, by setting different delay ratios for different organs such as eyeballs and eyelids, it accurately simulates the essential difference in neural reaction speed between human "conscious eyeball movement" and "almost instinctive blinking", making the details of the expression more realistic and credible. On the other hand, by making the delay ratio dynamically change according to the type of emotion (such as surprise, sadness), the robot's expression rhythm has the fit of emotional semantics, realizing the leap from "being able to make expressions" to "being able to make expressions with emotional rhythm", which greatly enhances the context perception ability of the interaction. (4) It can provide a data-driven objective and robust implementation path, that is, by randomly generating specific delay ratios within a preset range, it introduces natural fluctuations in robot expressions that conform to the randomness of human behavior, avoiding the sense of repetition and pattern caused by fixed parameters. Furthermore, by shifting the basis for setting core control parameters (minimum / maximum silence duration, or even the entire ratio range) from subjective experience to statistical analysis of massive amounts of real human expression data, and by using minimum / maximum values ​​or more robust lower / upper percentiles to determine parameter boundaries, it effectively eliminates the interference of abnormal data, making the rhythm imitated by the robot a common and typical human reaction pattern, rather than an individual exception, thereby ensuring the consistency and reliability of the anthropomorphic effect. (5) This solution ultimately achieves a leap from isolated optimization to system biomimicry. That is, by quantifying and controlling the ratio of "delay-execution", it systematically reshapes the generation sequence of robot expressions, making its rhythm, differences and emotions consistent with human expectations. This can significantly alleviate the "uncanny valley effect" caused by bizarre movements, making robot expressions no longer cold mechanical movements, but friendly and natural, and able to convey emotional warmth. This provides key emotional interaction technology support for a human-machine integrated society, which is convenient for practical application and promotion. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the humanoid robot facial expression control method provided in an embodiment of this application.

[0022] Figure 2 This is a schematic diagram of the structure of the humanoid robot facial expression control device provided in the embodiments of this application.

[0023] Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0025] It should be understood that although the terms "first" and "second", etc., may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object, without departing from the scope of the exemplary embodiments of the invention.

[0026] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, or A and B exist simultaneously. Another example is A, B and / or C, which can mean that any one of A, B, and C or any combination thereof exists. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone or A and B exist simultaneously. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.

[0027] Example like Figure 1 As shown, the anthropomorphic robot facial expression control method provided in the first aspect of this embodiment can be executed, but is not limited to, by a computer device with certain computing resources and communicatively connected to the various facial actuators of the anthropomorphic robot. For example, it can be executed by electronic devices such as servers, personal computers (PCs, referring to a type of multi-purpose computer suitable for personal use in terms of size, price, and performance; desktops, laptops, mini-laptops, tablets, and ultrabooks are all considered personal computers), smartphones, personal digital assistants (PDAs), or wearable devices. Figure 1 As shown, the humanoid robot facial expression control method includes, but is not limited to, the following steps S1 to S3.

[0028] S1. Receive facial expression trigger commands from external sources.

[0029] In step S1, the facial expression triggering instruction is used to trigger the execution of at least one facial expression action that can present the target expression (for example, in order to make a "surprised" expression, at least a first facial expression action for eye movement and a second facial expression action for eyelid movement need to be executed, etc.). It may be, but is not limited to, from an external network or control terminal, and specifically includes, but is not limited to, a unique identifier of the at least one facial expression action and the execution duration, minimum silence duration, maximum silence duration and / or the ratio range of silence duration to execution duration of each facial expression action in the at least one facial expression action.

[0030] S2. Based on the facial expression triggering instruction, determine at least one facial expression action that needs to be executed independently and the execution duration of each facial expression action in the at least one facial expression action, and plan a facial expression action sequence composed of the at least one facial expression action as follows: For each facial expression action, insert a silent delay with a corresponding silent duration before executing the corresponding action, wherein the silent duration refers to the duration from the time of receiving the facial expression triggering instruction to the start time of the execution of the corresponding action, and the ratio of the silent duration to the execution duration of the corresponding action belongs to a preset range.

[0031] In step S2, the execution duration refers to the time from the start of the corresponding action to its completion. For example, the execution duration of a blinking action (which is one of the facial expression actions) is 100ms to 400ms. The silence duration is used to simulate the biological characteristics of delayed activation of human muscle groups corresponding to the facial actuators under the drive of neural signals. For example, when the facial expression action is a blinking action, the corresponding silence duration is 50ms to 200ms (i.e., the minimum silence duration is 50ms and the maximum silence duration is 200ms). The ratio is used to overcome the mechanical and abrupt feeling caused by the immediate response of robot expressions in the prior art. That is, by associating the execution duration with the silence duration, a human-like temporal logic of 'stimulus-processing-response' can be given to the robot's facial expressions, making its actions present a sense of life with 'intention', significantly improving the naturalness and affinity of human-computer interaction, thereby effectively alleviating the 'uncanny valley effect' and achieving the purpose of simulating the natural reaction rhythm of humans.

[0032] In step S2, the preset range of the ratio can be exemplified as [0.11, 0.6], meaning the preset ratio of the silence delay time (which is the ratio between the time interval from the issuance of the command to the robot's facial expression and the time interval from the issuance of the command to the robot's action of completing the command) ranges from approximately 10% to 40%. In this embodiment, different facial expressions can correspond to different ratios (i.e., the corresponding preset ranges will also be different); for example, considering that most eye movements (such as tracking objects or shifting gaze) are "voluntary movements" controlled by consciousness, involving complex processing in the cerebral cortex, there is a significant cognitive delay between "deciding where to look" and "the eyes begin to move," making it possible to perfectly simulate this higher-level neural activity of "thinking before acting" by using a longer silence duration (i.e., a larger ratio); at the same time, blinking is mostly a rapid and spontaneous physiological activity (such as moisturizing and protection), or even a lower... Nonverbal communication of consciousness (such as eye contact) has shorter neural pathways and is closer to a reflex arc, allowing for the precise simulation of this near-instinctive and rapid reaction characteristic using a shorter period of silence (i.e., a smaller ratio). Furthermore, by differentiating the ratio, the robot is no longer uniformly "delayed in motion," but can exhibit the essential difference between "thinking eye movements" and "instinctive blinking," making the movement of the anthropomorphic robot in this embodiment more human-like. Specifically, the ratio corresponding to the first facial expression action used to make eye movements is greater than the ratio corresponding to the second facial expression action used to make eyelid movements.

[0033] In step S2, the ratio can also be configured as a parameter that can be dynamically adjusted based on emotional commands, so that the anthropomorphic robot's facial expressions can conform to emotional semantics (such as a quick reaction when surprised or a slow reaction when sad), achieving a higher level of context-aware interaction. Specifically, the ratio is preferably dynamically adjusted as follows: first, the emotional type is determined based on the facial expression trigger command; then, when the emotional type is a quick-response emotional type, the ratio is decreased, and when the emotional type is a slow-response emotional type, the ratio is increased. The quick-response emotional type refers to an expression type that can react quickly, such as a "surprised" expression; the slow-response emotional type refers to an expression type that reacts slowly, such as a "sad" or "thinking" expression.

[0034] In step S2, the silence duration can also be randomly configured based on the execution duration, minimum silence duration, maximum silence duration, and the range of the ratio, so as to introduce random micro-movements to simulate unconscious physiological activities. Preferably, the silence duration is dynamically generated as follows: first, the execution duration corresponding to the corresponding facial expression action is obtained. Minimum silence duration Maximum silence duration The range of values ​​for the ratio Then randomly generate a value belonging to the interval. Pure decimals The duration of silence was calculated. The minimum and maximum silence durations reflect the upper and lower limits of the biological characteristics of the delayed activation of the human muscle groups corresponding to the facial actuators under the drive of neural signals. In order to achieve accurate measurement of these upper and lower limits of biological characteristics, they can be obtained in advance based on real facial expression video data. That is, preferably, the minimum and maximum silence durations are obtained according to the following steps S211 to S213.

[0035] S211. Obtain multiple sets of real-person facial expression video data that are pre-stored and correspond to the facial expression triggering command.

[0036] In step S211, the multiple sets of real facial expression video data can be routinely extracted from a video database used to bind and store target expressions and real facial expression video data.

[0037] S212. For each video data in the plurality of real-person facial expression video data, firstly, analyze the corresponding video data to obtain at least one action intensity time sequence data corresponding to the at least one facial expression action, then determine the start time of the corresponding facial expression action based on the action intensity time sequence data, and calculate the relative time difference between the start time and the reference start time, and finally obtain at least one relative time difference corresponding to the at least one facial expression action.

[0038] In step S212, the real facial expression video data will contain multiple consecutive frames of real facial expression video images (e.g., 60 frames per second). This allows for frame-by-frame visual analysis processing, such as motion intensity recognition (e.g., first performing motion recognition based on existing algorithms, then performing conventional motion intensity measurement on the motion recognition results to obtain the motion intensity recognition result), resulting in the following set of values: left eye blink intensity: 0.1; right eye blink intensity: 0.1; left eyebrow drooping intensity: 0.3; right eyebrow drooping intensity: 0.2; mouth opening intensity: 0.4. Intensity of left corner of mouth lift: 0.5; etc. (The aforementioned left eye blink, right eye blink, left eyebrow droop, right eyebrow droop, mouth opening, and left corner of mouth lift intensity are the at least one facial expression action; the aforementioned values ​​represent the intensity of the corresponding specific facial muscle action, such as the degree of left eye blink and the degree of mouth lift, etc. These values ​​vary between 0 and 1, where 0 represents complete relaxation and 1 represents maximum contraction). Finally, for each action in the at least one facial expression action, all corresponding intensity values ​​are summarized in chronological order to obtain the corresponding action intensity time sequence data.

[0039] In step S212, the starting time refers to when the corresponding action begins. For example, it is necessary to determine when the action of opening the mouth changes from a closed state (i.e., the intensity value is close to 0) to an open state (i.e., the intensity value increases). This can be determined as follows: Set an intensity threshold value for the mouth opening action, such as 0.2; then continuously monitor the intensity time series data of the mouth opening action. Once it is found that it changes from below 0.2 to above 0.2 and continues to rise, it is considered that the action has "started", and this time point (e.g., 0.5 seconds from the start of the imitation) is recorded as the starting time corresponding to the mouth opening action. The reference start time can be specifically the earliest start time. For example, if the left eye blinking starts at 0.3 seconds, the mouth opening starts at 0.5 seconds, and the eyebrow drooping starts at 0.4 seconds, then the reference start time can be determined to be 0.3 seconds. Thus, the following set of relative time differences can be obtained: left eye blinking action: 0.0 seconds; mouth opening action: 0.2 seconds; eyebrow drooping action: 0.1 seconds; and so on.

[0040] S213. For each action in the at least one facial expression action, determine the corresponding minimum silence duration and maximum silence duration based on multiple relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data, wherein the minimum silence duration is determined based on the minimum value or lower percentile of the multiple relative time differences, and the maximum silence duration is determined based on the maximum value or upper percentile of the multiple relative time differences.

[0041] In step S213, the lower percentile refers to the value at a lower percentage after all data are arranged in ascending order (commonly represented by the 5th percentile, P5), which serves as the basis for determining the minimum silence duration, thereby eliminating unreasonably short delays (which may be due to measurement errors or accidental actions) and finding a reasonable and relatively small lower limit for delay. The upper percentile refers to the value at a higher percentage after all data are arranged in ascending order (commonly represented by the 95th percentile, P95), which serves as the basis for determining the maximum silence duration, thereby eliminating unreasonably long delays (which may be abnormal cases of excessively slow response) and finding a reasonable and relatively large upper limit for delay. Furthermore, when determining the minimum and maximum silence durations, the minimum value or the lower percentile can be scaled proportionally (possibly faster or slower) to determine the minimum silence duration, and the maximum value or the upper percentile can be scaled proportionally (possibly faster or slower) to determine the maximum silence duration.

[0042] In step S2, based on the principle of steps S211 to S213 above, the preset range of the ratio can also be accurately determined. Preferably, the preset range of the ratio is obtained according to the following steps S221 to S223.

[0043] S221. Obtain multiple sets of real-person facial expression video data that are pre-stored and correspond to the facial expression triggering command.

[0044] S222. For each video data in the plurality of real-person facial expression video data, firstly, analyze the corresponding video data to obtain at least one action intensity time sequence data corresponding to the at least one facial expression action, then determine the start time and completion time of the corresponding facial expression action based on the action intensity time sequence data, and calculate the first relative time difference between the start time and the reference start time, and calculate the second relative time difference between the completion time and the start time, and finally obtain at least one first relative time difference and at least one second relative time difference corresponding to the at least one facial expression action.

[0045] S223. For each action in the at least one facial expression action, based on multiple first relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data, determine the corresponding minimum silence duration and maximum silence duration, and average the multiple second relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data to obtain the corresponding average execution duration. Then, the result of dividing the minimum silence duration by the average execution duration is taken as the lower limit of the corresponding ratio, and the result of dividing the maximum silence duration by the average execution duration is taken as the upper limit of the corresponding ratio. The minimum silence duration is determined based on the minimum value or lower percentile of the multiple first relative time differences, and the maximum silence duration is determined based on the maximum value or upper percentile of the multiple first relative time differences.

[0046] The specific details of steps S221 to S223 above can be derived from steps S211 to S213 above, and will not be repeated here. In addition, when acquiring the multiple sets of real-person facial expression video data, the emotion type can be determined first according to the facial expression triggering command, and then the emotion type can be used to filter the multiple sets of real-person facial expression video data so that the upper and lower limits of the final ratio can fit the emotional semantics.

[0047] S3. Based on the planned facial expression action sequence, drive the humanoid robot and each facial actuator corresponding to each facial expression action to perform the corresponding facial expression action.

[0048] In step S3, a series of actions consisting of the corresponding facial expression movements performed by each facial actuator can present the target expression.

[0049] Therefore, based on the anthropomorphic robot facial expression control method described in steps S1 to S3 above, a robot facial expression control scheme that can simulate the rhythm of natural human reactions is provided. This involves first receiving an external facial expression trigger command, then planning a corresponding sequence of facial expression actions. Specifically, when planning each facial expression action in the sequence, a silent delay is actively inserted before the action is executed. The duration of this silent delay is the time from receiving the command to the start of the action, and its ratio to the execution time of the corresponding action is controlled within a preset range. Finally, the robot's facial actuator is driven to complete the expression based on this sequence. By actively introducing and quantifying the ratio of the "stimulus-response" delay to the action duration, the mechanical and abrupt feeling caused by instantaneous response in existing technologies is overcome. This simulates the natural cognitive processing rhythm of humans, making the robot's facial organ movements more consistent with human movement patterns and reaction rules, significantly improving the naturalness and lifelikeness of expressions, making the human-computer interaction experience more intimate and harmonious, and facilitating practical application and promotion.

[0050] like Figure 2 As shown, the second aspect of this embodiment provides a virtual device for implementing the humanoid robot facial expression control method described in the first aspect, including a trigger instruction receiving unit, an action sequence planning unit, and an actuator driving unit that are sequentially connected in communication. The trigger instruction receiving unit is used to receive facial expression trigger instructions from external sources; The action sequence planning unit is used to determine, based on the facial expression triggering instruction, at least one facial expression action that needs to be executed independently and the execution duration of each facial expression action in the at least one facial expression action, and to plan a facial expression action sequence composed of the at least one facial expression action in the following manner: for each facial expression action, a silent delay with a corresponding silent duration is inserted before the corresponding action is executed, wherein the silent duration refers to the duration from the time the facial expression triggering instruction is received to the start time of the execution of the corresponding action, and the ratio of the silent duration to the execution duration of the corresponding action belongs to a preset range; The actuator driving unit is used to drive the humanoid robot and each facial actuator corresponding to each facial expression action to perform the corresponding facial expression action based on the planned facial expression action sequence.

[0051] The working process, working details and technical effects of the aforementioned device provided in the second aspect of this embodiment can be found in the anthropomorphic robot facial expression control method described in the first aspect, and will not be repeated here.

[0052] like Figure 3 As shown, the third aspect of this embodiment provides a computer device for executing the humanoid robot facial expression control method as described in the first aspect. The device includes a storage module, a processing module, and a transceiver module connected in sequence. The storage module stores a computer program, the transceiver module sends and receives messages, and the processing module reads the computer program and executes the humanoid robot facial expression control method as described in the first aspect. Specifically, the storage module may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; the processing module may, but is not limited to, use a microprocessor of the STM32F105 series. Furthermore, the computer device may also include, but is not limited to, a power supply module, a display screen, and other necessary components.

[0053] The working process, working details and technical effects of the aforementioned computer device provided in the third aspect of this embodiment can be found in the anthropomorphic robot facial expression control method described in the first aspect, and will not be repeated here.

[0054] This fourth aspect of the embodiment provides a computer-readable storage medium storing instructions comprising the humanoid robot facial expression control method as described in the first aspect. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, perform the humanoid robot facial expression control method as described in the first aspect. The computer-readable storage medium refers to a data storage medium, and may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0055] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment can be found in the anthropomorphic robot facial expression control method described in the first aspect, and will not be repeated here.

[0056] This fifth aspect of the embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, implements the anthropomorphic robot facial expression control method as described in the first aspect. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0057] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for controlling facial expressions in a humanoid robot, characterized in that, include: Receive facial expression trigger commands from external sources; Based on the facial expression triggering command, at least one facial expression action that needs to be executed independently and the execution duration of each facial expression action in the at least one facial expression action are determined, and a facial expression action sequence composed of the at least one facial expression action is planned in the following manner: For each facial expression action, a silent delay with a corresponding silent duration is inserted before the corresponding action is executed, wherein the silent duration refers to the duration from the time of receiving the facial expression triggering command to the start time of the execution of the corresponding action, and the ratio of the silent duration to the execution duration of the corresponding action belongs to a preset range; Based on the planned sequence of facial expression movements, the anthropomorphic robot is driven to perform corresponding facial expression movements through its facial actuators, which are one-to-one with each of the facial expression movements.

2. The humanoid robot facial expression control method according to claim 1, characterized in that, The ratio corresponding to the first facial expression action used to make eye movements is greater than the ratio corresponding to the second facial expression action used to make eyelid movements.

3. The humanoid robot facial expression control method according to claim 1, characterized in that, The ratio is dynamically adjusted as follows: first, the emotion type is determined based on the facial expression trigger command; then, when the emotion type is a fast-response emotion type, the ratio is decreased, and when the emotion type is a slow-response emotion type, the ratio is increased.

4. The humanoid robot facial expression control method according to claim 1, characterized in that, The silence duration is dynamically generated as follows: Get the execution duration corresponding to the facial expression action. Minimum silence duration Maximum silence duration The range of values ​​for the ratio ; Randomly generate a value belonging to the interval Pure decimals The duration of silence was calculated. .

5. The humanoid robot facial expression control method according to claim 4, characterized in that, The minimum silence duration and the maximum silence duration are obtained as follows: Acquire multiple pre-stored real-person facial expression video data corresponding to the facial expression trigger command; For each video data in the multiple sets of real-person facial expression video data, firstly, analyze the corresponding video data to obtain at least one action intensity time sequence data that corresponds one-to-one with the at least one facial expression action. Then, determine the start time of the corresponding facial expression action based on the action intensity time sequence data, and calculate the relative time difference between the start time and the reference start time. Finally, obtain at least one relative time difference that corresponds one-to-one with the at least one facial expression action. For each action in the at least one facial expression action, a minimum silence duration and a maximum silence duration are determined based on multiple relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data. The minimum silence duration is determined based on the minimum value or lower percentile of the multiple relative time differences, and the maximum silence duration is determined based on the maximum value or upper percentile of the multiple relative time differences.

6. The humanoid robot facial expression control method according to claim 1, characterized in that, The preset range of the ratio is obtained in the following manner: Acquire multiple pre-stored real-person facial expression video data corresponding to the facial expression trigger command; For each video data in the multiple sets of real-person facial expression video data, firstly, analyze the corresponding video data to obtain at least one action intensity time sequence data corresponding to the at least one facial expression action. Then, determine the start time and completion time of the corresponding facial expression action based on the action intensity time sequence data, and calculate the first relative time difference between the start time and the reference start time, and calculate the second relative time difference between the completion time and the start time. Finally, obtain at least one first relative time difference and at least one second relative time difference corresponding to the at least one facial expression action. For each action in the at least one facial expression action, a minimum silence duration and a maximum silence duration are determined based on multiple first relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data. The multiple second relative time differences that correspond one-to-one with the multiple sets of real-person facial expression video data are then averaged to obtain a corresponding average execution duration. The result of dividing the minimum silence duration by the average execution duration is then used as the lower limit of the corresponding ratio, and the result of dividing the maximum silence duration by the average execution duration is used as the upper limit of the corresponding ratio. The minimum silence duration is determined based on the minimum or lower percentile of the multiple first relative time differences, and the maximum silence duration is determined based on the maximum or upper percentile of the multiple first relative time differences.

7. A humanoid robot facial expression control device, characterized in that, It includes a trigger instruction receiving unit, an action sequence planning unit, and an actuator driving unit that are connected in sequence. The trigger instruction receiving unit is used to receive facial expression trigger instructions from external sources; The action sequence planning unit is used to determine, based on the facial expression triggering instruction, at least one facial expression action that needs to be executed independently and the execution duration of each facial expression action in the at least one facial expression action, and to plan a facial expression action sequence composed of the at least one facial expression action in the following manner: for each facial expression action, a silent delay with a corresponding silent duration is inserted before the corresponding action is executed, wherein the silent duration refers to the duration from the time the facial expression triggering instruction is received to the start time of the execution of the corresponding action, and the ratio of the silent duration to the execution duration of the corresponding action belongs to a preset range; The actuator driving unit is used to drive the humanoid robot and each facial actuator corresponding to each facial expression action to perform the corresponding facial expression action based on the planned facial expression action sequence.

8. A computer device, characterized in that, It includes a storage module, a processing module, and a transceiver module that are sequentially connected in communication. The storage module is used to store computer programs, the transceiver module is used to send and receive messages, and the processing module is used to read the computer programs and execute the humanoid robot facial expression control method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that... The computer-readable storage medium stores instructions that, when executed on a computer, perform the humanoid robot facial expression control method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or the instructions are executed by the computer, they implement the humanoid robot facial expression control method as described in any one of claims 1 to 6.