Bionic robot facial expression synchronization method and device based on visual perception, equipment and medium

By combining visual perception and asynchronous thread control with hardware compensation parameters and a lightweight MLP model, the problems of lightweight and intelligent facial expression synchronization in bionic robots have been solved, achieving efficient facial expression synchronization.

CN121893229BActive Publication Date: 2026-05-29DIGITAL HUAXIA (SHENZHEN) TECHNOLOGY CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DIGITAL HUAXIA (SHENZHEN) TECHNOLOGY CO LTD
Filing Date
2026-03-24
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for synchronizing facial expressions in bionic robots rely on complex 3D simulation engines, which increases R&D costs and makes it difficult to achieve lightweight deployment. Furthermore, limitations in processing precision and assembly technology lead to differences in facial expressions, resulting in a lack of flexibility and intelligence.

Method used

By using a vision-based approach, the robot captures visual facial expressions using a pre-set camera, constructs a hardware compensation parameter vector, and combines asynchronous perception and driving threads with a lightweight MLP model and a visual closed-loop feedback mechanism to achieve end-to-end facial expression synchronization.

Benefits of technology

It has achieved lightweight bionic robot facial expression synchronization, improved the flexibility and intelligence of expression synchronization, reduced R&D costs, and solved the problem of individual differences caused by production and assembly errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121893229B_ABST
    Figure CN121893229B_ABST
Patent Text Reader

Abstract

The application discloses a visual perception-based bionic robot facial expression synchronization method and device, equipment and medium, and relates to the technical field of human-computer interaction. The method comprises the following steps: controlling a target bionic robot to execute a target expression control instruction for realizing a preset expected robot facial expression, and capturing a robot visual facial expression of the robot; calibrating and setting the target bionic robot based on the robot visual facial expression and the preset expected robot facial expression, so as to construct a hardware compensation parameter vector; executing a facial expression synchronization operation based on the hardware compensation parameter vector and a pre-trained target expression mapping model by asynchronously running a preset perception thread and a preset driving thread; the preset perception thread is used for collecting a target human face image in real time, and outputting a target steering wheel control vector through the target expression mapping model; and the preset driving thread is used for outputting a steering wheel control instruction at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector and the target steering wheel control vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and in particular to a method, apparatus, device, and medium for synchronizing facial expressions of bionic robots based on visual perception. Background Technology

[0002] As a crucial medium for human-computer interaction, bionic robots need to possess facial expressions similar to humans to achieve more natural and effective communication. However, current bionic robots still face numerous challenges in achieving facial expressions. Traditional bionic face synchronization typically relies on complex 3D simulation engines such as Blender (a 3D graphics software) as intermediaries. This requires creating a scaled-down 3D model for each robot and developing complex communication links, increasing R&D costs and making it difficult to meet the demands of lightweight device deployment. Furthermore, due to limitations in manufacturing precision and assembly processes, unavoidable physical errors occur during the production of bionic robots. This results in significant differences in facial expressions across different hardware entities using the same set of control parameters, making it difficult to achieve consistent facial movements. Simultaneously, existing expression control methods primarily rely on pre-programmed mechanical movements and limited expression templates, lacking flexibility and intelligence.

[0003] In summary, how to achieve a lightweight method for synchronizing facial expressions in bionic robots to improve the flexibility and intelligence of expression synchronization is a problem that urgently needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for synchronizing facial expressions in bionic robots based on visual perception, which can realize a lightweight method for synchronizing facial expressions in bionic robots to improve the flexibility and intelligence of expression synchronization. The specific solution is as follows:

[0005] In a first aspect, this application provides a method for synchronizing facial expressions in a bionic robot based on visual perception, comprising:

[0006] The target bionic robot is controlled to execute target expression control commands to achieve a preset desired robot facial expression, and during the execution process, the robot visual facial expression of the target bionic robot is captured by a preset camera;

[0007] The target bionic robot is calibrated and standardized based on the robot's visual facial expressions and the preset expected robot facial expressions in order to construct the hardware compensation parameter vector of the target bionic robot.

[0008] By asynchronously running a preset perception thread and a preset driving thread, the facial expression synchronization operation of the target bionic robot is performed based on the hardware compensation parameter vector and the pre-trained target expression mapping model. The preset perception thread is used to acquire the target face image in real time, and after extracting the target expression feature vector of the target face image, it outputs the corresponding target servo control vector through the target expression mapping model. The preset driving thread is used to output servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector.

[0009] The target expression mapping model is a model obtained by training a pre-generated training dataset; the training dataset consists of several expression feature vectors and corresponding servo control vectors; the expression feature vectors are obtained based on target expression parameters, which are preset expression parameters that conform to the anatomical logic of the human face.

[0010] Optionally, the process of generating the training dataset includes:

[0011] The target facial expression parameters are smoothed by an interpolation algorithm and then superimposed with fractal Brownian motion noise to obtain the processed target facial expression parameters.

[0012] Under preset anatomical constraints, the joint probability distribution of the processed target facial expression parameters is sampled to obtain the corresponding facial expression feature vector, and the facial expression feature vector is optimized to obtain the corresponding optimized facial expression feature vector.

[0013] The optimized facial expression feature vector is input into the target simulation tool to extract the corresponding servo control vector, thereby obtaining the training dataset composed of the servo control vector and the optimized facial expression feature vector.

[0014] Optionally, the step of sampling the joint probability distribution of the processed target expression parameters under preset anatomical constraints to obtain the corresponding expression feature vector, and optimizing the expression feature vector to obtain the corresponding optimized expression feature vector, includes:

[0015] Based on a preset anatomical correlation matrix, a joint probability distribution of each dimension variable in the processed target expression parameters is generated; the preset anatomical correlation matrix is ​​used to define the physical coupling ratio between the target expression parameters.

[0016] The joint probability distribution is sampled to obtain the corresponding facial expression feature vector;

[0017] Determine the target feature value in the facial expression feature vector; the target feature value is a feature value that exceeds a preset feature value range;

[0018] The target feature value is adjusted by a preset logarithmic compression function to control the target feature value within the preset feature value range, thereby obtaining the adjusted facial expression feature vector;

[0019] Gaussian white noise is superimposed on the adjusted facial expression feature vector, and then processed using a low-pass filter to obtain the optimized facial expression feature vector.

[0020] Optionally, the control target bionic robot executes target expression control commands to achieve a preset desired robot facial expression, and during execution, captures the robot visual facial expression of the target bionic robot through a preset camera, including:

[0021] The target bionic robot is controlled to execute a first target expression control command to achieve a first preset desired robot facial expression; the first preset desired robot facial expression corresponds to a human face expression.

[0022] During the execution of the first target facial expression control command by the target bionic robot, the first robot visual facial expression of the target bionic robot is captured by a preset camera.

[0023] Optionally, the calibration and standardization of the target bionic robot based on the robot's visual facial expression and the preset desired robot facial expression, to construct the hardware compensation parameter vector of the target bionic robot, includes:

[0024] Calculate the corresponding facial feature residual matrix based on the first robot visual facial expression and the first preset expected robot facial expression;

[0025] If there is a first target dimension in the facial feature residual matrix with a residual value not less than a preset residual threshold, then the hardware offset compensation parameter vector corresponding to the first target dimension is adjusted based on the preset regularization technique until the preset convergence condition is met.

[0026] The preset convergence conditions include that the residual value of the first target dimension is less than the preset residual threshold, and that after the hardware offset compensation parameter vector has been adjusted a preset number of times, the degree of change of the first robot visual facial expression is lower than the preset degree of change threshold.

[0027] Optionally, the control target bionic robot executes target expression control commands to achieve a preset desired robot facial expression, and during execution, captures the robot visual facial expression of the target bionic robot through a preset camera, including:

[0028] The target bionic robot is controlled to execute a second target facial expression control command to achieve a second preset desired robot facial expression; the second preset desired robot facial expression corresponds to a preset number of individual facial expressions.

[0029] During the process of the target bionic robot executing the second target expression control command, the second robot visual facial expression of the target bionic robot is captured by a preset camera;

[0030] Accordingly, the calibration and standardization of the target bionic robot based on the robot's visual facial expression and the preset desired robot facial expression, to construct the hardware compensation parameter vector of the target bionic robot, includes:

[0031] Based on the second robot's visual facial expression, the first facial expression amplitude of the target bionic robot is determined, and the second facial expression amplitude of the second preset desired robot facial expression is determined.

[0032] Based on the first facial expression amplitude and the second facial expression amplitude, determine the corresponding facial expression amplitude residual matrix;

[0033] If there is a second target dimension in the facial expression amplitude residual matrix with a residual value not less than a preset residual threshold, then the hardware gain compensation parameter vector corresponding to the second target dimension is adjusted based on the preset regularization technique until the preset convergence condition is met.

[0034] The preset convergence conditions include that the residual value of the second target dimension is less than the preset residual threshold, and that after the hardware gain compensation parameter vector has been adjusted a preset number of times, the degree of change of the second robot's visual facial expression is lower than the preset degree of change threshold.

[0035] Optionally, the process of outputting servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector includes:

[0036] Within the time interval between two consecutive outputs of the target servo control vector by the preset sensing thread, the current servo state of the target bionic robot and the latest target servo control vector output by the preset sensing thread are determined based on the preset output frequency.

[0037] Based on the current servo state and the latest target servo control vector, a transition frame-filling servo control vector is generated using a linear interpolation algorithm;

[0038] The frame-complementing servo control vector is corrected using the hardware compensation parameter vector, and corresponding frame-complementing servo control commands are generated and output based on the correction results.

[0039] Secondly, this application provides a vision-perception-based bionic robot facial expression synchronization device, comprising:

[0040] The data capture module is used to control the target bionic robot to execute target expression control commands to achieve a preset desired robot facial expression, and to capture the robot visual facial expression of the target bionic robot through a preset camera during the execution process;

[0041] The calibration module is used to calibrate the target bionic robot based on the robot's visual facial expression and the preset expected robot facial expression, so as to construct the hardware compensation parameter vector of the target bionic robot.

[0042] The synchronization implementation module is used to perform facial expression synchronization operations of the target bionic robot based on the hardware compensation parameter vector and the pre-trained target expression mapping model by asynchronously running a preset perception thread and a preset drive thread. The preset perception thread is used to acquire target face images in real time, and after extracting the target expression feature vector of the target face image, output the corresponding target servo control vector through the target expression mapping model. The preset drive thread is used to output servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector.

[0043] The target expression mapping model is a model obtained by training a pre-generated training dataset; the training dataset consists of several expression feature vectors and corresponding servo control vectors; the expression feature vectors are obtained based on target expression parameters, which are preset expression parameters that conform to the anatomical logic of the human face.

[0044] Thirdly, this application provides an electronic device, comprising:

[0045] Memory, used to store computer programs;

[0046] A processor is used to execute the computer program to implement the aforementioned method for synchronizing facial expressions of a bionic robot based on visual perception.

[0047] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned method for synchronizing facial expressions of a bionic robot based on visual perception.

[0048] In this application, a target bionic robot is controlled to execute target facial expression control commands to achieve a preset desired robot facial expression. During execution, a preset camera captures the robot's visual facial expression. Based on the robot's visual facial expression and the preset desired robot facial expression, the target bionic robot is calibrated to construct a hardware compensation parameter vector. A preset perception thread and a preset driving thread are run asynchronously to perform facial expression synchronization operations on the target bionic robot based on the hardware compensation parameter vector and a pre-trained target expression mapping model. The preset perception thread is used to collect target facial expression data in real time. A facial image is processed, and after extracting the target facial expression feature vector, the corresponding target servo control vector is output through the target facial expression mapping model. The preset drive thread is used to output servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector. The target facial expression mapping model is a model trained on a pre-generated training dataset. The training dataset consists of several facial expression feature vectors and corresponding servo control vectors. The facial expression feature vectors are obtained based on target facial expression parameters, which are preset facial expression parameters that conform to the anatomical logic of the human face. As can be seen from the above, this application controls the target bionic robot to execute target expression control commands containing preset desired robot facial expressions. Simultaneously, during command execution, a preset camera captures the robot's visual facial expressions. Based on these visual facial expressions and the preset desired robot facial expressions, the target bionic robot is calibrated to construct a hardware compensation parameter vector adapted to the robot. Subsequently, by asynchronously running a preset perception thread and a preset drive thread, combined with the hardware compensation parameter vector and a target expression mapping model pre-trained based on a training dataset consisting of several expression feature vectors and their corresponding servo control vectors, the target bionic robot's facial expression synchronization operation is executed. The preset perception thread is responsible for real-time acquisition of target face images, extracting target expression feature vectors from the images, and outputting the corresponding target servo control vectors via the target expression mapping model. The preset drive thread, based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vectors, outputs servo control commands at a preset output frequency. Furthermore, the expression feature vectors in the training dataset are obtained based on preset target expression parameters that conform to human facial anatomy.In this way, through the process described in this application, the conversion from facial feature vectors to servo control vectors is directly achieved using the target facial expression mapping model, eliminating the need for a complex 3D simulation engine as an intermediary. This simplifies the system structure, reduces R&D costs, and enables the deployment of lightweight equipment. The dual-thread asynchronous control architecture, combined with a frame interpolation algorithm, effectively solves the problems of mechanical vibration and motion jumps caused by the mismatch between visual perception and drive frequency, ensuring that the smoothness of the hardware execution end is not affected by fluctuations at the perception end. Through a parameter correction mechanism, control parameters can be dynamically adjusted according to the actual performance of different robots, effectively solving the problem of individual differences caused by production assembly errors, reducing manual debugging costs, and thus realizing a lightweight bionic robot facial expression synchronization method to improve the flexibility and intelligence of expression synchronization. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0050] Figure 1 This is a flowchart of a method for synchronizing facial expressions of a bionic robot based on visual perception, as disclosed in this application.

[0051] Figure 2 This is a system logic block diagram of a bionic robot facial expression synchronization method disclosed in this application;

[0052] Figure 3 This is a schematic diagram of the head structure of a bionic robot disclosed in this application;

[0053] Figure 4 This is a schematic diagram illustrating the generation process of a training dataset as disclosed in this application;

[0054] Figure 5 This is a schematic diagram of the operation of a dual-threaded asynchronous architecture disclosed in this application;

[0055] Figure 6 This is a schematic diagram of the structure of a vision-perception-based bionic robot facial expression synchronization device disclosed in this application.

[0056] Figure 7 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] Current bionic robots still face numerous challenges in achieving facial expressions. Traditional bionic face synchronization typically relies on complex 3D simulation engines like Blender as intermediaries, requiring the creation of scaled 3D models for each robot and the development of complex communication links, increasing R&D costs and making it difficult to meet the demands of lightweight device deployment. Furthermore, due to limitations in manufacturing precision and assembly processes, unavoidable physical errors occur during the production of bionic robots, leading to significant differences in facial expressions across different hardware entities using the same set of control parameters, making it difficult to achieve consistent facial movements. Simultaneously, existing expression control methods primarily rely on pre-programmed mechanical movements and limited expression templates, lacking flexibility and intelligence.

[0059] To overcome the aforementioned technical problems, this application provides a visual perception-based bionic robot facial expression synchronization method, which can realize a lightweight bionic robot facial expression synchronization method to improve the flexibility and intelligence of expression synchronization.

[0060] See Figure 1 As shown, this embodiment of the invention discloses a method for synchronizing facial expressions of a bionic robot based on visual perception, including:

[0061] Step S11: Control the target bionic robot to execute target expression control commands to achieve the preset desired robot facial expression, and capture the robot visual facial expression of the target bionic robot through a preset camera during the execution process.

[0062] In this embodiment, due to unavoidable assembly errors during the production of the bionic robot head, such as servo motor mounting angle deviation, mechanical link length error, and inconsistent silicone surface resistance, traditional fixed parameters... The facial expression feature vectors actually exhibited by the robot when applied to different entities With the target expectation vector There are discrepancies, therefore, adaptive calibration based on visual feedback is required for the target bionic robot. First, the target bionic robot is controlled to execute target facial expression control commands containing a preset desired robot facial expression. Simultaneously, during command execution, a preset camera captures the robot's visual facial expressions to obtain... .

[0063] It should be noted that, such as Figure 2 The diagram shows a system logic block diagram of a biomimetic robot facial expression synchronization method provided in this application. Specifically, it includes a camera module, an MLP (Multi-Layer Perceptron) mapping model module, a linear compensation layer, and a dual-thread asynchronous drive module. The MLP model is used to directly map ARKit (an augmented reality development kit for facial capture and expression recognition) facial feature vectors to servo control vectors. By replacing the traditional 3D engine with a lightweight MLP deep learning model, an end-to-end mapping from the ARKit facial feature space to hardware servo control vectors is constructed, and a visual closed-loop feedback mechanism is combined to achieve automated calibration of hardware parameters. In the process of generating the training dataset for the MLP model, multidimensional action coefficients are jointly sampled using the anatomical correlation matrix M to ensure the integrity of the feature space coverage under anatomical constraints. A nonlinear viscoelastic reflection algorithm is used to process sampled values ​​exceeding the defined [0,1] range, and a damping coefficient controls the continuity of data distribution at the boundaries. A drive execution thread independent of the visual perception thread is established, and by maintaining buffer frames and performing temporal interpolation calculations, non-equidistant perception signals are transformed into equidistant, smooth servo control sequences. Facial features of both humans and robots are simultaneously acquired via a camera, and the hardware compensation parameter vector in the linear compensation layer is automatically corrected based on the difference between the two in the ARKit feature space. Figure 3 The diagram shows a schematic of the head structure of a bionic robot provided in this application. It uses 27 servo motors to simulate the tension of 27 key muscles in its silicone facial skin, such as eyebrows, nose, mouth, and chin. Simultaneously, servo motors inside the eyeballs control eyeball movement. Correspondingly, the hardware compensation parameter vector is a 27-dimensional vector data used to represent the compensation coefficients of the 27 servo motors, where each element is a float number.

[0064] It should be noted that the automatic calibration process of the target bionic robot includes two stages: static reference calibration (calibration) and dynamic travel measurement. Therefore, the target expression control command includes a first target expression control command and a second target expression control command, corresponding to the static reference calibration stage and the dynamic travel measurement stage, respectively. Specifically, when the target expression control command is the first target expression control command, the processing flow for obtaining the robot's visual facial expression is as follows: the target bionic robot is controlled to execute the first target expression control command to achieve a first preset desired robot facial expression; the first preset desired robot facial expression corresponds to a human face expression; during the execution of the first target expression control command by the target bionic robot, the first robot visual facial expression of the target bionic robot is captured by a preset camera. That is, in the static reference calibration stage, the servo control vector corresponding to the ARKi vector of the default expression is issued to control the target bionic robot to execute the first target expression control command corresponding to a single human face expression (i.e., the aforementioned default expression) and containing the first preset desired robot facial expression. During the robot's execution of this command, N frames of robot facial data are continuously captured by a preset camera to obtain its first robot visual facial expression. When the target expression control instruction is the second target expression control instruction, the processing flow for acquiring the robot's visual facial expression is as follows: The target bionic robot is controlled to execute the second target expression control instruction to achieve a second preset desired robot facial expression; the second preset desired robot facial expression corresponds to a preset number of individual facial expressions; during the execution of the second target expression control instruction by the target bionic robot, the second robot visual facial expression of the target bionic robot is captured by a preset camera. That is, during the dynamic travel measurement phase, a fixed expression sequence is issued to control the target bionic robot to execute the second target expression control instruction containing a second preset desired robot facial expression that corresponds to a preset number of individual facial expressions (i.e., several facial expressions corresponding to the aforementioned fixed expression sequence), and during the robot's execution of this instruction, its second robot visual facial expression is captured by a preset camera. In this way, this embodiment accurately obtains the correspondence between the visual data of the robot's actual facial presentation under a single expression and the preset expected data, providing comparative data for subsequent calibration and standardization of the single expression dimension, ensuring the pertinence and accuracy of the hardware compensation parameter vector construction, and making the presentation effect of the bionic robot's single expression more closely resemble the expression features of a real human face; it also batch obtains the correspondence between the visual data of the robot's actual facial presentation and the preset expected data under multiple expression scenarios, providing comparative data for subsequent comprehensive calibration and standardization of multiple expression dimensions, ensuring the comprehensiveness and adaptability of the hardware compensation parameter vector construction, and making the effect of the bionic robot when switching and presenting multiple expressions more closely resemble the expression features and change patterns of a real human face.

[0065] Step S12: Based on the robot's visual facial expression and the preset expected robot facial expression, calibrate and standardize the target bionic robot to construct the hardware compensation parameter vector of the target bionic robot.

[0066] In this embodiment, the target bionic robot is calibrated based on the captured visual facial expressions of the robot and the preset expected facial expressions of the robot, and a hardware compensation parameter vector adapted to the target bionic robot is constructed accordingly.

[0067] It should be noted that the hardware compensation parameter vector includes a hardware offset compensation parameter vector Offset and a hardware gain compensation parameter vector Gain. That is, each servo motor corresponds to two compensation parameters: offset and gain. The robot visual facial expression includes a first robot visual facial expression and a second robot visual facial expression, which correspond to the determination of Offset and Gain, respectively. The process for determining the hardware offset compensation parameter vector in the static benchmark calibration stage is as follows: Calculate the corresponding facial feature residual matrix based on the first robot visual facial expression and the first preset expected robot facial expression; if there is a first target dimension in the facial feature residual matrix with a residual value not less than a preset residual threshold, then adjust the hardware offset compensation parameter vector corresponding to the first target dimension based on a preset regularization technique until a preset convergence condition is met; wherein, the preset convergence condition includes that the residual value of the first target dimension is less than the preset residual threshold, and that after a preset number of adjustments, the degree of change of the first robot visual facial expression is less than a preset degree of change threshold. That is, based on the first robot's visual facial expression and the first preset expected robot facial expression, the facial feature residual matrix corresponding to the 52-dimensional ARKit features is calculated, and the specific calculation formula is as follows:

[0068] ;

[0069] in, The facial feature residual matrix; Corresponding to the facial expression feature vector exhibited by the robot in the first robot's visual facial expression; This corresponds to the target expectation vector in the first preset expected robot facial expression. After obtaining the facial feature residual matrix, if there is a first target dimension in the matrix with a residual value not less than the preset residual threshold Threshold, the hardware offset compensation parameter vector corresponding to that dimension is adjusted based on the preset regularization technique. Specifically, a step-increment method is adopted, adjusting the offset of the channel in reverse with a step size of 1μs (pulse width signal), and continuing to run until the preset convergence condition is met.

[0070] It should be further noted that the process for determining the hardware gain compensation parameter vector in the dynamic travel measurement stage is as follows: Based on the second robot's visual facial expression, determine the first facial expression amplitude of the target bionic robot, and determine the second facial expression amplitude of the second preset desired robot facial expression; based on the first facial expression amplitude and the second facial expression amplitude, determine the corresponding facial expression amplitude residual matrix; if there is a second target dimension in the facial expression amplitude residual matrix with a residual value not less than a preset residual threshold, then adjust the hardware gain compensation parameter vector corresponding to the second target dimension based on a preset regularization technique until a preset convergence condition is met; wherein, the preset convergence condition includes that the residual value of the second target dimension is less than the preset residual threshold, and that after a preset number of adjustments, the degree of change in the second robot's visual facial expression is lower than a preset degree of change threshold. That is, based on the second robot's visual facial expression, the actual first facial expression amplitude of the target bionic robot is determined, and simultaneously, the second facial expression amplitude corresponding to the second preset expected robot facial expression is determined. Then, a facial expression amplitude residual matrix is ​​calculated based on the two expression amplitudes. The calculation process of the facial expression amplitude residual matrix can be found in the above-described facial feature residual matrix calculation process, and will not be repeated here. If there is a second target dimension in the facial expression amplitude residual matrix with a residual value not less than a preset residual threshold (Threshold), the hardware gain compensation parameter vector corresponding to the second target dimension is adjusted specifically based on a preset regularization technique. For example, if a given parameter has reached 1.0, but visual capture only reaches 0.8, the gain coefficient compensation for that channel is increased until a preset convergence condition is met. It is understood that, in addition to using a fixed expression sequence issued by the system, this embodiment can also use a preset camera to simultaneously capture the target human face and the bionic robot head for calibration with the same input source, in order to achieve real-time dynamic travel measurement.

[0071] It should be noted that after calibration, this embodiment can automatically write the converged optimal hardware compensation parameter vector into a local configuration file in a specific format, such as JSON (a lightweight data exchange format). During subsequent startups, the hardware-specific calibration parameters will be directly loaded, enabling rapid plug-and-play between the model and the hardware. Furthermore, to prevent parameter oscillations caused by visual noise, this embodiment introduces a time integral filtering mechanism during the determination of the hardware compensation parameter vector, specifically by setting a threshold: setting the mean absolute error. The parameters can be adjusted according to the actual situation, corresponding to the preset residual threshold; a smoothing parameter is introduced during the adjustment of the hardware compensation parameter vector, and a smoothing parameter is adopted. The regularization constraint correction magnitude, also known as the aforementioned preset regularization technique, avoids abrupt parameter changes. Its corresponding calculation formula is as follows:

[0072] ;

[0073] in, The learning rate; Characterizing the loss function; The hardware compensation parameter vector representing the t-th iteration; The partial derivative of the loss function with respect to the hardware compensation parameter vector is used to characterize the loss function.

[0074] It is understandable that in actual production, the inability to further improve similarity is often not an algorithmic problem, but rather due to servo motor stall, linkage reaching its limit, or the physical tension of the silicone skin reaching its limit. This embodiment proposes a mechanical saturation detection mechanism to determine convergence: if the system continuously modifies parameters in the same way for M consecutive calibration cycles, but the visual feedback change is less than the minimum threshold σ, then the degree of freedom is determined to have reached its physical limit. Once mechanical saturation is detected, the system will no longer forcibly correct the residual and consider the calibration successful. That is, the preset convergence condition includes accuracy convergence: Physical constraint convergence (i.e., the change in the robot's visual facial expression before and after the adjustment approaches zero): , The degree of change in the robot's visual facial expression is characterized by dual convergence to avoid damage to the servo motors due to excessive pursuit of perfect consistency. Therefore, the preset convergence conditions in the static benchmark calibration stage include that the residual value of the first target dimension is less than a preset residual threshold, and that after the hardware offset compensation parameter vector is adjusted a preset number of times, the degree of change in the first robot's visual facial expression is less than a preset change threshold. The preset convergence conditions in the dynamic travel measurement stage include that the residual value of the second target dimension is less than a preset residual threshold, and that after the hardware gain compensation parameter vector is adjusted a preset number of times, the degree of change in the second robot's visual facial expression is less than a preset change threshold. In this way, this embodiment determines compensation parameters for offset and gain respectively. By accurately quantifying and directionally regularizing the deviation of a single facial feature and the deviation of multiple facial amplitudes, it can specifically correct the hardware offset error and amplitude gain error in the facial expression presentation of the bionic robot. This effectively compensates for the expression presentation errors caused by hardware transmission and structural assembly, making the subsequent facial expression control of the robot more in line with the preset expectations and improving the accuracy and fidelity of expression presentation. By using regularization technology and relying on dual convergence conditions, the effect and boundary of parameter adjustment are strictly controlled, avoiding damage to the servo motor due to excessive pursuit of perfect consistency. This achieves a refined and dimensional construction of the hardware compensation parameter vector, effectively improving the adaptability of compensation parameters to robot hardware characteristics and providing a hardware compensation basis for the accurate synchronization of subsequent expressions. After calibration, the converged optimal hardware compensation parameter vector is persistently saved, and the hardware-specific calibration parameters are directly loaded during subsequent startup, realizing fast plug-and-play between the model and the hardware.

[0075] Step S13: By asynchronously running a preset perception thread and a preset drive thread, the facial expression synchronization operation of the target bionic robot is performed based on the hardware compensation parameter vector and the pre-trained target expression mapping model. The preset perception thread is used to acquire the target face image in real time, and after extracting the target expression feature vector of the target face image, it outputs the corresponding target servo control vector through the target expression mapping model. The preset drive thread is used to output servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector. The target expression mapping model is a model obtained by training based on a pre-generated training dataset. The training dataset consists of several expression feature vectors and corresponding servo control vectors. The expression feature vectors are obtained based on target expression parameters, which are preset expression parameters that conform to the anatomical logic of the human face.

[0076] In this embodiment, the facial expression synchronization operation of the target bionic robot is performed by asynchronously running a preset perception thread and a preset driving thread, combined with the hardware compensation parameter vector and the MLP model (i.e., the target expression mapping model) trained in advance based on the training dataset with MSE (Mean Squared Error) as the objective function. Specifically, the preset perception thread is responsible for real-time acquisition of target face images, extracting target expression feature vectors from the images, and outputting the corresponding target servo control vectors via the target expression mapping model. The preset driving thread, based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vectors, runs independently at a preset output frequency (e.g., 60Hz), outputting servo control commands to the control board at a constant speed. The target expression mapping model is used to obtain a nonlinear mapping from ARKit to servo vectors, mapping virtual ARKit parameters to physical servo vectors (e.g., 32-dimensional vectors). The training dataset consists of several expression feature vectors and their corresponding servo control vectors, and is a training dataset constructed based on expression mapping requirements. All expression feature vectors are obtained based on preset target expression parameters that conform to human facial anatomy; these target expression parameters are ARKit expression parameters.

[0077] It should be pointed out that, such as Figure 4The diagram illustrates the generation process of a training dataset provided in this application. The generation process of the training dataset is as follows: the target expression parameters are smoothed using an interpolation algorithm, and fractal Brownian motion noise is superimposed to obtain processed target expression parameters; under preset anatomical constraints, the joint probability distribution of the processed target expression parameters is sampled to obtain corresponding expression feature vectors, and the expression feature vectors are optimized to obtain corresponding optimized expression feature vectors; the optimized expression feature vectors are input into a target simulation tool to extract corresponding servo control vectors, resulting in the training dataset composed of the servo control vectors and the optimized expression feature vectors. In other words, to ensure that the 1.2 million parameter MLP model accurately covers the complex real-world facial expression space, rather than wasting feature representation capabilities in an ineffective mechanical motion space, a hierarchical randomization algorithm is employed. First, macroscopic and microscopic motions are superimposed. At the macroscopic level, an interpolation algorithm smooths the target expression parameters to generate a smooth expression transition trajectory. At the microscopic level, fractal Brownian motion (fBm) noise is superimposed, with multiple layers of noise at different frequencies and amplitudes simulating varying degrees of facial muscle tremors, resulting in the processed target expression parameters. It is understandable that human facial movements are not isolated movements of individual muscle groups, but rather coordinated movements constrained by anatomical structures. For example, the degree of eyelid opening and closing is physiologically strongly correlated with the angle of eyeball rotation. If ARKit features (such as left and right eyelid closure) are generated independently and randomly, when applying anatomical linkage constraints (such as the need for feature subtraction or difference in binocular coordinated movement), according to probability theory, the mean of the feature distribution will tend to 0, resulting in feature cancellation. This leads to missing sampling points for extreme expressions or specific linked expressions, causing distribution drift. Assuming two related features (such as left eyelid closure degree)... Right eyelid closure degree Independent, has:

[0078] Mean: ;

[0079] variance: ;

[0080] Covariance: ;

[0081] in, , The mean values ​​representing the degree of closure of the left and right eyelids are respectively used. A numerical value representing the mean; , The variances representing the degree of closure of the left and right eyelids, respectively. A numerical value representing variance; Characterizing covariance. In anatomy, the characteristics of co-movement are subtracted: According to the properties of linear expectation: It is known that this will lead to data in Interval stacking occurs, while data is insufficient in extreme intervals, leading to missing model gradients. To address this, after obtaining the processed target expression parameters, under preset anatomical constraints, the parameters are sampled by joint probability distribution to obtain the corresponding expression feature vector and optimized. The optimized expression feature vector is then input into the target simulation tool, i.e., the Blender system, to extract the corresponding servo control vector, i.e., the 27-dimensional servo command. This constructs a dual high-precision training dataset (ARKit, servo vector) consisting of the optimized expression feature vector and the servo control vector, thereby eliminating the uncanny valley effect and abnormal mechanical linkage at the source.

[0082] It should be further noted that the processing flow for obtaining the expression feature vector and optimizing it to obtain the corresponding optimized expression feature vector is as follows: Based on a preset anatomical correlation matrix, a joint probability distribution of each dimension variable in the processed target expression parameters is generated; the preset anatomical correlation matrix is ​​used to define the physical coupling ratio between the target expression parameters; the joint probability distribution is sampled to obtain the corresponding expression feature vector; the target feature value in the expression feature vector is determined; the target feature value is a feature value that exceeds a preset feature value range; the target feature value is adjusted by a preset logarithmic compression function to control the target feature value within the preset feature value range to obtain the adjusted expression feature vector; Gaussian white noise is superimposed on the adjusted expression feature vector and processed using a low-pass filter to obtain the optimized expression feature vector. That is, a preset anatomical correlation matrix M is introduced to define the target expression parameters, i.e., the physical coupling ratio between the 55-dimensional motion coefficients of the face. Independent dimension sampling is abandoned, and instead, based on the preset anatomical correlation matrix, a joint probability distribution of the variables of each dimension of the processed target expression parameters is generated. The joint probability distribution is sampled to obtain the corresponding expression feature vector, avoiding the mean zeroing phenomenon caused by feature cancellation and ensuring that the model has sufficient sample density in the linked expression region. The mathematical explanation of the above process is as follows: Setting a multi-dimensional vector of ARKit expression features. During the generation process, the coupling relationship between preset dimensions of the covariance matrix Σ (derived from the preset anatomical correlation matrix) ensures the application of constraint operators. In other words, after representing the anatomical constraint relationship between two related features, the feature space still covers the complete domain interval, avoiding the phenomenon of the mean returning to zero and ensuring the accuracy of the model in expressing linked expressions.

[0083] It is understandable that the ARKit feature domain is [0, 1]. Traditional hard truncation or linear reflection can lead to singularity accumulation at the boundaries of the probability density function (PDF). Due to the complex viscoelasticity of human facial tissue, it exhibits significant nonlinear stiffness enhancement when displacement approaches its limit. Simple linear mapping cannot simulate this complex mechanical feedback, resulting in a lack of subtle nuances in the trained model at extreme facial expressions. This embodiment introduces a nonlinear viscoelastic damping mapping algorithm based on biomechanical properties. This algorithm simulates the nonlinear tensile resistance of skin tissue through a logarithmic compression function. Specifically, after obtaining the facial feature vector, target feature values ​​exceeding the preset feature value range (i.e., [0, 1]) are selected, which are the out-of-bounds original feature values ​​(…). Instead of performing hard truncation, it "soft-locks" the target feature values ​​within the effective range according to a nonlinear mapping function, i.e., a preset logarithmic compression function. Specifically, it adjusts the target feature values ​​to make them fall within the preset feature value range, resulting in an adjusted facial expression feature vector. This ensures that the generated training data maintains a smooth gradient even at extreme positions, effectively solving the problem of probability density accumulation at boundaries. This allows the model to maintain high sensitivity and linear adjustment capability even with extreme expressions, simulating the elastic resistance of facial muscles reaching their physiological limits. This makes the generated action sequences visually more consistent with biological logic and effectively overcomes the uncanny valley effect. The specific formula for the preset logarithmic compression function is as follows:

[0084] ;

[0085] in, This is a preset organizational stiffness coefficient used to adjust the compression ratio near the boundary; The adjusted facial expression feature vector represents the corrected feature value corresponding to the target feature value. After obtaining the adjusted facial expression feature vector, Gaussian white noise is superimposed on it and processed using a low-pass filter to simulate the noise characteristics of the camera-acquired signal in actual use and the fluctuations of the MediaPipe (a machine learning solution framework) algorithm. This ensures that the data distribution during offline training is highly aligned with the input distribution during online inference, resulting in the optimized facial expression feature vector.

[0086] It should be noted that in the field of biomimetic control, the latency of visual perception algorithms (such as MediaPipe) and fluctuations in the physical frame rate of camera hardware can cause the perceived signal to exhibit a non-uniform distribution on the time axis. Directly sending this signal to a high-response-frequency servo system can cause mechanical vibration and abrupt changes in motion. To reduce latency and hardware impact and avoid servo drive blockage, this embodiment achieves deep decoupling between the perception frequency and the drive frequency through a dual-thread asynchronous architecture, specifically including a preset drive thread and a preset perception thread. The preset drive thread acts as a higher-level high-frequency periodic thread; the preset perception thread acts as a lower-level non-periodic thread. Due to changes in ambient lighting or fluctuations in algorithm computing power, the time interval between its output ARKit features and servo control vectors is not constant. Figure 5 The diagram illustrates the process flow of a dual-thread asynchronous architecture provided in this application. First, a preset perception thread receives data in real-time, performing image acquisition, MediaPipe feature extraction, and mapping model inference. A preset drive thread calculates interpolation vectors to achieve stable robot control. The operation flow of the preset drive thread is as follows: Within the time interval between two consecutive outputs of the target servo control vector by the preset perception thread, the current servo state of the target bionic robot and the latest target servo control vector output by the preset perception thread are determined based on a preset output frequency. Based on the current servo state and the latest target servo control vector, a transition frame-filled servo control vector is generated using a linear interpolation algorithm. The transition frame-filled servo control vector is corrected using the hardware compensation parameter vector, and corresponding transition frame-filled servo control commands are generated and output based on the corrected results. That is, the preset driving thread does not directly transmit the output of the preset sensing thread, but maintains a buffer between the current position frame and the target desired frame in real time. During the time interval between two consecutive outputs of the target servo control vector by the preset sensing thread, i.e., the update interval, the current servo state of the target bionic robot and the latest target servo control vector output by the preset sensing thread are first determined according to the preset output frequency. Then, combining the current servo state and the latest target servo control vector, a frame-supplementing servo control vector for the intermediate position is generated through a linear interpolation algorithm to ensure a smooth transition of the servo output pulses and eliminate visual jumps caused by frame loss or low frame rate at the sensing end. Subsequently, the frame-supplementing servo control vector is used as input, and the hardware compensation parameter vector and the correction function established through the visual closed loop are used to correct the frame-supplementing servo control vector. Based on the correction result, the corresponding frame-supplementing servo control command is generated and output through linear transformation. The specific formula of the correction function is as follows:

[0087] ;

[0088] in, The corrected servo parameters are characterized. In addition, this embodiment can filter high-frequency noise in the sensing signal through low-pass filtering, limit the servo's displacement per unit time through velocity and acceleration constraints to protect the physical mechanical structure, and perform deadzone processing to filter frequent jitter caused by minute noise, extending the servo's lifespan. Thus, this embodiment adopts a dual-thread asynchronous control architecture, combined with a time-delay compensation-based frame interpolation algorithm, ensuring efficient parallel processing of facial expression acquisition and servo command output. The frame interpolation algorithm solves problems such as insufficient camera hardware and low frame rate leading to inability to detect / control in real time, effectively solving mechanical vibration and motion jump problems caused by mismatch between visual perception and drive frequency, ensuring that the smoothness of the hardware execution end is not affected by fluctuations at the sensing end; by correcting the hardware-level expression presentation error through hardware compensation parameter vectors, the accuracy, smoothness, and real-time performance of the bionic robot's facial expression synchronization can be effectively improved, allowing the robot's facial expressions to more closely resemble the expression states and rhythms of real human faces; an anatomically constrained data generation algorithm is adopted, introducing anatomical correlation matrices and multi-dimensional variable joint distribution sampling technology, utilizing Bleu... The nder simulation environment acts as a truth generator, inputting ARKit facial expression parameters to calculate the servo control vector under ideal conditions, thereby constructing a dataset. This effectively solves the problems of feature cancellation and distribution drift caused by the independent random generation of feature dimensions in traditional methods, making the generated dataset more consistent with human anatomy and improving the model's ability to express complex facial expressions. By superimposing fractal Brownian motion noise and Gaussian white noise and employing nonlinear viscoelastic reflection processing technology, the generated facial expressions are made more natural and realistic, with long-range correlation and subtle tremors, improving the quality of the bionic robot's facial expression performance. The conversion from facial feature vectors to servo control vectors is directly achieved using an MLP deep learning model, eliminating the need for a complex 3D simulation engine as an intermediary, greatly simplifying the system structure, reducing R&D costs, and enabling the deployment of lightweight equipment.

[0089] As can be seen from the above, the embodiments of this application control the target bionic robot to execute target expression control commands containing preset desired robot facial expressions. Simultaneously, during command execution, a preset camera captures the robot's visual facial expressions. Based on these visual facial expressions and the preset desired robot facial expressions, the target bionic robot is calibrated to construct a hardware compensation parameter vector adapted to the robot. Subsequently, by asynchronously running a preset perception thread and a preset drive thread, combined with the hardware compensation parameter vector and a target expression mapping model pre-trained based on a training dataset consisting of several expression feature vectors and their corresponding servo control vectors, the target bionic robot's facial expression synchronization operation is executed. The preset perception thread is responsible for real-time acquisition of target face images, extracting target expression feature vectors from the images, and outputting the corresponding target servo control vectors via the target expression mapping model. The preset drive thread, based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vectors, outputs servo control commands at a preset output frequency. Furthermore, the expression feature vectors in the training dataset are obtained based on preset target expression parameters that conform to human facial anatomy. In this way, through the process described in the embodiments of this application, the conversion from facial feature vectors to servo control vectors is directly achieved using the target facial expression mapping model, without the need for a complex 3D simulation engine as an intermediary. This simplifies the system structure, reduces R&D costs, and enables the deployment of lightweight equipment. The dual-thread asynchronous control architecture, combined with frame interpolation algorithms, effectively solves the problems of mechanical vibration and motion jumps caused by the mismatch between visual perception and drive frequency, ensuring that the smoothness of the hardware execution end is not affected by fluctuations at the perception end. Through the parameter correction mechanism, the control parameters can be dynamically adjusted according to the actual performance of different robots, effectively solving the problem of individual differences caused by production assembly errors, reducing manual debugging costs, and thus realizing a lightweight bionic robot facial expression synchronization method to improve the flexibility and intelligence of expression synchronization.

[0090] Accordingly, see Figure 6 As shown in the illustration, this application also provides a vision-perception-based bionic robot facial expression synchronization device, comprising:

[0091] The data capture module 11 is used to control the target bionic robot to execute target expression control commands to achieve a preset desired robot facial expression, and to capture the robot visual facial expression of the target bionic robot through a preset camera during the execution process.

[0092] The calibration module 12 is used to calibrate the target bionic robot based on the robot's visual facial expression and the preset expected robot facial expression, so as to construct the hardware compensation parameter vector of the target bionic robot.

[0093] The synchronization implementation module 13 is used to perform facial expression synchronization operations of the target bionic robot based on the hardware compensation parameter vector and the pre-trained target expression mapping model by asynchronously running a preset perception thread and a preset drive thread; the preset perception thread is used to acquire the target face image in real time, and after extracting the target expression feature vector of the target face image, output the corresponding target servo control vector through the target expression mapping model; the preset drive thread is used to output servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector and the target servo control vector.

[0094] The target expression mapping model is a model obtained by training a pre-generated training dataset; the training dataset consists of several expression feature vectors and corresponding servo control vectors; the expression feature vectors are obtained based on target expression parameters, which are preset expression parameters that conform to the anatomical logic of the human face.

[0095] In some specific embodiments, the vision-perception-based bionic robot facial expression synchronization device may specifically include:

[0096] The noise superposition unit is used to smooth the target expression parameters through an interpolation algorithm and superimpose fractal Brownian motion noise to obtain the processed target expression parameters.

[0097] The vector optimization submodule is used to sample the joint probability distribution of the processed target expression parameters under preset anatomical constraints to obtain the corresponding expression feature vector, and optimize the expression feature vector to obtain the corresponding optimized expression feature vector.

[0098] The vector input unit is used to input the optimized facial expression feature vector into the target simulation tool to extract the corresponding servo control vector, thereby obtaining the training dataset composed of the servo control vector and the optimized facial expression feature vector.

[0099] In some specific implementations, the vector optimization submodule may specifically include:

[0100] The distribution generation unit is used to generate the joint probability distribution of each dimension variable in the processed target expression parameters based on a preset anatomical correlation matrix; the preset anatomical correlation matrix is ​​used to define the physical coupling ratio between the target expression parameters.

[0101] A distributed sampling unit is used to sample the joint probability distribution to obtain the corresponding facial expression feature vector;

[0102] A feature value determination unit is used to determine a target feature value in the facial expression feature vector; the target feature value is a feature value that exceeds a preset feature value range.

[0103] The feature value adjustment unit is used to adjust the target feature value through a preset logarithmic compression function to control the target feature value within the preset feature value range, thereby obtaining the adjusted facial expression feature vector;

[0104] The filtering unit is used to superimpose Gaussian white noise onto the adjusted facial expression feature vector and process it using a low-pass filter to obtain an optimized facial expression feature vector.

[0105] In some specific embodiments, the data capture module 11 may specifically include:

[0106] The first instruction execution unit is used to control the target bionic robot to execute a first target expression control instruction for achieving a first preset desired robot facial expression; the first preset desired robot facial expression corresponds to a human face expression.

[0107] The first data capture unit is used to capture the first robot visual facial expression of the target bionic robot through a preset camera during the process of the target bionic robot executing the first target expression control command.

[0108] In some specific embodiments, the calibration module 12 may specifically include:

[0109] The matrix calculation unit is used to calculate the corresponding facial feature residual matrix based on the first robot visual facial expression and the first preset expected robot facial expression.

[0110] The first parameter adjustment unit is used to adjust the hardware offset compensation parameter vector corresponding to the first target dimension based on the preset regularization technique if there is a first target dimension in the facial feature residual matrix with a residual value not less than a preset residual threshold, until the preset convergence condition is met.

[0111] The preset convergence conditions include that the residual value of the first target dimension is less than the preset residual threshold, and that after the hardware offset compensation parameter vector has been adjusted a preset number of times, the degree of change of the first robot visual facial expression is lower than the preset degree of change threshold.

[0112] In some specific embodiments, the data capture module 11 may specifically include:

[0113] The second instruction execution unit is used to control the target bionic robot to execute a second target expression control instruction for achieving a second preset desired robot facial expression; the second preset desired robot facial expression corresponds to a preset number of individual facial expressions.

[0114] The second data capture unit is used to capture the second robot visual facial expression of the target bionic robot through a preset camera during the process of the target bionic robot executing the second target expression control command.

[0115] Accordingly, the calibration module 12 may specifically include:

[0116] An amplitude determination unit is used to determine the first facial expression amplitude of the target bionic robot based on the second robot's visual facial expression, and to determine the second facial expression amplitude of the second preset desired robot facial expression.

[0117] The matrix determination unit is used to determine the corresponding facial expression amplitude residual matrix based on the first facial expression amplitude and the second facial expression amplitude.

[0118] The second parameter adjustment unit is used to adjust the hardware gain compensation parameter vector corresponding to the second target dimension based on the preset regularization technique if there is a second target dimension in the facial expression amplitude residual matrix with a residual value not less than a preset residual threshold, until the preset convergence condition is met.

[0119] The preset convergence conditions include that the residual value of the second target dimension is less than the preset residual threshold, and that after the hardware gain compensation parameter vector has been adjusted a preset number of times, the degree of change of the second robot's visual facial expression is lower than the preset degree of change threshold.

[0120] In some specific embodiments, the synchronization implementation module 13 may specifically include:

[0121] The vector determination unit is used to determine the current servo state of the target bionic robot and the latest target servo control vector output by the preset sensing thread within the time interval between two consecutive outputs of the target servo control vector by the preset sensing thread, based on a preset output frequency.

[0122] The vector generation unit is used to generate a frame-splitting servo control vector for transition based on the current servo state and the latest target servo control vector using a linear interpolation algorithm.

[0123] The vector correction unit is used to correct the frame-filling servo control vector using the hardware compensation parameter vector, so as to generate and output the corresponding frame-filling servo control command based on the obtained correction result.

[0124] Furthermore, embodiments of this application also disclose an electronic device, Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the vision-perception-based bionic robot facial expression synchronization method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0125] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0126] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0127] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the vision-perception-based bionic robot facial expression synchronization method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0128] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned visual perception-based bionic robot facial expression synchronization method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0129] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0130] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0131] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0132] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0133] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for synchronizing facial expressions in a bionic robot based on visual perception, characterized in that, include: The target bionic robot is controlled to execute target expression control commands to achieve a preset desired robot facial expression, and during the execution process, the robot visual facial expression of the target bionic robot is captured by a preset camera; The target bionic robot is calibrated and standardized based on the robot's visual facial expressions and the preset expected robot facial expressions in order to construct the hardware compensation parameter vector of the target bionic robot. By asynchronously running a preset perception thread and a preset driving thread, the facial expression synchronization operation of the target bionic robot is performed based on the hardware compensation parameter vector and the pre-trained target expression mapping model. The preset perception thread is used to acquire the target face image in real time, and after extracting the target expression feature vector of the target face image, it outputs the corresponding target servo control vector through the target expression mapping model. The preset driving thread is used to output servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector. The target expression mapping model is a model obtained by training a pre-generated training dataset; the training dataset consists of several expression feature vectors and corresponding servo control vectors; the expression feature vectors are obtained based on target expression parameters, which are preset expression parameters that conform to the anatomical logic of the human face; The process of generating the training dataset includes: The target facial expression parameters are smoothed by an interpolation algorithm and then superimposed with fractal Brownian motion noise to obtain the processed target facial expression parameters. Under preset anatomical constraints, the joint probability distribution of the processed target facial expression parameters is sampled to obtain the corresponding facial expression feature vector, and the facial expression feature vector is optimized to obtain the corresponding optimized facial expression feature vector. The optimized facial expression feature vector is input into the target simulation tool to extract the corresponding servo control vector, thereby obtaining the training dataset composed of the servo control vector and the optimized facial expression feature vector.

2. The method for synchronizing facial expressions of a bionic robot based on visual perception according to claim 1, characterized in that, Under preset anatomical constraints, the process involves sampling the joint probability distribution of the processed target facial expression parameters to obtain a corresponding facial expression feature vector, and then optimizing the facial expression feature vector to obtain a corresponding optimized facial expression feature vector. This includes: Based on a preset anatomical correlation matrix, a joint probability distribution of each dimension variable in the processed target expression parameters is generated; the preset anatomical correlation matrix is ​​used to define the physical coupling ratio between the target expression parameters. The joint probability distribution is sampled to obtain the corresponding facial expression feature vector; Determine the target feature value in the facial expression feature vector; the target feature value is a feature value that exceeds a preset feature value range; The target feature value is adjusted by a preset logarithmic compression function to control the target feature value within the preset feature value range, thereby obtaining the adjusted facial expression feature vector; Gaussian white noise is superimposed on the adjusted facial expression feature vector, and then processed using a low-pass filter to obtain the optimized facial expression feature vector.

3. The method for synchronizing facial expressions of a bionic robot based on visual perception according to claim 1, characterized in that, The controlled target bionic robot executes target expression control commands to achieve a preset desired robot facial expression, and during execution, captures the robot's visual facial expression through a preset camera, including: The target bionic robot is controlled to execute a first target expression control command to achieve a first preset desired robot facial expression; the first preset desired robot facial expression corresponds to a human face expression. During the execution of the first target facial expression control command by the target bionic robot, the first robot visual facial expression of the target bionic robot is captured by a preset camera.

4. The method for synchronizing facial expressions of a bionic robot based on visual perception according to claim 3, characterized in that, The calibration and standardization of the target bionic robot based on the robot's visual facial expressions and the preset expected robot facial expressions, to construct the hardware compensation parameter vector of the target bionic robot, includes: Calculate the corresponding facial feature residual matrix based on the first robot visual facial expression and the first preset expected robot facial expression; If there is a first target dimension in the facial feature residual matrix with a residual value not less than a preset residual threshold, then the hardware offset compensation parameter vector corresponding to the first target dimension is adjusted based on the preset regularization technique until the preset convergence condition is met. The preset convergence conditions include that the residual value of the first target dimension is less than the preset residual threshold, and that after the hardware offset compensation parameter vector has been adjusted a preset number of times, the degree of change of the first robot visual facial expression is lower than the preset degree of change threshold.

5. The method for synchronizing facial expressions of a bionic robot based on visual perception according to claim 1, characterized in that, The controlled target bionic robot executes target expression control commands to achieve a preset desired robot facial expression, and during execution, captures the robot's visual facial expression through a preset camera, including: The target bionic robot is controlled to execute a second target facial expression control command to achieve a second preset desired robot facial expression; the second preset desired robot facial expression corresponds to a preset number of individual facial expressions. During the process of the target bionic robot executing the second target expression control command, the second robot visual facial expression of the target bionic robot is captured by a preset camera; Accordingly, the calibration and standardization of the target bionic robot based on the robot's visual facial expression and the preset expected robot facial expression, to construct the hardware compensation parameter vector of the target bionic robot, includes: Based on the second robot's visual facial expression, the first facial expression amplitude of the target bionic robot is determined, and the second facial expression amplitude of the second preset desired robot facial expression is determined. Based on the first facial expression amplitude and the second facial expression amplitude, determine the corresponding facial expression amplitude residual matrix; If there is a second target dimension in the facial expression amplitude residual matrix with a residual value not less than a preset residual threshold, then the hardware gain compensation parameter vector corresponding to the second target dimension is adjusted based on the preset regularization technique until the preset convergence condition is met. The preset convergence conditions include that the residual value of the second target dimension is less than the preset residual threshold, and that after the hardware gain compensation parameter vector has been adjusted a preset number of times, the degree of change of the second robot's visual facial expression is lower than the preset degree of change threshold.

6. The method for synchronizing facial expressions of a bionic robot based on visual perception according to any one of claims 1 to 5, characterized in that, The process of outputting servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector includes: Within the time interval between two consecutive outputs of the target servo control vector by the preset sensing thread, the current servo state of the target bionic robot and the latest target servo control vector output by the preset sensing thread are determined based on the preset output frequency. Based on the current servo state and the latest target servo control vector, a transition frame-filling servo control vector is generated using a linear interpolation algorithm; The frame-complementing servo control vector is corrected using the hardware compensation parameter vector, and corresponding frame-complementing servo control commands are generated and output based on the correction results.

7. A bionic robot facial expression synchronization device based on visual perception, characterized in that, include: The data capture module is used to control the target bionic robot to execute target expression control commands to achieve a preset desired robot facial expression, and to capture the robot visual facial expression of the target bionic robot through a preset camera during the execution process; The calibration module is used to calibrate the target bionic robot based on the robot's visual facial expression and the preset expected robot facial expression, so as to construct the hardware compensation parameter vector of the target bionic robot. The synchronization implementation module is used to perform facial expression synchronization operations of the target bionic robot based on the hardware compensation parameter vector and the pre-trained target expression mapping model by asynchronously running a preset perception thread and a preset drive thread. The preset perception thread is used to acquire target face images in real time, and after extracting the target expression feature vector of the target face image, output the corresponding target servo control vector through the target expression mapping model. The preset drive thread is used to output servo control commands at a preset output frequency based on a preset frame interpolation algorithm, the hardware compensation parameter vector, and the target servo control vector. The target expression mapping model is a model obtained by training a pre-generated training dataset; the training dataset consists of several expression feature vectors and corresponding servo control vectors; the expression feature vectors are obtained based on target expression parameters, which are preset expression parameters that conform to the anatomical logic of the human face; The vision-perception-based bionic robot facial expression synchronization device includes: The noise superposition unit is used to smooth the target expression parameters through an interpolation algorithm and superimpose fractal Brownian motion noise to obtain the processed target expression parameters. The vector optimization submodule is used to sample the joint probability distribution of the processed target expression parameters under preset anatomical constraints to obtain the corresponding expression feature vector, and optimize the expression feature vector to obtain the corresponding optimized expression feature vector. The vector input unit is used to input the optimized facial expression feature vector into the target simulation tool to extract the corresponding servo control vector, thereby obtaining the training dataset composed of the servo control vector and the optimized facial expression feature vector.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the visual perception-based bionic robot facial expression synchronization method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the visual perception-based bionic robot facial expression synchronization method as described in any one of claims 1 to 6.