Digital human face binding generation method based on expression capture

By collecting and optimizing digital human facial binding data in real time, the problems of low efficiency and poor adaptability in existing technologies have been solved, achieving naturalness and dynamic expression of digital human facial expressions, which is suitable for film and television, virtual live streaming and interactive games.

CN121214518AActive Publication Date: 2025-12-26ZHEJIANG VERSATILE MEDIA
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511390476.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-26
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing digital human facial binding methods rely on manual adjustments or preset templates, resulting in low efficiency, poor adaptability, and an inability to capture subtle facial movements, affecting the realism of expressions and user experience.

Method used

Facial data is collected in real time using facial expression capture devices, features are extracted and fitted, initial binding rules are generated, and binding optimization is performed to drive the facial movements of the digital human model.

Benefits of technology

It achieves naturalness and dynamic expression of digital human facial expressions, improves generation efficiency, adapts to different facial contours and detail designs, and enhances the application experience in film and television, virtual live streaming and interactive games.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121214518A_ABST
    Figure CN121214518A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital human face, and discloses a digital human face binding generation method based on expression capture. The method comprises the following steps: acquiring facial expression data of a target object in real time through an expression capturing device, and obtaining facial subtle motion information; performing feature extraction on the collected data, and filtering redundant information to obtain an expression feature vector set reflecting facial motion key features; fitting a face continuous movement track according to the set, and generating an initial binding rule conforming to a real human body movement rule; mapping the initial face binding data to a face topological structure of the digital human model, and generating initial face binding data of an adaptive model; systematically optimizing the initial data, correcting a motion parameter problem, and obtaining optimized face binding data; and driving the face movement of the digital human model by using the optimized data. According to the method, the expression of the digital human is more natural, the movement is smoother, the manual operation is reduced, the binding efficiency is improved, and the method is suitable for scenes such as movies and televisions, virtual live broadcast and interactive games.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital human face, in particular to a digital human face binding generation method based on expression capture. BACKGROUND

[0002] With the rapid penetration of digital technology in the fields of film and television production, virtual live broadcast, interactive games, etc., the realism and real-time performance of digital human facial expressions have become one of the core indicators for measuring the application effect of digital humans. Currently, digital human face binding is mainly achieved by manual adjustment or semi-automatic method based on preset templates. In the process of manual binding, technicians need to set the bone nodes and muscle movement parameters of the digital human face one by one according to experience, which not only requires a lot of time and labor cost, but also is prone to inconsistent binding results due to individual experience differences, making it difficult to meet the efficiency requirements of large-scale digital human production. Although the semi-automatic binding method based on preset templates reduces manual operation to some extent, the templates are usually only applicable to digital human models with specific facial topological structures. When facing digital humans with different facial contours and different detailed features, a large number of modifications need to be made to the templates to adapt to them, and the flexibility is poor. At the same time, the existing binding methods often lack accurate processing of real-time expression data when associating real human facial expressions with digital human facial movements, and can only achieve simple expression mapping, failing to capture subtle changes in facial movements such as eye corners and mouth corners, resulting in stiff and unnatural digital human facial expressions. After generating the initial binding data, the existing technology has relatively single means for optimizing the binding effect, mainly focusing on adjusting the movement range of individual bone nodes, and failing to systematically optimize the coordination of the overall facial topological structure and the smoothness of the movement trajectory, making the digital human prone to uncoordinated facial component movements, lag, and other situations when performing complex expression conversion, further affecting the realism of the digital human expression and the user's interactive experience. The existence of these problems restricts the widespread application of digital human technology in fields with high requirements for expression realism, such as virtual social interaction, film and television character creation, etc. SUMMARY

[0003] The purpose of the present application is to provide a digital human face binding generation method based on expression capture to solve the problems raised in the background art.

[0004] To achieve the above purpose, the present application provides a digital human face binding generation method based on expression capture, which comprises: real-time facial expression data of a target object is collected by an expression capture device; feature extraction is performed on the real-time facial expression data to obtain an expression feature vector set; performing face motion trajectory fitting according to the expression feature vector set to generate initial binding rules; mapping the initial binding rules to a face topology structure of a digital human model to generate initial face binding data; performing binding optimization processing on the initial face binding data to generate optimized face binding data; driving face motion of the digital human model by using the optimized face binding data.

[0005] Preferably, real-time facial expression data of a target object is collected in real time by an expression capture device, including: obtaining a sequence of facial key point coordinates of the target object by the expression capture device; comparing the sequence of facial key point coordinates with a preset standard expression template to obtain expression offset data; labeling the expression offset data as real-time facial expression data.

[0006] Preferably, the initial face binding data is subjected to binding optimization processing to generate optimized face binding data, including: calling face skeleton constraint parameters of the digital human model; performing matching degree calculation on the initial face binding data and the face skeleton constraint parameters to obtain binding deviation data; adjusting the initial binding rules according to the binding deviation data to generate optimized binding rules; updating the initial face binding data based on the optimized binding rules to generate optimized face binding data.

[0007] Preferably, the initial binding rules are adjusted according to the binding deviation data to generate optimized binding rules, including: extracting key motion axis deviation values in the binding deviation data; comparing the key motion axis deviation values with a preset deviation threshold to locate abnormal binding nodes; reconstructing motion weight distribution of the abnormal binding nodes to generate optimized binding rules.

[0008] Preferably, the motion weight distribution of the abnormal binding nodes is reconstructed to generate optimized binding rules, including: collecting historical motion smoothness data of the abnormal binding nodes; performing fusion calculation on the historical motion smoothness data and the current key motion axis deviation values to obtain node optimization coefficients; redistributing motion weight values of the abnormal binding nodes according to the node optimization coefficients to generate optimized binding rules.

[0009] Preferably, historical motion smoothness data of the abnormal binding node is collected, including: Obtaining a motion trajectory change amount of the abnormal binding node in consecutive time frames; Calculating a distribution dispersion of the motion trajectory change amount to obtain the historical motion smoothness data.

[0010] Preferably, the historical motion smoothness data is fused with a current key motion axis deviation value to obtain a node optimization coefficient, including: Normalizing the historical motion smoothness data to a preset smoothness reference range; Weighted summing the normalized historical motion smoothness data and the current key motion axis deviation value to obtain the node optimization coefficient.

[0011] Preferably, the motion weight value of the abnormal binding node is redistributed according to the node optimization coefficient to generate an optimized binding rule, including: Inputting the node optimization coefficient into a preset weight mapping table to match a target weight adjustment amount; Updating the initial motion weight value of the abnormal binding node by using the target weight adjustment amount to generate the optimized binding rule.

[0012] Preferably, after driving the facial motion of the digital human model by using the optimized facial binding data, including: Real-time monitoring facial motion deformation data of the digital human model; If the facial motion deformation data exceeds a preset deformation tolerance range, re-executing the step of collecting real-time facial expression data of a target object by using an expression capture device in real time.

[0013] Preferably, after real-time monitoring facial motion deformation data of the digital human model, including: Performing deformation feature analysis on the facial motion deformation data to generate expression error data; Comparing the expression error data with a preset expression tolerance threshold; If the expression error data exceeds the preset expression tolerance threshold, triggering dynamic updating of the optimized facial binding data, and re-executing the step of generating initial facial binding data based on the latest collected real-time facial expression data.

[0014] Compared with the prior art, the present application has the following advantages: The facial expression data of the target object is collected in real time by the expression capture device, which can directly obtain the dynamic information of the real human facial movement, breaking the limitations of traditional binding methods relying on preset templates or artificial experience, making the data source of digital human facial binding closer to real human expressions, and laying a foundation for subsequent generation of natural digital human facial movements. The real-time collection method can capture subtle changes in facial movements, such as slight contraction of facial muscles and slight displacement of organs. The integration of these subtle information can make the digital human facial expression more delicate and avoid the problem of stiff expression caused by the inability to capture subtle expressions in traditional methods. The facial expression data is extracted to obtain a set of expression feature vectors. This process can filter and refine the massive raw data collected, remove redundant information, and retain key features related to facial movement, making the subsequent facial movement trajectory fitting more targeted. By quantitatively representing the expression features in the form of feature vectors, abstract facial expressions can be converted into calculable and analyzable data, providing a reliable basis for precise fitting of facial movement trajectories, and thus making the generated initial binding rules more consistent with real facial movement rules and reducing the workload of subsequent binding adjustments. The facial movement trajectory fitting is performed according to the set of expression feature vectors to generate initial binding rules, which are then mapped to the facial topology structure of the digital human model to generate initial facial binding data. This process realizes the precise association between real facial expressions and digital human facial structures. The fitted movement trajectory can reflect the continuous change process of facial movement, avoiding the problem of incoherent movement caused by simple mapping in traditional binding, making the digital human facial movement more consistent with human physiological movement rules. At the same time, mapping based on the digital human facial topology structure can ensure that the initial binding data is compatible with the structural characteristics of the digital human model. Whether it is a digital human model with different facial contours or different detail designs, it can obtain matching initial binding data through the mapping process, improving the flexibility and adaptability of the method. The initial facial binding data is optimized to generate optimized facial binding data. The optimization process is not limited to individual node adjustment, but systematically optimizes the coordination of the overall topology structure of the digital human face and the smoothness of the movement trajectory. The optimization process can correct possible movement parameter inconsistencies and trajectory stalls in the initial binding data, allowing digital human facial components to work together during movement and avoiding movement disconnection or conflict between facial components. The optimized binding data can drive the digital human to maintain smooth movement during complex expression conversion, further enhancing the realism of digital human facial expressions. The facial expression of the finally presented digital human is not only close to the real human expression, but also has good dynamic performance effect by optimizing the facial motion of the data-driven digital human model. In the film and television production scene, the expression of the digital human role can be more lively, and the infectivity of the film and television work is enhanced; in the virtual live scene, the expression of the virtual host can follow the real host in real time, and the viewing experience of the audience is improved; in the interactive game scene, the game role can make natural expression response according to the expression feedback of the player, and the immersion of the player is enhanced. In addition, the whole method process does not need a large amount of manual intervention, from data acquisition, feature extraction to binding generation and optimization, a relatively automatic processing mechanism is formed, the artificial operation cost is effectively reduced, the generation efficiency of the digital human facial binding is improved, the demand of large-scale digital human production can be met, and the application and popularization of the digital human technology in more fields are promoted. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 The working principle diagram of the digital human facial binding generation method based on expression capture is described. Figure 2 The flowchart of the binding optimization process is described. Figure 3 The flowchart of the reconstructed motion weight distribution is described. Figure 4 The flowchart of the expression error triggered dynamic update is described. DETAILED DESCRIPTION

[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0017] Please refer to Figure 1 The present application provides a digital human facial binding generation method based on expression capture, which comprises the cooperative operation of an expression capture device, a data processing unit and a digital human model rendering engine.

[0018] The expression capture device adopts a high-precision optical marker point system or a depth camera array to collect facial muscle motion information at a rate of more than 60 frames per second. The data processing unit is equipped with a multi-thread feature extraction algorithm and a motion trajectory fitting module, and the digital human model rendering engine supports real-time skeletal driving and deformation grid processing.

[0019] The facial expression capture device collects real-time facial expression data of the target object in real time. The data stream is transmitted to the feature extraction module through a high-speed data interface, and the spatio-temporal features are extracted using a convolutional neural network architecture, outputting a set of 256-dimensional expression feature vectors. The trajectory fitting module uses a nonlinear regression algorithm to calculate the parameterized trajectory of facial muscle movement based on the expression feature vector set, generating an initial binding rule. The rule contains the motion equation of each control point of the face and the weight distribution parameter. The mapping engine associates the initial binding rule with the face topology of the digital human model, and generates the initial face binding data through the vertex mapping algorithm, which contains the vertex displacement matrix and the bone rotation amount. The binding optimization processor performs smoothness verification and physical constraint adjustment on the initial face binding data, and outputs the optimized face binding data. The rendering engine drives the face mesh deformation and bone rotation of the digital human model using the optimized face binding data, realizing real-time expression reproduction.

[0020] Embodiment 1: refer to Figure 2 The implementation of this embodiment relies on the cooperative work of a high-precision optical capture system, a data processing unit, and a three-dimensional modeling software. The optical capture system arranges eight infrared high-speed cameras around the target object, with a frame rate of 120 frames per second and a resolution of no less than 1280x1024 pixels. Forty-two reflective marker points are pasted on the face of the target object, which are distributed according to the facial muscle anatomy, focusing on covering the motion areas of major expression muscle groups such as orbicularis oculi, orbicularis oris, zygomaticus major, and frontalis. The three-dimensional coordinate data of each marker point is calculated in real time through multi-view camera triangulation, forming a continuous sequence of facial key point coordinates.

[0021] After the coordinate sequence is transmitted to the data processing unit, it is compared with the preset standard expression template. The standard expression template contains the reference coordinate values of the 42 marker points in the neutral expression state, which is obtained by averaging the facial data of the target object in the static state. The comparison process calculates the displacement of each marker point in three coordinate axes, generating an expression offset data set containing 126 displacement parameters. This data set is labeled as real-time facial expression data, with a time resolution consistent with the acquisition frame rate.

[0022] The feature extraction module performs dimensionality reduction on the real-time facial expression data. Principal component analysis is used to identify the main patterns of expression movement, extracting 20 main feature components from the original 126-dimensional data. These feature components correspond to independent motion patterns such as eyebrow raising, mouth stretching, and eye shrinking in different facial regions. Each feature component contains a set of feature vectors and corresponding variance contribution rates, which together constitute the expression feature vector set.

[0023] The facial motion trajectory fitting module receives the expression feature vector set and generates initial binding rules using time series analysis methods. The module first establishes the correspondence between the digital human facial control points and the real facial marker points. The digital human model has 126 pre-set control points, which are distributed in accordance with the real marker points. The motion trajectory function of each control point is calculated by a nonlinear regression algorithm, and the function parameters include displacement amplitude, motion speed, acceleration and other dynamic characteristics. At the same time, the motion correlation weight between each control point is calculated to form a complete initial binding rule.

[0024] The initial binding rule is mapped to the facial topology structure of the digital human model. The digital human model adopts a hierarchical architecture, including a skeleton layer, a muscle layer and an epidermis layer. The skeleton layer is composed of 86 bone nodes, each node has a rotation range and motion constraint. The mapping process first converts the control point motion trajectory into the rotation parameters of the bone nodes, and calculates the bone rotation angle by inverse kinematics algorithm. Then the muscle contraction parameters are calculated according to the physical characteristics of the muscle layer, and finally the motion is transmitted to the 3524 grid vertices of the epidermis layer through the vertex displacement mapping table to generate the initial facial binding data.

[0025] The binding optimization processing stage calls the facial skeleton constraint parameters of the digital human model. These parameters include the maximum rotation angle, motion damping coefficient, skeleton linkage relationship and other physical constraints of each bone node. The system calculates the matching degree of the initial facial binding data with these constraint parameters, and uses the constraint satisfaction algorithm to detect whether there is a bone motion that exceeds the allowed range. For each bone node, calculate the deviation value of its actual motion parameters and constraint parameters to generate a binding deviation data set containing 86 node deviation data. According to the binding deviation data, the system adjusts the initial binding rule through an iterative optimization algorithm. The adjustment process focuses on correcting the motion weight distribution parameters, reducing the motion influence weight of nodes with large deviation, and increasing the weight proportion of nodes that meet the constraints. The optimized binding rule recalculates the bone motion parameters and vertex displacement data to generate optimized facial binding data. The whole process maintains real-time running characteristics, and the data delay is controlled within 50 milliseconds, ensuring the real-time performance of expression driving.

[0026] All data processing modules use multi-thread parallel computing architecture. The expression capture device is connected to the data processing unit through a gigabit Ethernet interface, and the data transmission uses UDP protocol to ensure real-time performance. The data processing unit is equipped with a multi-core processor and a dedicated graphics computing card to ensure smooth operation of complex calculations. The digital human model rendering engine supports real-time skeletal animation and mesh deformation, and the rendering frame rate is synchronized with the acquisition frame rate.

[0027] The system realizes a complete pipeline from real expression capture to digital human driving, each link is precisely designed to ensure the accuracy and efficiency of data transmission. The expression capture accuracy reaches sub-millimeter level, and the natural degree of digital human expression reproduction is highly consistent with real expression. The whole system runs stably and can adapt to the capture needs of different expression intensity, from subtle expression changes to exaggerated expression actions can be accurately captured and reproduced.

[0028] Embodiment 2: see Figure 3 , focusing on the deep analysis and targeted optimization of binding deviation data, the core is to identify abnormal nodes in the motion transmission process, and based on the comprehensive evaluation of historical motion characteristics and current deviation, to realize the reconstruction of motion weight distribution. After processing the initial facial binding data, the system enters the binding optimization stage. The binding deviation data is first input into the analysis module, the task of this module is to extract the key motion axis deviation value. The identification of the key motion axis depends on the pre-definition of the importance of the digital human facial skeletal system function. For example, the zygomatic rotation axis dominates the mouth corner and cheek area movement of smile and chewing action, the mandibular opening and closing axis controls the opening and closing of the mouth, and the brow bone displacement axis directly affects the key features of the eyebrow raising and wrinkling expressions. The system calculates the projection deviation in each motion axis direction, that is, calculates the component size of the binding deviation data vector in each preset skeletal axis direction. These component values are marked as key motion axis deviation values, which quantify the degree of deviation of the current binding data from the ideal state in each main motion dimension.

[0029] These calculated key motion axis deviation values are compared one by one with the preset deviation threshold value. The preset deviation threshold value is not a fixed value, but an interval range dynamically adjusted according to the physical properties of the skeletal nodes, the elastic modulus simulation parameters of the muscles, and the historical motion performance. For example, the threshold range of the node controlling the mouth corner movement may be set narrower than that of the forehead node, because subtle mouth corner deviation is more visually obvious. The comparison process uses interval judgment logic, when the deviation value of a certain motion axis continuously and significantly falls outside its corresponding threshold interval, the system determines that the skeletal node associated with this motion axis is abnormal, and marks it as an abnormal binding node. The system records the unique identifier of the node, the motion axis it belongs to, and the specific value and direction of the deviation.

[0030] After locating the abnormal binding node, the system initiates the reconstruction process of its motion weight distribution. The first step of reconstruction is to collect the historical motion smoothness data of the node. The system accesses the motion trajectory database of the node within a time window, for example, 300 consecutive frames in the past. These data record the position sequence of the node in three-dimensional space. The motion trajectory variation is calculated by analyzing the position difference between consecutive frames, and then the velocity and acceleration sequence is obtained. The distribution dispersion is obtained by statistical analysis of the acceleration sequence, for example, calculating its standard deviation or root mean square value. The larger the value, the more unstable and jittering the historical motion, the worse the smoothness; on the contrary, it indicates that the motion history is smooth and stable. The quantitative index obtained is recorded as the historical motion smoothness data of the node.

[0031] The system fuses the collected historical motion smoothness data with the current key motion axis deviation value triggered by the abnormal alarm. Before that, the historical motion smoothness data needs to be normalized, which is linearly mapped to a pre-set smoothness reference range, so that data of different nodes and different dimensions are comparable. The current key motion axis deviation value is also standardized to eliminate the influence of the absolute value size in different axes. The fusion calculation adopts the weighted summation method, and the weight distribution strategy is pre-defined. A common strategy is to give a higher weight to the historical smoothness, emphasizing the node's consistent motion performance; at the same time, give the current deviation a proper weight, reflecting its immediate abnormality. The result of weighted summation is converted by a non-linear function (such as Sigmoid function), and finally an optimization coefficient of the node between 0 and 1 is output. The coefficient comprehensively reflects the adjustment intensity required by the node based on its historical performance and current situation. The higher the coefficient, the more the weight distribution of the node needs to be corrected.

[0032] According to the calculated node optimization coefficient, the system re-distributes the motion weight value of the abnormal binding node. The system maintains a pre-set weight mapping table inside, which essentially defines the functional relationship between the optimization coefficient and the specific weight adjustment amount. The mapping relationship is usually designed as: the higher the optimization coefficient, the larger the weight adjustment amount (usually the downward correction amplitude). The system obtains the target weight adjustment amount by looking up the table or calculating the function. Subsequently, the system updates the motion weight value of the abnormal binding node in the initial binding rule using the target weight adjustment amount. The update process needs to consider the weight sum constraint of the node and other associated nodes, so sometimes local range weight re-normalization is needed to ensure the coordination of the overall weight distribution. The updated weight value is written into the optimized binding rule, thereby completing the correction of the specific abnormal node. The whole process is iterated for each detected abnormal node until all significant deviations are processed, and a set of fully optimized binding rules is finally generated.

[0033] Example 3: Focuses on the quantitative analysis of historical motion data of anomalously bound nodes and the specific calculation method for fusing it with the current state. The core of this process lies in using precise mathematical processing to transform the node's past motion stability characteristics and immediate deviation information into a standardized coefficient that can guide weight adjustments. After locating an anomalously bound node, the system immediately accesses the node's motion history database. This database stores the node's three-dimensional coordinate sequence in the model's local coordinate system within the most recent N frames (e.g., N=300) using a circular buffer structure, denoted as... The subscript t represents the frame number. The calculation of the change in motion trajectory first obtains the instantaneous velocity sequence through first-order difference. Then, the instantaneous acceleration sequence is obtained through second-order difference. To comprehensively evaluate the stability of the motion, the system calculates the magnitude of the acceleration sequence. And analyze the statistical properties of this scalar sequence, including its distribution dispersion. This is obtained by calculating the standard deviation of the modulus sequence. Its value directly reflects the drastic nature of the historical acceleration changes at that node; a larger standard deviation indicates more abrupt and uneven motion. This dispersion value... After a logarithmic transformation function Processing (wherein) It is a very small constant (to prevent taking the logarithm with respect to zero), and the result is linearly mapped to a closed interval between zero and one, finally yielding normalized historical motion smoothness data. The mapping function is ,in and It is a normalized boundary value obtained statistically from a large amount of prior data, used to unify smoothness data of different nodes and different magnitudes to the same comparison benchmark. Thus, a standardized index describing the past motion stability of this node is formed. It is then generated.

[0034] At the same time, the system processes the current critical motion axis deviation values. This value is derived from the projection deviation along a specific motion axis calculated in Example 2. This is to ensure it is compatible with normalized smoothness data. Perform fusion, deviation value Standardization is also required. The system uses the Z-score standardization method, i.e. ,in and It is the deviation value of the abnormally bound node in its normal motion history. the mean and standard deviation of the historical motion smoothness data. This step converts the current bias value into a scalar with a mean of zero and a standard deviation of one, eliminating the influence of absolute value differences between different motion axes, and making it a relative indicator of the current bias relative to its own historical level.

[0035] Subsequently, the system will fuse the normalized historical motion smoothness data with the normalized current key motion axis bias value to obtain the final node optimization coefficient . The fusion is performed in the form of weighted summation, and its calculation relationship is defined by the following formula:

[0036] Wherein: represents the final calculated node optimization coefficient, whose value range is compressed between zero and one. represents the normalized historical motion smoothness data calculated in the previous step. represents the current key motion axis bias value after Z-score standardization. is the fusion weight given to the historical smoothness data, which is a preset constant greater than zero, and its value determines the influence of historical motion performance on the final optimization coefficient. is the fusion weight given to the current bias value, which is also a preset constant greater than zero, and its value determines the influence of the immediate abnormal situation on the final optimization coefficient. is a bias term used to adjust the offset of the function. represents the natural exponential function. The role of this Sigmoid function is to map the result of weighted summation nonlinearly to the range of zero to one. When the weighted sum result is very large, it tends to one infinitely, indicating that strong correction is needed; when the weighted sum result is very small, it tends to zero infinitely, indicating that only slight correction or no correction is needed; in the intermediate stage, the change of provides a smooth transition. The specific values of weights and are usually set according to experience during system initialization, for example, they can be set as , , which means that the historical motion smoothness occupies a more important position than the current instantaneous bias in evaluation, emphasizing the consistent motion characteristics of the node. The calculated node optimization coefficient As a comprehensive quantitative output, it will be directly passed to the weight adjustment module in Example 4 as the core input parameter of the query target weight adjustment amount, thus completing the closed loop from data analysis to parameter adjustment. The whole calculation process emphasizes the balanced consideration of historical behavior and current state, aiming to generate an optimization instruction that respects both the long-term movement rules of the node and effectively corrects short-term deviations.

[0037] Example 4: Involves converting the calculated node optimization coefficient into actual weight adjustment operations, and on this basis, establishing a continuous face movement deformation monitoring and system feedback mechanism. The core of this process is to use a predefined mapping relationship to convert abstract analysis results into specific parameter adjustment instructions, and through real-time monitoring to ensure the long-term stability and accuracy of digital human facial movement. After the system obtains the optimization coefficient κ describing the adjustment strength required for the abnormal binding node, it needs to be converted into an executable motion weight adjustment amount. This conversion is completed by querying a pre-defined weight mapping table. This mapping table is a static data structure, usually loaded into memory from a configuration file during system initialization. Its essence is a discrete function lookup table that defines the correspondence between optimization coefficients and weight adjustment amounts. The construction of this table is based on a large amount of prior experimental data, and its design principle is: the lower the optimization coefficient, the smoother the historical movement of the node, the smaller the current deviation, and therefore the smaller the weight adjustment amplitude required, or even no adjustment or slight enhancement; the higher the optimization coefficient, the more likely the historical movement of the node to have jitter and the current deviation to be significant, thus requiring more intensive downward correction of its weight to suppress its excessive movement influence.

[0038] The weight mapping table contains several records, each consisting of two fields: optimization coefficient interval and corresponding target weight adjustment amount. The system uses the nearest neighbor interpolation algorithm for querying. It compares the calculated precise optimization coefficient value with the median value of each optimization coefficient interval in the mapping table, and selects the adjustment amount corresponding to the nearest interval as the final target weight adjustment amount. This process ensures that even if the coefficient value does not exactly match the endpoint of a certain interval, a most reasonable approximate adjustment value can be found. The target weight adjustment amount is a signed floating-point number, with the sign indicating the adjustment direction (positive for increasing weight, negative for decreasing weight) and the absolute value indicating the adjustment amplitude.

[0039] After obtaining the target weight adjustment amount, the system applies it to the current motion weight value of the abnormal binding node. The update operation is a direct arithmetic addition: new weight value = original weight value + target weight adjustment amount. However, since there are mutual correlations and constraints between the weights of multiple nodes in the digital human facial binding system (for example, the total sum of the weights of multiple nodes that control the same facial area needs to be maintained constant), the change in the weight of a single node may disrupt this balance. To solve this problem, after updating the weights of all marked abnormal nodes, the system starts a local weight renormalization process. This process recalculates the weights of a group of nodes that have a linkage relationship with the abnormal node, so that the total sum of the weights of the group of nodes remains unchanged, while maintaining the original weight proportion relationship between the nodes as much as possible. In this way, the parameters of the abnormal node are corrected, and the coordination and stability of the overall weight distribution are maintained. All updated weight values are written into the new binding rule data structure, thereby generating the optimized binding rule.

[0040] After driving the digital human model using the optimized facial binding data, the system does not terminate work, but starts a parallel real-time monitoring process. The deformation monitoring module continuously analyzes the vertex data of the digital human model's facial mesh. It calculates the three-dimensional displacement vector of each vertex by comparing the driven vertex position with the reference position in the neutral state. The module does not focus on all vertices, but focuses on a set of pre-defined key deformation feature points, which are usually located in areas with the most expressive changes, such as eye corners, mouth corners, and eyebrow tips.

[0041] The system sets a preset deformation tolerance range for each key deformation feature point. This is a three-dimensional boundary value that defines the maximum displacement allowed for the vertex on the X, Y, and Z axes. The size of the tolerance range is differentiated according to the expression sensitivity and visual importance of the area where the vertex is located. The monitoring logic is: if the displacement of any key feature point exceeds the tolerance range set for it on any axis, it is considered a deformation overrun event.

[0042] Once a deformation overrun event is detected, the system automatically triggers a process reset. It keeps all current parameter configurations intact but reinitializes the data acquisition process. This means the expression capture device is instructed to start a new round of real-time facial expression data acquisition immediately, and the newly collected data will go through the entire processing pipeline again, experiencing feature extraction, track fitting, initial binding data generation from mapping, and optimization based on the current optimized binding rules. The purpose is to use the latest facial expression input data to generate an updated facial binding data based on the current optimization rules, in order to correct the observed excessive deformation. This is a closed-loop control strategy that adjusts the input process through the output result (facial deformation) feedback, thereby maintaining the stability and authenticity of the system output, see Table 1.

[0043] Table 1: Optimization coefficient and weight adjustment mapping table.

[0044]

[0045] The entire implementation process embodies a complete closed loop from parameter optimization to effect verification to feedback adjustment. The weight mapping table provides a bridge from analysis to execution, while real-time deformation monitoring ensures that the system can dynamically adapt to various situations, maintaining the high quality and stability of digital human facial expression performance. Through this mechanism, the system can to some extent self-correct the misfit problems caused by model errors, capture noise or sudden large-scale expressions.

[0046] Example 5: see Figure 4 , constitutes the closed-loop feedback and dynamic calibration mechanism of the entire facial binding system, the core of which is the continuous and fine evaluation of the driving results and the systematic self-adjustment based on the evaluation results. This mechanism ensures that the facial expression output of the digital human can maintain high consistency with the real capture source for a long time, and can adapt to the subtle changes of the capture environment or facial movement. While optimizing the facial binding data-driven digital human model for real-time facial movement, a high-precision deformation monitoring and analysis system starts to operate in parallel. This system continuously captures the vertex motion data of the digital human face mesh at a sampling frequency not lower than the driving frame rate. Its focus is not on all mesh vertices, but on a set of key deformation feature points defined in advance according to facial anatomy and expressionology. These points are usually located in areas where muscle traction effect is most significant and visual performance is most sensitive, such as eyebrow arch vertex, eye tail split point, nasal labial sulcus inflection point, corner of the mouth, and mandibular protuberance point, etc. For each feature point, the monitoring system records its three-dimensional displacement trajectory relative to the neutral expression reference position, forming a set of deformation data sets that change continuously over time.

[0047] The system starts the morphing feature analysis process. This process aims to extract higher-level, more semantic expression motion pattern information from the massive underlying vertex displacement data. The analysis process uses dimension reduction and feature extraction techniques such as principal component analysis to map the high-dimensional vertex displacement vector space to a low-dimensional feature space. Each principal dimension of this low-dimensional space usually corresponds to a basic facial motion pattern, such as "eyebrow overall lifting", "corners of the mouth stretching horizontally" or "nasal alar opening" and so on. By calculating the projection coefficients of the current morphing data on these principal dimensions, the system generates a set of compact and expressive expression error data. This set of data is no longer a measure of geometric displacement, but is transformed into a quantitative deviation of the current driven expression from the ideal expression in multiple motion patterns.

[0048] The generated expression error data is immediately sent to a threshold comparison module, which stores a set of pre-set expression tolerance threshold vectors. This set of thresholds is not a single global numerical value, but a set of tolerance ranges corresponding to the dimensions of the expression error data, each dimension has its own independent tolerance upper limit. The tolerance value of each dimension is obtained by statistical analysis of a large number of high-quality expression data samples, reflecting the normal deviation range of the specific motion pattern in natural expression. The comparison process is carried out dimension by dimension, and the system strictly compares each component of the error data with its corresponding pre-set tolerance threshold.

[0049] The comparison result directly determines the subsequent behavior of the system. If the error values of all dimensions fall within their corresponding tolerance thresholds, the system determines that the current facial binding state is good and does not need to be intervened, and continues to perform the monitoring task. However, if the expression error data of any one or more dimensions exceeds the pre-set expression tolerance threshold set for it, a flag control signal is immediately triggered - the dynamic update process of the optimized facial binding data.

[0050] The start of the dynamic update process means that the system switches from a passive monitoring state to an active calibration state. The core of this process is incremental update and minimum disturbance. The system does not reset or discard the existing optimized binding rules, but uses them as the basis for the new round of calculations. The system instructs the expression capture device to immediately provide the latest real-time facial expression data stream, which reflects the current most realistic expression state of the target object. These newly collected data are directly sent to the initial facial binding data generation module, which uses the existing, previously optimized and adjusted binding rules as parameters to process the fresh input data and quickly generate a new set of initial facial binding data.

[0051] The new initial facial rigging data then undergoes the same optimization process as before, but this time the optimization is based on the current latest facial expression input and the existing optimization rules, aiming to fine-tune the rigging parameters to correct the monitored expression errors. Finally, an updated optimized facial rigging data is output. This new data is seamlessly switched to the driving engine, replacing the old data set, thus driving the digital human model to produce more accurate expression performance. The entire dynamic updating process, from triggering to application completion, is designed to be completed in a very short time, aiming to achieve fast online correction of expression deviations, minimize interference with user experience, and maintain the coherence and realism of digital human facial animation. This mechanism makes the system adaptive and robust, effectively dealing with situations such as capture device drift, environmental light changes, or subtle changes in user facial expression style that may occur during long-term operation.

[0052] It should be noted that the relational terms herein such as first and second, and the like, are used solely to distinguish one from another entity or action, without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0053] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A method for generating digital human facial bindings based on expression capture, characterized in that, include: Real-time facial expression data of the target object is collected using facial expression capture equipment; Feature extraction is performed on the real-time facial expression data to obtain a set of facial expression feature vectors; Facial motion trajectories are fitted based on the set of facial expression feature vectors to generate initial binding rules; Initial facial binding data is generated by mapping the initial binding rules to the facial topology of the digital human model. The initial facial binding data is subjected to binding optimization processing to generate optimized facial binding data; The optimized facial binding data is used to drive facial movements in a digital human model.

2. The digital human face binding generation method based on expression capture according to claim 1, characterized in that, Real-time facial expression data of the target subject is collected using facial expression capture devices, including: The facial key point coordinate sequence of the target object is obtained through the facial expression capture device; The facial key point coordinate sequence is compared with a preset standard expression template to obtain expression offset data; The expression offset data is labeled as real-time facial expression data.

3. The digital human face binding generation method based on expression capture according to claim 2, characterized in that, The initial facial binding data is subjected to binding optimization processing to generate optimized facial binding data, including: Call the facial bone constraint parameters of the digital human model; The matching degree between the initial facial binding data and the facial bone constraint parameters is calculated to obtain binding deviation data; The initial binding rules are adjusted based on the binding deviation data to generate optimized binding rules; The initial facial binding data is updated based on the optimized binding rules to generate optimized facial binding data.

4. The digital human face binding generation method based on expression capture according to claim 3, characterized in that, Adjusting the initial binding rules based on the binding deviation data to generate optimized binding rules includes: Extract the key motion axis deviation values ​​from the binding deviation data; The deviation value of the key motion axis is compared with a preset deviation threshold to locate abnormal binding nodes; The motion weight allocation of the abnormally bound nodes is reconstructed to generate optimized binding rules.

5. The digital human face binding generation method based on expression capture according to claim 4, characterized in that, Reconstructing the motion weight allocation of the abnormally bound nodes to generate optimized binding rules includes: Collect historical motion smoothness data of the abnormally bound nodes; The historical motion smoothness data is fused with the current critical motion axis deviation value to obtain the node optimization coefficient; The motion weight values ​​of the abnormally bound nodes are reassigned based on the node optimization coefficients to generate optimized binding rules.

6. The digital human face binding generation method based on expression capture according to claim 5, characterized in that, Collect historical motion smoothness data of the abnormally bound nodes, including: Obtain the change in the motion trajectory of the abnormally bound node in consecutive time frames; The dispersion of the change in the motion trajectory is calculated to obtain historical motion smoothness data.

7. The digital human face binding generation method based on expression capture according to claim 6, characterized in that, The historical motion smoothness data is fused with the current critical motion axis deviation value to obtain the node optimization coefficient, including: The historical motion smoothness data is normalized to a preset smoothness reference range; The normalized historical motion smoothness data is weighted and summed with the current critical motion axis deviation value to obtain the node optimization coefficient.

8. The digital human face binding generation method based on expression capture according to claim 7, characterized in that, Based on the node optimization coefficients, the motion weight values ​​of the abnormally bound nodes are reassigned to generate optimized binding rules, including: The node optimization coefficients are input into a preset weight mapping table to match the target weight adjustment amount; The initial motion weight value of the abnormally bound node is updated using the target weight adjustment amount to generate an optimized binding rule.

9. The digital human face binding generation method based on expression capture according to claim 8, characterized in that, After using the optimized facial rigging data to drive facial movements in a digital human model, the process includes: Real-time monitoring of facial motion deformation data of the digital human model; If the facial motion deformation data exceeds the preset deformation tolerance range, the step of collecting real-time facial expression data of the target object through the expression capture device is repeated.

10. The digital human face binding generation method based on expression capture according to claim 9, characterized in that, After real-time monitoring of the facial motion deformation data of the digital human model, the process includes: Deformation feature analysis is performed on the facial motion deformation data to generate expression error data; The facial expression error data is compared with a preset facial expression tolerance threshold; If the expression error data exceeds the preset expression tolerance threshold, the dynamic update of the optimized facial binding data is triggered, and the step of generating the initial facial binding data is re-executed based on the latest collected real-time facial expression data.

Citation Information

Patent Citations

  • Method and device for generating expression animation

    CN114219880A

  • Face binding method and device, equipment and storage medium

    CN115393532A

  • Virtual image expression generation method and system based on real feeling technology

    CN119295683A

  • Virtual human design and application platform and method based on artificial intelligence, equipment and medium

    CN120339470A

  • Digital human generation method based on multi-modal large model

    CN120543710A