A method and system for dynamic reconstruction of craniofacial based on multi-modal data fusion

By using a multimodal data fusion method, combining static images and electromyographic signals, high-precision dynamic craniofacial reconstruction was achieved, solving the problems of low model accuracy and low computational efficiency in existing technologies. This generated a dynamic craniofacial model that conforms to physiological characteristics and is suitable for multiple application scenarios.

CN120236011BActive Publication Date: 2025-11-21青峰宇
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510322526.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-11-21
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

In existing technologies, craniofacial dynamic reconstruction methods rely on single-modal data, which makes it difficult to accurately model facial movements and biomechanical mechanisms. This results in models lacking physical realism and individual adaptability, and also incurring high computational costs, making it difficult to achieve high-precision dynamic reconstruction efficiently.

Method used

A high-precision craniofacial dynamic model was generated by employing a multimodal data fusion method, combining static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals, through rigid and non-rigid registration, optical flow analysis, muscle activation feature extraction, and biomechanical driving strategies.

Benefits of technology

It improves the accuracy and robustness of craniofacial dynamic modeling, generates dynamic craniofacial models that are more physiologically consistent, and provides objective model evaluation criteria, which are suitable for applications such as medical surgical planning, facial expression disorder assessment, and virtual human construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236011B_ABST
    Figure CN120236011B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical image processing, and discloses a craniomaxillary dynamic reconstruction method and system based on multi-modal data fusion, which comprises the following steps: arranging multi-modal data acquisition devices in a target craniomaxillary region, acquiring static CT images, static MRI images, dynamic expression video sequences and surface electromyography signals; preprocessing the static CT images, the static MRI images, the dynamic expression video sequences and the surface electromyography signals; inputting the preprocessed data into a multi-scale finite element model and simulating the coupling relationship between muscle contraction force and skin deformation by adopting a biomechanical driving strategy to generate a dynamic craniomaxillary model; and fusing the geometric error and motion consistency score of real data by adopting a linear regression method to output a comprehensive reconstruction quality index. The application can solve the problems of low dynamic modeling accuracy and poor robustness in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, specifically to a method and system for dynamic craniofacial reconstruction based on multimodal data fusion. Background Technology

[0002] Dynamic reconstruction of the craniofacial region has significant applications in facial expression analysis, personalized surgical planning, medical imaging research, and virtual human modeling. Existing methods primarily rely on single-modal data, such as static CT or MRI images, which only provide anatomical information of craniofacial tissues and lack accurate modeling of facial movements. Furthermore, while video-based facial reconstruction methods can acquire information on dynamic facial deformation, they struggle to deeply analyze biomechanical mechanisms such as muscle contraction and skeletal support, resulting in models lacking physical realism and individual adaptability.

[0003] In recent years, multimodal craniofacial modeling methods have faced numerous challenges, including: Registration and fusion of multimodal images: CT and MRI images have different characteristics, and achieving high-precision alignment and integration of skeletal and soft tissue information remains a challenge. Representation of facial dynamics: Efficiently extracting the motion trajectories of key facial points and combining them with deep learning methods to enhance motion feature representation is crucial for improving modeling accuracy. Modeling of muscle activation features: Electromyographic signals are subject to noise interference; accurately extracting muscle activation features and reasonably mapping them into the finite element model is a core issue in achieving biomechanical drive. Physical simulation and computational efficiency: Traditional finite element methods have high computational costs; improving computational efficiency while ensuring simulation accuracy is a challenge for dynamic reconstruction.

[0004] To address the aforementioned issues, this invention proposes a craniofacial dynamic reconstruction method based on multimodal data fusion. This method comprehensively utilizes static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals, combined with a biomechanical driving strategy, to construct a high-precision craniofacial dynamic model, thereby improving the accuracy and robustness of modeling.

[0005] Therefore, there is an urgent need for a method to construct a high-precision craniofacial dynamic model by utilizing static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals, combined with biomechanical driving strategies, in order to improve the accuracy and robustness of modeling. Summary of the Invention

[0006] In view of this, the present invention proposes a method and system for craniofacial dynamic reconstruction based on multimodal data fusion, aiming to solve the problems of low accuracy and poor robustness of current craniofacial dynamic modeling.

[0007] This invention proposes a craniofacial dynamic reconstruction method based on multimodal data fusion, comprising:

[0008] Multimodal data acquisition devices were deployed in the target craniofacial region to acquire static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals.

[0009] Rigid and non-rigid registration are performed on static CT images and static MRI images to generate a fused static three-dimensional anatomical model;

[0010] Optical flow analysis and keyframe extraction are performed on dynamic facial expression video sequences to obtain facial motion trajectory data;

[0011] Filtering and time-frequency analysis of surface electromyography signals were performed to extract muscle activation features;

[0012] The fused static 3D anatomical model, facial motion trajectory data, and muscle activation features are input into a multi-scale finite element model, and a biomechanical driving strategy is used to simulate the coupling relationship between muscle contraction force and skin deformation to generate a dynamic craniofacial model.

[0013] A linear regression method is used to fuse geometric error and motion consistency scores from real data to output a comprehensive reconstruction quality index.

[0014] Furthermore, the multimodal data acquisition device includes a CT scanner, an MRI device, a high-speed camera, and a wireless surface electromyography sensor array.

[0015] Furthermore, the specific content of performing rigid and non-rigid registration on static CT images and static MRI images to generate a fused static three-dimensional anatomical model is as follows: the rigid registration algorithm based on mutual information is used to align the skeletal structure on the static CT images and static MRI images to obtain skeletal details;

[0016] A non-rigid registration algorithm based on B-spline was used to align soft tissue contours in static CT and static MRI images to obtain soft tissue information.

[0017] By integrating skeletal details and soft tissue information through a weighted fusion strategy, a high-precision static 3D model is generated.

[0018] Furthermore, the specific content of performing optical flow analysis and keyframe extraction on the dynamic expression video sequence to obtain facial motion trajectory data is as follows: the displacement of facial key points is extracted from the dynamic expression video sequence using an improved Lucas-Kanade optical flow method to generate original motion trajectory data. Then, a time series clustering algorithm is used to filter keyframes in the original motion trajectory data to construct a sparse representation of the facial motion trajectory. Finally, facial texture features in the dynamic expression video sequence are extracted through a convolutional neural network and fused with the sparse representation of the facial motion trajectory to obtain facial motion trajectory data.

[0019] Furthermore, the facial expression texture features in the dynamic facial expression video sequence are extracted by a convolutional neural network and fused with the sparse representation of facial motion trajectory. The specific content of the facial motion trajectory data is as follows: optical flow timestamps are used to align the sparse representation of facial motion trajectory with the facial expression texture features in time and space. Then, the weights are dynamically adjusted by the cross attention module in the convolutional neural network to enhance the modeling strength. Finally, the sparse representation of facial motion trajectory and facial expression texture features are concatenated into a joint feature vector to obtain the facial motion trajectory data.

[0020] Furthermore, the specific steps for filtering and time-frequency analysis of the surface electromyography (EMG) signal to extract muscle activation features are as follows: First, wavelet packet transform is used to remove power frequency interference and motion artifacts from the EMG signal; then, independent component analysis is used to separate the target muscle group signal from the EMG signal; finally, time-domain and frequency-domain features are extracted to obtain muscle activation features.

[0021] Furthermore, the specific content of inputting the fused static three-dimensional anatomical model, facial motion trajectory data, and muscle activation features into a multi-scale finite element model and simulating the coupling relationship between muscle contraction force and skin deformation using a biomechanical driving strategy to generate a dynamic craniofacial model is as follows: the static three-dimensional anatomical model is discretized into nonlinear hyperelastic soft tissue elements, a stiffness matrix K is constructed, muscle activation features are converted into muscle contraction force through a muscle force model, facial motion trajectory data is used as boundary conditions, the Lagrange multiplier method is used to embed the dynamic equations, the Newmark-β method is used to solve the dynamic equations, the nodal displacement u is obtained by balancing computational efficiency and stability, the obtained nodal displacement u is mapped to the skin surface, and a dynamic craniofacial model is generated through radial basis function interpolation.

[0022] Furthermore, the specific content of the method of using linear regression to fuse the geometric error and motion consistency scores of real data and outputting the comprehensive reconstruction quality index is as follows: the Hausdorff distance between the vertices of the static three-dimensional anatomical model and the real anatomical landmarks is calculated as the geometric error score, the facial motion trajectory data consistency score is quantified by the dynamic time warping algorithm, and the geometric error score and motion consistency score are fused by linear regression to generate the reconstruction quality index.

[0023] Furthermore, the linear regression method is used to fuse geometric error scores and motion consistency scores to generate a reconstruction quality index, expressed as:

[0024] Q = α·S geo +β·S motion ;

[0025] Where Q is the reconstruction quality index, α is the first weighting coefficient, β is the second weighting coefficient, and S... geo S is used to score the geometric error.motion Scoring for motor consistency.

[0026] On the other hand, the present invention also proposes a craniofacial dynamic reconstruction system based on multimodal data fusion, the system comprising:

[0027] The data acquisition module is configured to acquire static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals;

[0028] The data processing module is configured to perform rigid and non-rigid registration on static CT and static MRI images to generate a fused static three-dimensional anatomical model; perform optical flow analysis and keyframe extraction on dynamic facial expression video sequences to obtain facial motion trajectory data; and perform filtering and time-frequency analysis on surface electromyography signals to extract muscle activation features.

[0029] The dynamic modeling module is configured to simulate the coupling relationship between muscle contraction force and skin deformation by integrating the fused static 3D anatomical model, facial motion trajectory data, and muscle activation features to generate a dynamic craniofacial model.

[0030] The evaluation output module is configured to use linear regression to fuse geometric error and motion consistency scores from real data, and output a comprehensive reconstruction quality index.

[0031] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a method and system for dynamic craniofacial reconstruction based on multimodal data fusion. It employs rigid and non-rigid registration of static CT and MRI images, aligns skeletal structures based on mutual information algorithms, and aligns soft tissues based on B-spline algorithms, simultaneously preserving skeletal details and soft tissue information, thus improving the accuracy of static three-dimensional anatomical models. This invention uses an improved Lucas-Kanade optical flow method to extract facial keypoint displacements, combines this with a time-series clustering algorithm to extract keyframes, and further utilizes convolutional neural networks to fuse facial texture features, achieving accurate representation of facial motion trajectories. This invention removes power frequency interference and motion artifacts from electromyographic signals through wavelet packet transform, uses independent component analysis (ICA) to separate target muscle group signals, and extracts time-domain and frequency-domain features to accurately obtain muscle activation features, providing high-quality input data for biomechanical drives. This invention adopts… This invention discretizes facial tissue using a nonlinear hyperelastic finite element model, constructs a stiffness matrix K, converts muscle activation features into muscle contraction forces through a muscle force model, and combines facial motion trajectory data as boundary conditions. It embeds the dynamic equations using the Lagrange multiplier method and solves them using the Newmark-β method, balancing computational efficiency and stability to generate a more physiologically accurate dynamic craniofacial model. The invention calculates the geometric error score between the vertices of the static 3D anatomical model and real anatomical landmarks using Hausdorff distance, quantifies the consistency score of facial motion trajectories using a dynamic time warping algorithm, and fuses these two indicators using linear regression to generate a comprehensive reconstruction quality index, thus providing an objective model evaluation standard. This invention can be applied to multiple scenarios such as medical surgical planning, facial expression disorder assessment, personalized prosthesis design, and virtual human construction, providing a precise craniofacial dynamic modeling solution for the medical and computer graphics fields. Attached Figure Description

[0032] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0033] Figure 1 This is a flowchart of a craniofacial dynamic reconstruction method based on multimodal data fusion according to an embodiment of the present invention;

[0034] Figure 2 This is a structural block diagram of a craniofacial dynamic reconstruction system based on multimodal data fusion, according to an embodiment of the present invention. Detailed Implementation

[0035] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0036] See Figure 1 As shown, this embodiment of the invention provides a method for dynamic craniofacial reconstruction based on multimodal data fusion, including:

[0037] S1: Deploy multimodal data acquisition devices in the target craniofacial region to acquire static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals;

[0038] S2: Perform rigid and non-rigid registration on static CT and static MRI images to generate a fused static three-dimensional anatomical model;

[0039] S3: Perform optical flow analysis and keyframe extraction on dynamic facial expression video sequences to obtain facial motion trajectory data;

[0040] S4: Filter and perform time-frequency analysis on the surface electromyography signal to extract muscle activation features;

[0041] S5: Input the fused static 3D anatomical model, facial motion trajectory data and muscle activation features into the multi-scale finite element model and use a biomechanical driving strategy to simulate the coupling relationship between muscle contraction force and skin deformation to generate a dynamic craniofacial model.

[0042] S6: The geometric error and motion consistency scores of the real data are fused using linear regression to output a comprehensive reconstruction quality index.

[0043] Furthermore, the multimodal data acquisition device includes a CT scanner, MRI equipment, a high-speed camera, and a wireless surface electromyography sensor array.

[0044] Furthermore, the specific content of rigid and non-rigid registration of static CT images and static MRI images to generate a fused static three-dimensional anatomical model is as follows: the rigid registration algorithm based on mutual information is used to align the skeletal structure of static CT images and static MRI images to obtain skeletal details;

[0045] A non-rigid registration algorithm based on B-spline was used to align soft tissue contours in static CT and static MRI images to obtain soft tissue information.

[0046] By integrating skeletal details and soft tissue information through a weighted fusion strategy, a high-precision static 3D model is generated.

[0047] Specifically, firstly, static CT and static MRI images are preprocessed, including grayscale normalization, histogram matching, and noise suppression, to reduce contrast differences between different imaging modalities and improve the robustness and accuracy of subsequent registration.

[0048] Next, a rigid registration algorithm based on mutual information (MI) is employed to align the skeletal structures in CT and MRI images. This method, based on information theory principles, achieves image alignment by maximizing the mutual information between CT and MRI images, thus overcoming the limitations of traditional gray-scale difference-based registration methods on multimodal images. During rigid registration, a gradient descent optimization strategy is used to iteratively adjust the rotation and translation parameters to ensure alignment of skeletal details between CT and MRI images, thereby improving registration accuracy and obtaining clear skeletal structural information.

[0049] Then, for soft tissue information, a non-rigid registration algorithm based on B-spline is employed. Since CT images primarily highlight skeletal information, while MRI images offer higher resolution for soft tissues, non-rigid registration is necessary to accurately match soft tissue morphology after aligning the skeletal structure. The B-spline deformation model constructs a low-dimensional control point mesh and freely deforms the image to align the soft tissue contours, compensating for individual differences in facial anatomy. This method utilizes a multi-scale optimization strategy, first performing large-scale global alignment, then progressively optimizing local details to improve the precision of soft tissue registration and ensure the integrity of soft tissue contour information.

[0050] Finally, a weighted fusion strategy was employed to integrate skeletal details and soft tissue information, generating a high-precision static 3D anatomical model. This primarily utilized a Gaussian kernel-based weighting method, assigning higher weights to CT images in skeletal regions and higher weights to MRI images in soft tissue regions. This ensured that the fused model maintained clear geometric details in the skeletal structure while exhibiting more realistic anatomical morphology in the soft tissue regions.

[0051] It should be noted that rigid registration ensures the consistency of skeletal structure, non-rigid registration compensates for individual differences in soft tissues, and weighted fusion optimizes information integration, making the overall registration result more accurate. Combining the high-resolution skeletal information from CT and the soft tissue details from MRI, the fusion model can accurately represent the structure of hard tissues (such as the skull) and clearly present the morphology of soft tissues (such as muscles and fat).

[0052] Furthermore, optical flow analysis and keyframe extraction are performed on the dynamic expression video sequence to obtain the specific content of facial motion trajectory data. The modified Lucas-Kanade optical flow method is used to extract the displacement of facial key points from the dynamic expression video sequence to generate the original motion trajectory data. Then, a time series clustering algorithm is used to filter the keyframes in the original motion trajectory data to construct a sparse representation of the facial motion trajectory. Finally, the facial texture features in the dynamic expression video sequence are extracted through a convolutional neural network and fused with the sparse representation of the facial motion trajectory to obtain the facial motion trajectory data.

[0053] Specifically, firstly, the dynamic expression video sequence is preprocessed, including denoising, grayscale normalization, and face region alignment, to reduce the impact of illumination changes and camera angle on optical flow calculation and improve the robustness of key point detection.

[0054] Next, an improved Lucas-Kanade optical flow method is used to extract displacement information of facial key points and generate raw motion trajectory data. The Lucas-Kanade optical flow method, based on the assumption of gray-level consistency within local windows, can effectively track the motion trajectory of facial feature points. However, traditional methods are prone to failure under large displacements or rapid movements; therefore, a multi-scale pyramid structure is introduced to enhance the stability of optical flow estimation. Simultaneously, regularization constraints and edge-preserving strategies are combined to avoid drift problems caused by occlusion or noise of facial key points, improving the continuity and accuracy of the motion trajectory.

[0055] Then, time series analysis was performed on the original motion trajectory data, and a keyframe selection method based on time series clustering was adopted. Since there are many redundant frames during facial expression changes, directly processing the complete motion trajectory would increase computational cost and potentially introduce noise. Therefore, the motion amplitude and acceleration changes of key points in each frame were first calculated. Similar trajectories were clustered using the Dynamic Time Warping (DTW) method, and representative keyframes were selected using the K-means clustering algorithm to construct a sparse representation of the facial motion trajectory. This process preserves the main features of expression changes while reducing data dimensionality and improving the computational efficiency of subsequent analysis.

[0056] Subsequently, facial expression texture features are extracted from dynamic facial expression video sequences using a convolutional neural network (CNN). Facial expressions involve not only geometric deformations (such as keypoint displacement) but also complex texture information (such as wrinkles and shadow variations). Therefore, deep convolutional networks such as ResNet or EfficientNet are employed to perform multi-scale feature extraction on video frames, capturing facial expression details at different levels. Furthermore, a self-attention mechanism is incorporated to enhance the weights of key regions (such as eyebrows, corners of the mouth, and the area around the eyes) during feature fusion, thereby improving the expressive power of facial expression texture information.

[0057] Finally, the sparse representation of facial motion trajectories and facial expression texture features are fused to obtain the final facial motion trajectory data. Specifically, a cross-attention mechanism is used to dynamically weight the data from the two modalities to ensure the synergistic effect of geometric motion features and facial expression texture features. Finally, a feature concatenation strategy is used to combine the motion trajectory and texture information into a unified feature vector, providing high-quality input data for subsequent craniofacial dynamic modeling.

[0058] Furthermore, the facial expression texture features in the dynamic facial expression video sequence are extracted by a convolutional neural network and fused with the sparse representation of facial motion trajectory. The specific content of the facial motion trajectory data is as follows: optical flow timestamps are used to align the sparse representation of facial motion trajectory with the facial expression texture features in time and space. Then, the weights are dynamically adjusted by the cross attention module in the convolutional neural network to enhance the modeling strength. Finally, the sparse representation of facial motion trajectory and facial expression texture features are concatenated into a joint feature vector to obtain the facial motion trajectory data.

[0059] Furthermore, the surface electromyography (EMG) signal is filtered and subjected to time-frequency analysis to extract muscle activation features. The specific steps are as follows: First, wavelet packet transform is used to remove power frequency interference and motion artifacts from the surface EMG signal; then, independent component analysis (ICA) is used to separate the target muscle group signal from the surface EMG signal; finally, time-domain and frequency-domain features are extracted to obtain muscle activation features.

[0060] Furthermore, the fused static 3D anatomical model, facial motion trajectory data, and muscle activation features are input into a multi-scale finite element model, and a biomechanical driving strategy is used to simulate the coupling relationship between muscle contraction force and skin deformation to generate a dynamic craniofacial model. The specific content is as follows: the static 3D anatomical model is discretized into nonlinear hyperelastic soft tissue elements, a stiffness matrix K is constructed, muscle activation features are converted into muscle contraction force through a muscle force model, facial motion trajectory data is used as boundary conditions, the Lagrange multiplier method is used to embed it into the dynamic equation, the Newmark-β method is used to solve the dynamic equation, the nodal displacement u is obtained by balancing computational efficiency and stability, the obtained nodal displacement u is mapped to the skin surface, and a dynamic craniofacial model is generated by radial basis function interpolation.

[0061] Specifically, firstly, the static three-dimensional anatomical model is discretized using the finite element method. An adaptive meshing strategy is employed to divide the craniofacial region into finite element meshes of different scales. Larger elements are used for the skeletal region to improve computational efficiency, while smaller tetrahedral or hexahedral elements are used for the soft tissue region to ensure accurate simulation of complex geometries. For the soft tissue portion, a nonlinear hyperelastic constitutive model (such as the Mooney-Rivlin or Neo-Hookean model) is introduced to describe the stress-strain relationship of the skin and muscle tissue during deformation, and the stiffness matrix K of the overall system is constructed.

[0062] Then, a biomechanical muscle force model is constructed, converting the muscle activation features extracted from surface electromyography signals into the corresponding muscle contractile force. Specifically, based on the Hill muscle model or a finite element-based active muscle force model, the muscle contractile force Fm is calculated and applied to the corresponding finite element elements. This model can consider the nonlinear stress-strain characteristics of muscles, as well as the dynamic relationship between muscle length, activation level, and force, thereby improving the realism of muscle actuation. Furthermore, to ensure the transmission effect of muscle contraction on skin deformation, a fiber-guided interpolation method is used to construct the muscle stress field, allowing muscle force to influence surrounding soft tissues anisotropically.

[0063] Next, facial motion trajectory data is introduced as boundary conditions into the finite element dynamics equations. Combined with skeletal constraints, this ensures that facial muscle movements conform to real physical constraints. The Lagrange multiplier method is used to enforce boundary constraints into the dynamics equations, avoiding numerical instabilities that may be introduced by traditional penalty function methods and ensuring coordination between facial motion trajectories and muscle forces.

[0064] Subsequently, the Newmark-β method was used to solve the dynamic equations, balancing computational efficiency and stability to obtain the displacement of each finite element node. The Newmark-β method is an implicit integration method that effectively handles large deformation problems and suppresses high-frequency numerical oscillations, resulting in a smoother facial deformation process that conforms to physiological characteristics. To further improve computational stability during the solution process, nonlinear iterative methods (such as the Newton-Raphson iteration) were employed to optimize displacement convergence. A multi-scale solution strategy was also combined, with pre-calculation performed on a coarse mesh to accelerate convergence and on a fine mesh to improve local accuracy.

[0065] Finally, the solved nodal displacements *u* are mapped onto the skin surface, and a dynamic craniofacial model is generated through radial basis function (RBF) interpolation. RBF interpolation can smoothly propagate the displacement information of the finite element mesh while maintaining facial geometric continuity, thus generating a craniofacial model with high-precision dynamic deformation effects. Furthermore, texture mapping techniques can be combined to finely model the textural changes such as skin wrinkles and stretching during facial expression changes, making the final dynamic craniofacial model more realistic in both visual and physical representation.

[0066] It should be noted that the nonlinear hyperelastic soft tissue model can accurately describe the real physical properties of facial tissues and improve the simulation effect of the influence of muscle contraction on skin deformation. The muscle force model based on electromyography signals is adopted so that the muscle contraction process can realistically reflect facial expression movements of different intensities, improving the physiological rationality of the dynamic craniofacial model. The Newmark-β method combined with Newton-Raphson iteration is used to ensure the numerical stability of the solution process of facial dynamics equations, and the multi-scale solution method accelerates the calculation and improves the simulation efficiency.

[0067] Furthermore, the geometric error and motion consistency scores of the real data are fused using linear regression to output the comprehensive reconstruction quality index. The specific content is as follows: the Hausdorff distance between the vertices of the static three-dimensional anatomical model and the real anatomical landmarks is calculated as the geometric error score; the consistency score of the facial motion trajectory data is quantified by the dynamic time warping algorithm; and the geometric error score and motion consistency score are fused using linear regression to generate the reconstruction quality index.

[0068] Furthermore, a linear regression method is used to fuse geometric error scores and motion consistency scores to generate a reconstruction quality index, expressed as:

[0069] Q = α·S geo +β·S motion ;

[0070] Where Q is the reconstruction quality index, α is the first weighting coefficient, β is the second weighting coefficient, and S...geo S is used to score the geometric error. motion Scoring for motor consistency.

[0071] See Figure 2 As shown, this embodiment of the invention also provides a craniofacial dynamic reconstruction system based on multimodal data fusion, the system comprising:

[0072] The data acquisition module is configured to acquire static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals;

[0073] The data processing module is configured to perform rigid and non-rigid registration on static CT and static MRI images to generate a fused static three-dimensional anatomical model; perform optical flow analysis and keyframe extraction on dynamic facial expression video sequences to obtain facial motion trajectory data; and perform filtering and time-frequency analysis on surface electromyography signals to extract muscle activation features.

[0074] The dynamic modeling module is configured to simulate the coupling relationship between muscle contraction force and skin deformation by integrating the fused static 3D anatomical model, facial motion trajectory data, and muscle activation features to generate a dynamic craniofacial model.

[0075] The evaluation output module is configured to use linear regression to fuse geometric error and motion consistency scores from real data, and output a comprehensive reconstruction quality index.

[0076] Therefore, this invention provides a method and system for dynamic craniofacial reconstruction based on multimodal data fusion. It employs rigid and non-rigid registration of static CT and MRI images, aligns skeletal structures using a mutual information algorithm, and aligns soft tissues using a B-spline algorithm. This simultaneously preserves skeletal details and soft tissue information, improving the accuracy of the static three-dimensional anatomical model. The invention uses an improved Lucas-Kanade optical flow method to extract facial keypoint displacements, combines this with a time-series clustering algorithm to extract keyframes, and further utilizes a convolutional neural network to fuse facial texture features, achieving accurate representation of facial motion trajectories. The invention removes power frequency interference and motion artifacts from electromyographic signals using wavelet packet transform, separates target muscle group signals using independent component analysis (ICA), and extracts time-domain and frequency-domain features to accurately obtain muscle activation features, providing high-quality input data for biomechanical drives. The invention also employs nonlinear hyperelasticity... This invention discretizes facial tissue using a finite element method, constructs a stiffness matrix K, converts muscle activation features into muscle contraction forces through a muscle force model, and combines facial motion trajectory data as boundary conditions. It embeds the dynamic equations using the Lagrange multiplier method and solves them using the Newmark-β method, balancing computational efficiency and stability to generate a more physiologically accurate dynamic craniofacial model. The invention calculates the geometric error score between the vertices of the static 3D anatomical model and real anatomical landmarks using Hausdorff distance, quantifies the consistency score of facial motion trajectories using a dynamic time warping algorithm, and fuses these two indicators using linear regression to generate a comprehensive reconstruction quality index, thus providing an objective model evaluation standard. This invention can be applied to multiple scenarios such as medical surgical planning, facial expression disorder assessment, personalized prosthesis design, and virtual human construction, providing a precise craniofacial dynamic modeling solution for the medical and computer graphics fields.

[0077] This invention can be used in multiple application scenarios such as medical surgery planning, facial expression disorder assessment, personalized prosthesis design, and virtual human construction. For example, it can capture the user's facial muscle movements (such as subtle changes in eyebrows and corners of the mouth) through cameras or sensors, combine them with deep learning algorithms to identify emotional states (such as anger, sadness, and surprise), and at the same time, use generative adversarial networks (GAN) or 3D modeling technology to generate micro-expressions (such as raising eyebrows and pursing lips during dialogue) that match the semantic content of the virtual character.

[0078] This invention can be applied to virtual humans, such as virtual assistants, virtual idols, and corporate customer service avatars. These virtual humans can recognize user emotions through micro-expressions and dynamically adjust dialogue strategies (such as shortening responses when they detect user impatience).

[0079] It can also be applied to mental health and education, such as in psychological counseling robots, where micro-expressions can be used to identify users' hidden emotions (such as unnatural twitching of the corners of the mouth when forcing a smile) and generate empathetic dialogue.

[0080] When applied to language learning assistants, virtual teachers display encouraging / affirming micro-expressions when correcting pronunciation, enhancing learning motivation;

[0081] It can also be applied to social and entertainment purposes, such as: users uploading selfies to generate "dynamic virtual clones with expressive faces", sending voice messages with micro-expressions on social platforms, or game NPCs triggering different plot branches based on the player's emotions (such as fear expressions triggering a comforting plot).

[0082] This invention can also be applied to dynamic micro-expression monitoring technology. By capturing and analyzing brief, unconscious micro-expressions (typically lasting 0.04 to 0.5 seconds), it can reveal hidden emotions and true psychological states. With advancements in artificial intelligence, computer vision, and neuroscience, the following is an analysis of its potential areas of deep application and specific scenarios:

[0083] 1. Mental Health and Clinical Diagnosis

[0084] Early screening for depression / anxiety: Quantifying emotional fluctuations through micro-expressions during patient conversations (such as brief drooping of the corners of the mouth and slight tightening of the eyebrows), combined with AI models to predict the risk of mental illness, and assisting doctors in developing intervention plans.

[0085] Autism intervention: Real-time monitoring of autistic children's micro-responses to social stimuli (such as facial muscle tremors when avoiding eye contact) and adjustment of rehabilitation training content.

[0086] Post-traumatic stress disorder (PTSD) assessment: During exposure therapy, monitor patients' fear micro-expressions in response to triggers (such as nasal flaring, rapid eye movement) and dynamically adjust the intensity of treatment.

[0087] 2. Medical diagnostic assistance

[0088] Early diagnosis of ALS: assisting in the screening of neurodegenerative diseases through minor motor impairments of facial muscles (such as unilateral delay of the corner of the mouth when attempting to smile).

[0089] Pain management: Quantify patients' hidden pain expressions after surgery (such as brief eyebrow raising) to optimize analgesic dosage.

[0090] Drug side effect monitoring: Capturing micro-expressions of disgust (such as temporary nasal wrinkles) in subjects after taking medication during clinical trials to quickly identify potential problems.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A method for dynamic craniofacial reconstruction based on multimodal data fusion, characterized in that, include: Multimodal data acquisition devices were deployed in the target craniofacial region to acquire static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals. Rigid and non-rigid registration are performed on static CT images and static MRI images to generate a fused static three-dimensional anatomical model; Optical flow analysis and keyframe extraction are performed on dynamic facial expression video sequences to obtain facial motion trajectory data; Filtering and time-frequency analysis of surface electromyography signals were performed to extract muscle activation features; The fused static 3D anatomical model, facial motion trajectory data, and muscle activation features are input into a multi-scale finite element model, and a biomechanical driving strategy is used to simulate the coupling relationship between muscle contraction force and skin deformation to generate a dynamic craniofacial model. A linear regression method is used to fuse geometric error and motion consistency scores from real data to output a comprehensive reconstruction quality index. The specific content of performing optical flow analysis and keyframe extraction on dynamic expression video sequences to obtain facial motion trajectory data is as follows: the improved Lucas-Kanade optical flow method is used to extract the displacement of facial key points in the dynamic expression video sequence to generate original motion trajectory data. Then, a time series clustering algorithm is used to filter keyframes in the original motion trajectory data to construct a sparse representation of facial motion trajectory. Finally, facial texture features in the dynamic expression video sequence are extracted through a convolutional neural network and fused with the sparse representation of facial motion trajectory to obtain facial motion trajectory data. Finally, facial texture features are extracted from dynamic facial expression video sequences using a convolutional neural network and fused with the sparse representation of facial motion trajectories to obtain the specific content of facial motion trajectory data: Optical flow timestamps are used to align the sparse representation of facial motion trajectories with facial texture features in a spatiotemporal manner. Then, the weights are dynamically adjusted through the cross-attention module in the convolutional neural network to enhance the modeling strength. Finally, the sparse representation of facial motion trajectories and facial texture features are concatenated into a joint feature vector to obtain facial motion trajectory data. The specific steps for generating a dynamic craniofacial model are as follows: The static three-dimensional anatomical model, facial motion trajectory data, and muscle activation features are input into a multi-scale finite element model, and a biomechanical driving strategy is used to simulate the coupling relationship between muscle contraction force and skin deformation. This involves discretizing the static three-dimensional anatomical model into nonlinear hyperelastic soft tissue elements, constructing a stiffness matrix K, converting muscle activation features into muscle contraction force using a muscle force model, using facial motion trajectory data as boundary conditions, embedding the Lagrange multiplier method into the dynamic equations, solving the dynamic equations using the Newmark-β method, balancing computational efficiency and stability to obtain the nodal displacement u, mapping the obtained nodal displacement u onto the skin surface, and generating the dynamic craniofacial model through radial basis function interpolation.

2. The method for dynamic craniofacial reconstruction based on multimodal data fusion according to claim 1, characterized in that, The multimodal data acquisition device includes a CT scanner, an MRI device, a high-speed camera, and a wireless surface electromyography sensor array.

3. The method for dynamic craniofacial reconstruction based on multimodal data fusion according to claim 2, characterized in that, The specific content of performing rigid and non-rigid registration of static CT images and static MRI images to generate a fused static three-dimensional anatomical model is as follows: the rigid registration algorithm based on mutual information is used to align the skeletal structure of the static CT images and static MRI images to obtain skeletal details; A non-rigid registration algorithm based on B-spline was used to align soft tissue contours in static CT and static MRI images to obtain soft tissue information. By integrating skeletal details and soft tissue information through a weighted fusion strategy, a high-precision static 3D model is generated.

4. The method for dynamic craniofacial reconstruction based on multimodal data fusion according to claim 1, characterized in that, The specific steps for filtering and time-frequency analysis of surface electromyography (EMG) signals to extract muscle activation features are as follows: First, wavelet packet transform is used to remove power frequency interference and motion artifacts from the EMG signals; then, independent component analysis is used to separate the target muscle group signals from the EMG signals; finally, time-domain and frequency-domain features are extracted to obtain muscle activation features.

5. The method for dynamic craniofacial reconstruction based on multimodal data fusion according to claim 1, characterized in that, The specific content of the method of using linear regression to fuse geometric error and motion consistency scores of real data to output a comprehensive reconstruction quality index is as follows: the Hausdorff distance between the vertices of the static three-dimensional anatomical model and the real anatomical landmarks is calculated as the geometric error score, the consistency score of facial motion trajectory data is quantified by dynamic time warping algorithm, and the geometric error score and motion consistency score are fused by linear regression to generate the reconstruction quality index.

6. The method for dynamic craniofacial reconstruction based on multimodal data fusion according to claim 5, characterized in that, The linear regression method is used to fuse geometric error scores and motion consistency scores to generate a reconstruction quality index, expressed as follows: in, To reconstruct the quality index, As the first weighting coefficient, This is the second weighting coefficient. Scoring for geometric errors, Scoring for motor consistency.

7. A craniofacial dynamic reconstruction system based on multimodal data fusion, implemented in any one of claims 1-6, characterized in that, The system includes: The data acquisition module is configured to acquire static CT images, static MRI images, dynamic facial expression video sequences, and surface electromyography signals; The data processing module is configured to perform rigid and non-rigid registration on static CT and static MRI images to generate a fused static three-dimensional anatomical model; perform optical flow analysis and keyframe extraction on dynamic facial expression video sequences to obtain facial motion trajectory data; and perform filtering and time-frequency analysis on surface electromyography signals to extract muscle activation features. The dynamic modeling module is configured to simulate the coupling relationship between muscle contraction force and skin deformation by integrating the fused static 3D anatomical model, facial motion trajectory data, and muscle activation features to generate a dynamic craniofacial model. The evaluation output module is configured to use linear regression to fuse geometric error and motion consistency scores from real data, and output a comprehensive reconstruction quality index.

Citation Information

Patent Citations

  • Face model generation method and device, electronic equipment and readable storage medium

    CN116863044A

  • Image three-dimensional post-processing method and system based on multiple modes

    CN119048694A