A method, system, device, and medium for grading facial paralysis by integrating multimodal data.
By combining multimodal data fusion and prior knowledge features, the problem of inaccurate facial paralysis diagnosis caused by single modal features is solved, and high-precision graded diagnosis of facial paralysis is achieved in complex environments.
Patent Information
- Application Number
- CN202511384069.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing intelligent grading and diagnostic methods for facial paralysis rely on single modal features, which makes it difficult to cope with complex real-world scenarios. Furthermore, the lack of data resources leads to inaccurate diagnostic results, preventing their widespread clinical application.
A multimodal data fusion method is adopted, combining prior knowledge-based manual facial paralysis features and depth visual features. Through heterogeneous feature interaction attention mechanism and multi-scale feature fusion, a two-way information interaction channel between dynamic facial features and static symmetry features is constructed to achieve complementarity between dynamic and static features.
It improves the accuracy of facial paralysis diagnosis and grading, and can accurately reflect the muscle motor dysfunction of facial paralysis patients in complex environments, generating more comprehensive and interpretable diagnostic evidence.
Smart Images

Figure CN120876479B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, and in particular relates to a method, system, device and medium for classifying facial paralysis by fusing multimodal data. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Artificial intelligence-based intelligent grading and assisted diagnosis of facial paralysis not only improves work efficiency compared to traditional manual diagnosis but also largely eliminates subjective differences in diagnostic results, thus possessing significant research value and application prospects. However, current intelligent grading and diagnostic methods for facial paralysis are still in the academic research stage and have not yet been applied clinically. A major reason is that most existing methods simply view the entire diagnostic process as a mapping from a single modal feature, such as facial image features, to the output of the severity level of facial paralysis, and use deep neural networks to fit this mapping relationship. However, real-world scenarios are complex and diverse; factors such as race, gender, and skin color can interfere with the intelligent diagnostic results, and information from a single modal feature often cannot fully reflect the situation. Furthermore, training a robust and generalizable network for facial paralysis diagnosis requires a large amount of accurately labeled sample data, but currently available data resources are still very scarce compared to actual needs. Collecting and accurately labeling a large amount of patient data in real-world scenarios is extremely difficult, inefficient, and fails to cover all situations, even raising issues related to patient privacy and security. However, both in terms of methodology and data, existing methods based on single modal features are difficult to obtain accurate diagnostic results for facial paralysis and lack credibility. In the medical field, credibility is an important prerequisite for ensuring widespread application and user trust. Therefore, existing methods can only be tested in laboratory or simple real-world scenarios and are difficult to cope with complex real-world scenarios, making them unsuitable for practical clinical application.
[0004] In clinical practice, when diagnosing facial paralysis, doctors typically combine various medical characteristics, such as the patient's static facial appearance and dynamic facial movements, for a comprehensive assessment. This indicates that facial paralysis diagnosis does not rely solely on a single modality of features from facial image data. Therefore, in the process of diagnosing facial paralysis, relying solely on limited facial image data may not be sufficient to accurately assess the severity of the patient's condition.
[0005] While there are currently methods for triaging facial paralysis based on multimodal data, traditional multimodal fusion methods, such as early-stage, late-stage, and hybrid fusion, often result in information redundancy during feature fusion. Furthermore, there is modal heterogeneity with significant representational differences between different modalities, leading to relatively low model learning efficiency. In addition, even when using multimodal data, some computer vision-based methods remain weak in capturing subtle changes in facial muscle movement, and the accuracy of key point detection and texture analysis is significantly affected in complex environments.
[0006] Therefore, how to effectively extract and combine the unique information contained in various modal data to enable it to cope with complex real-world scenarios and obtain accurate diagnostic and grading results for facial paralysis is a problem that needs to be solved. Summary of the Invention
[0007] To overcome the shortcomings of the prior art, this invention provides a method, system, device, and medium for grading facial paralysis by integrating multimodal data. It introduces manual facial paralysis features based on prior knowledge to construct dynamic facial features, and considers facial asymmetry to construct static symmetrical features. This effectively integrates information from multiple modal data and improves the accuracy of facial paralysis diagnosis and grading.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides a method for classifying facial paralysis by fusing multimodal data, comprising:
[0010] Based on dynamic videos of a target individual performing standardized facial movements, depth visual features containing spatiotemporal information are extracted.
[0011] Based on the extracted deep visual features containing spatiotemporal information, a heterogeneous feature interaction attention mechanism is used to combine the manual facial paralysis features based on prior knowledge to perform correlation modeling to obtain a bimodal representation. The bimodal representation is then fused to obtain dynamic facial features.
[0012] Multi-scale features are extracted from the static image characteristics and key point features of a single frame facial image of a target individual. Small-scale features representing local asymmetry are gradually transferred to medium-scale and large-scale features. Feature fusion is then performed adaptively through dynamic weight allocation to obtain static symmetrical features.
[0013] A two-way information interaction channel is constructed between dynamic facial features and static symmetry features to achieve complementarity between the two features. The complementary dynamic facial features and static symmetry features are then fused together, and the facial paralysis classification result of the target individual is obtained based on the fused features.
[0014] Secondly, the present invention provides a facial paralysis grading system that integrates multimodal data, comprising:
[0015] The extraction unit is configured to extract depth visual features containing spatiotemporal information from dynamic videos of a target individual performing standardized facial movements.
[0016] The dynamic facial feature unit is configured to: obtain a bimodal representation by using a heterogeneous feature interaction attention mechanism based on the extracted deep visual features containing spatiotemporal information and combining them with the manual facial paralysis features based on prior knowledge for correlation modeling; and fuse the bimodal representation to obtain dynamic facial features.
[0017] The static symmetric feature unit is configured to: extract multi-scale features from the static image characteristics and key point features of a single frame facial image of the target individual, gradually transfer the small-scale features representing local asymmetry to the medium-scale and large-scale features, and perform feature fusion adaptively through dynamic weight allocation to obtain static symmetric features;
[0018] The grading unit is configured to: construct a two-way information interaction channel between dynamic facial features and static symmetry features, realize the complementarity of dynamic facial features and static symmetry features, fuse the complementary dynamic facial features and static symmetry features, and obtain the facial paralysis grading result of the target individual based on the fused features.
[0019] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0020] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0021] The above one or more technical solutions have the following beneficial effects:
[0022] In this invention, a priori knowledge-based manual facial paralysis feature is introduced. This priori knowledge-based manual facial paralysis feature directly corresponds to the core anatomical region for facial paralysis diagnosis, enabling feature extraction to focus on pathological information and significantly enhancing the clinical interpretability of the features.
[0023] In this invention, a heterogeneous feature interaction attention mechanism is used to jointly model manual facial paralysis features based on prior knowledge with deep visual features containing spatiotemporal information. This allows the extracted dynamic facial features to more accurately reflect the muscle motor dysfunction of facial paralysis patients. Multi-scale features are extracted from the static image characteristics and key point features of a single-frame facial image of the target individual. Considering the complex multi-scale characteristics of motor dysfunction in facial paralysis patients, small-scale features representing local asymmetry are gradually transferred to medium-scale and large-scale features. Dynamic weight allocation is used for adaptive feature fusion to obtain static symmetrical features, ensuring that the extracted static features take into account both global and local asymmetries and adapt to complex case scenarios. A bidirectional information interaction channel is constructed between dynamic facial features and static symmetrical features, simulating the logic of clinicians combining static facial appearance and dynamic movements to judge the condition. This allows the generated fused features to simultaneously contain complete pathological information on spatial structural asymmetry and motor abnormalities, resulting in a more comprehensive diagnostic basis and improved diagnostic grading accuracy.
[0024] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0025] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0026] Figure 1(a) is a flowchart of the dataset processing in Embodiment 1 of the present invention;
[0027] Figure 1(b) is a schematic diagram of the annotation results of the standard 68 key points and sample face in Embodiment 1 of the present invention;
[0028] Figure 2 This is a schematic diagram of the overall network architecture of the multimodal dual-stream cross-attention fusion network in Embodiment 1 of the present invention;
[0029] Figure 3(a) is a schematic diagram of key points in the eyebrow and nose areas during facial paralysis assessment in Embodiment 1 of the present invention;
[0030] Figure 3(b) is a schematic diagram of key points in the eye area during facial paralysis assessment in Embodiment 1 of the present invention;
[0031] Figure 3(c) is a schematic diagram of key points in the oral region during facial paralysis assessment in Embodiment 1 of the present invention;
[0032] Figure 3(d) is a schematic diagram of key points of the whole face in the facial paralysis assessment in Embodiment 1 of the present invention;
[0033] Figure 4 This is a schematic diagram of the network structure of the gating fusion mechanism in Embodiment 1 of the present invention;
[0034] Figure 5 This is a schematic diagram of the multi-scale attention mechanism network structure in Embodiment 1 of the present invention;
[0035] Figure 6 This is a schematic diagram of the network structure of the cross-modal fusion module in Embodiment 1 of the present invention. Detailed Implementation
[0036] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0037] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0038] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0039] Example 1
[0040] This embodiment discloses a facial paralysis classification method that integrates multimodal data, including:
[0041] Based on dynamic videos of a target individual performing standardized facial movements, depth visual features containing spatiotemporal information are extracted.
[0042] Based on the extracted deep visual features containing spatiotemporal information, a heterogeneous feature interaction attention mechanism is used to combine the manual facial paralysis features based on prior knowledge to perform correlation modeling to obtain a bimodal representation. The bimodal representation is then fused to obtain dynamic facial features.
[0043] Multi-scale features are extracted from the static image characteristics and key point features of a single frame facial image of a target individual. Small-scale features representing local asymmetry are gradually transferred to medium-scale and large-scale features. Feature fusion is then performed adaptively through dynamic weight allocation to obtain static symmetrical features.
[0044] A two-way information interaction channel is constructed between dynamic facial features and static symmetry features to achieve complementarity between the two features. The complementary dynamic facial features and static symmetry features are then fused together, and the facial paralysis classification result of the target individual is obtained based on the fused features.
[0045] This embodiment introduces pre-knowledge-based manual facial paralysis features, which directly correspond to the core anatomical regions for facial paralysis diagnosis. This allows feature extraction to focus on pathologically relevant information, significantly enhancing the clinical interpretability of the features. Through a heterogeneous feature interaction attention mechanism, pre-knowledge-based manual facial paralysis features are jointly modeled with depth visual features containing spatiotemporal information, enabling the extracted dynamic facial features to more accurately reflect the muscle motor dysfunction of facial paralysis patients. For the static image characteristics and key point features of a single frame of the target individual's facial image, a cross-fusion Transformer encoder network is used to obtain multi-scale features. Considering the complex multi-scale characteristics of motor dysfunction in facial paralysis patients, small-scale features representing local asymmetry are progressively transferred to medium-scale and large-scale features. Dynamic weight allocation is used for adaptive feature fusion to obtain static symmetrical features, ensuring that the extracted static features take into account both global and local asymmetries, adapting to complex case scenarios. A bidirectional information interaction channel is constructed between dynamic facial features and static symmetrical features, simulating the logic of clinicians combining static facial appearance and dynamic movements to jointly judge the condition. This ensures that the generated fused features simultaneously contain complete pathological information on spatial structural asymmetry and motor abnormalities, providing a more comprehensive diagnostic basis.
[0046] In the design and training of facial paralysis assessment models, sufficient and accurately labeled data is often crucial. Currently, the field of facial paralysis assessment requires large-scale, well-labeled datasets to support model training, evaluation, and benchmarking. However, the number of existing publicly available datasets is limited, and they suffer from inconsistent labeling standards and inconsistent labeling, lacking a unified benchmark dataset for comparison and validation between different methods.
[0047] To address this challenge, this embodiment reorganizes and annotates two public datasets: MEEI and AFLFP. Table 1 summarizes the details of these datasets.
[0048] It should be noted that the face images in the attached figures of this embodiment are from the AFLFP public dataset. In the usage requirements of this dataset, some numbered face images are allowed to be used publicly. The attached figures of this embodiment use face image number 5. The corresponding URL of the dataset is: https: / / github.com / Yifan313 / AFLFP / blob / main / End_User_License_Agreement.pdf.
[0049] The Massachusetts Eye and Ear Institute (MEEI) dataset is a laboratory-controlled facial paralysis dataset consisting of videos of 60 subjects performing eight different facial movements. To standardize assessment criteria, the severity of facial paralysis assessed using the eFACE scale in the original dataset was re-labeled according to the HB scale, categorizing facial paralysis severity into six HB levels from 1 to 6, with HB=6 representing the most severe. Furthermore, to reduce labeling and judgment errors, the data was further divided into mild (1≤HB score≤2), moderate (3≤HB score≤4), and severe (5≤HB score≤6), with 20 subjects in each category.
[0050] The Facial Landmark Annotation for Facial Nerve Palsy (AFLFP) dataset is a diverse, lab-controlled dataset containing videos of 255 participants performing six different facial movements (such as raising eyebrows, slightly closing eyes, and smiling). After re-annotation, participants were categorized into mild, moderate, and severe cases based on the severity of facial paralysis. To ensure data balance and remove low-quality data, 40 participants were ultimately retained for each category, with each participant performing six or fewer facial movements.
[0051] Table 1: Overview of the Facial Paralysis Dataset
[0052]
[0053] In this embodiment, keyframes were selected, facial key points were labeled, and data preprocessing was performed on the MEEI and AFLFP facial paralysis datasets. The specific process is shown in Figure 1(a).
[0054] (1) Data labeling.
[0055] To comprehensively capture the dynamic changes in facial movements, 32 keyframes were manually extracted from each subject's facial movement video. These keyframes covered the entire process from a static facial state to the gradual increase in movement intensity to a peak, and finally back to a static state. To minimize potential errors, the selected keyframes from each video were manually checked again to ensure they could fully represent the entire process of a facial movement.
[0056] Subsequently, all keyframes were manually annotated independently, with 68 facial key points marked for each facial image. To improve annotation efficiency, the facial key points were first preliminarily detected using the key point detection software Emotrics. Based on this, annotators manually corrected inaccurate or missing key points. Finally, each annotated image underwent manual review to ensure the reliability of the annotation results. Figure 1(b) shows the annotation results of the model and example sample face with the standard 68 key points.
[0057] This embodiment collected 15,328 and 22,144 facial image keyframes from the MEEI and AFLFP datasets, respectively, and preprocessed them for a facial paralysis grading assessment task. Specifically, face detection was performed on each facial image frame, and irrelevant information such as the background was cropped to retain only the facial portion. The cropped image size was adjusted to 112×112, and image enhancement strategies such as flipping, scaling, and random cropping were applied to increase the diversity of training data and improve the model's generalization ability.
[0058] This embodiment proposes a multimodal dual-stream cross-attention fusion network for facial paralysis grading. The multimodal dual-stream cross-attention fusion network comprises three core modules: a dynamic feature extraction module, a static feature extraction module, and a cross-modal fusion network. The dynamic feature extraction module processes continuous frame-by-frame facial video sequences to capture temporal anomalies in facial muscle movements. The static feature extraction module analyzes single-frame images at peak facial movement moments to quantify bilateral facial asymmetry. The cross-modal fusion network deeply fuses complementary information from dynamic and static features through an attention mechanism. The dual-stream architecture design of the multimodal dual-stream cross-attention fusion network achieves, for the first time, collaborative modeling of dynamic muscle movement anomalies and static facial asymmetry in facial paralysis diagnosis.
[0059] The core pathological feature of facial paralysis is muscle motor dysfunction caused by facial nerve damage. Dynamic video, by continuously recording the entire process of a patient performing standardized facial movements such as raising eyebrows, closing eyes, and showing teeth, can accurately capture abnormal muscle movement patterns that cannot be shown in static images. Therefore, this embodiment designs a two-branch dynamic feature extraction module, using continuous frame video of the patient performing standardized facial movements as input, and capturing pathological movement features through multi-level processing.
[0060] This embodiment introduces manual facial paralysis features constructed based on prior clinical knowledge and integrates them with deep visual features containing spatiotemporal information. This allows dynamic facial feature extraction to focus on key areas of facial paralysis assessment, significantly enhancing the interpretability and pathological relevance of feature expression.
[0061] Specifically, input a series of frame videos of standardized facial movements. Where T represents the number of frames in the video sequence, and H and W correspond to the height and width of a single frame, respectively. A ResNet-50 pre-trained on the MS-CELEB dataset was used as the face image feature extractor to encode the spatial information of each frame. Simultaneously, a 30-dimensional hand-drawn facial paralysis feature extractor was used to calculate the hand-drawn facial paralysis features. Thirty hand-drawn facial paralysis features calculated based on facial key points were integrated, as shown in Tables 2-5, illustrating the key points involved and the calculation methods for these features.
[0062] This study quantifies marked facial key points through mathematical modeling to measure the degree of facial asymmetry. These features cover four major anatomical regions: the brow, eye, nose, and mouth, and incorporate whole-face motion features as a comprehensive evaluation dimension.
[0063] To minimize the impact of head tilt angle, tilt correction is first performed, based on the numbering sequence of the standard 68 key point model. and Calculate the rotation matrix using the reference key points. ,use Perform affine transformations on all key points:
[0064]
[0065] in, These are the x and y coordinates of a key point before the affine transformation; These are the new x and y coordinates of the corresponding key points after affine transformation; M is the affine transformation matrix, which is a 2×3 matrix. , , , These represent the values of a corresponding row and column in the transformation matrix M, respectively. The first number in the lower right corner represents the row, and the second number represents the column.
[0066] Then, as shown in Figures 3(a)-3(d), following the steps in Tables 2-5, 30 geometric indicators such as angle and distance are calculated as non-semantic values to quantify specific facial organs and the interaction between different organs.
[0067] Table 2:
[0068]
[0069] Table 3:
[0070]
[0071] Table 4:
[0072]
[0073] Table 5:
[0074]
[0075] in, , ... The identifiers represent geometric metrics 1 through 30; uppercase letters A through X represent the straight-line distance between specific key points; P18... These represent the facial key points at the corresponding numbered locations; , Represents a key facial feature. and It is a point x and y coordinates and It is a point The x and y coordinates; Representative point , The angle between the vector formed and the positive X-axis; Represents the slope, indicating , The degree of inclination of the line connecting two points relative to the horizontal line; EORl , EOR Represents the ratio of the openness of the left and right eyes; Representative point , The Euclidean distance between them; The perimeter of the closed curve formed by points 49, 50, 51, 52, 58, 59, and 60; The perimeter of the closed curve formed by points 52, 53, 54, 55, 56, 57, and 58; and Represents the starting and ending points of a closed curve. It is used to calculate the perimeter of a closed curve formed between two points.
[0076] Figures 3(a)-3(d) provide a detailed description of the 30 symmetry measurements used in the facial paralysis assessment. The numbers represent the sequence numbers of the facial key points, consistent with the numbering order of the standard 68 key point model in Figure 1(b). In addition, the left and right pupils are additionally numbered 69 and 70.
[0077] Subsequently, the spatial information and facial paralysis features extracted from each frame are input into independent temporal convolutional networks (TCNs) to capture temporal dynamics and obtain depth visual features containing spatiotemporal information. and the characteristics of facial paralysis based on prior knowledge ,in, Represents the number of video frames. , These are the dimensions of the video features and the dimensions of the handcrafted features, respectively.
[0078] Given deep visual features containing spatiotemporal information and hand-made facial paralysis features based on prior knowledge, a heterogeneous feature interaction attention mechanism is used to model the association between the two features.
[0079] Specifically, the two feature vectors are first concatenated, and then a joint feature representation is obtained through a fully connected layer:
[0080]
[0081] Where FC stands for fully connected layer. The dimension represents the joint feature representation.
[0082] The joint feature representation is then fed into the joint interactive attention framework of the individual features, allowing the high-dimensional information in the deep visual features and the clinical quantitative indicators in the handcrafted features to guide and validate each other, forming a refined bimodal representation. and .
[0083] The above process can be represented as follows:
[0084] Characterization using bimodality Calculate the cross-correlation matrix with video features and handcrafted features respectively:
[0085]
[0086]
[0087] in, , representing the learnable weight matrix between video features and joint features; , represents the learnable weight matrix between handcrafted features and joint features, and the superscript T indicates transpose.
[0088] Attention weights are generated based on the cross-correlation matrix. :
[0089]
[0090]
[0091] in, , The weight matrix is a learnable matrix. express function.
[0092] And by combining residual connections, enhanced features are obtained:
[0093]
[0094]
[0095] in, , represents the learnable weight matrix for video features and the learnable weight matrix for handcrafted features, respectively.
[0096] Finally, the attention features are repeatedly input into the heterogeneous feature interaction attention, and gradually refined to obtain the final bimodal representation, namely the video modal representation. and manual modal characterization :
[0097]
[0098]
[0099] in, This represents the number of times recursive interaction attention calculations are performed.
[0100] In this embodiment, as Figure 4 As shown, the gated fusion module first stitches together the video modal representation along the feature dimension. and manual modal characterization Forming combined features Subsequently, the combined features are processed through a weighted computation network (WCN): the weighted computation network consists of two fully connected layers, which generate dynamic weight parameters through nonlinear transformation. These weighting parameters adaptively adjust the contribution strength of features from both modalities, prioritizing the retention of key features with clinical discriminative power while suppressing noise and redundant information.
[0101] The above process can be represented as:
[0102]
[0103]
[0104] in and Represents the learnable weight matrix. and For bias terms, The final extracted dynamic facial features, Indicates multiplication.
[0105] To construct a spatiotemporally complementary feature representation system, this embodiment designs a static feature extraction module to extract highly discriminative spatial features from a single frame of facial image, which complements the dynamic temporal features and enhances the model's ability to model spatial details.
[0106] Specifically, the ResNet-50 network, consistent with the dynamic feature extraction module, is used as the backbone network to obtain static image features. Simultaneously, key point features are obtained from the readily available facial landmark extractor, MobileFaceNet. Where P is the number of feature blocks and D is the dimension of each feature block. This refers to the feature transfer of static images. Key features The input is fed into a cross-fusion Transformer encoder network to achieve scale-adaptive geometry-appearance alignment, resulting in fused features at large, medium, and small scales. 、 、 .
[0107] Considering the complex multi-scale characteristics of motor dysfunction in facial paralysis patients, ranging from overall contour symmetry to local muscle abnormalities, and the significant differences in facial imaging scale in real-world scenarios, this embodiment proposes a multi-scale attention mechanism, such as... Figure 5 As shown, multi-resolution features are adaptively fused through dynamic weight allocation.
[0108] Specifically, the multi-scale attention mechanism achieves feature fusion through two deconvolution upsampling operations and multi-head cross-attention: first, it upsamples the smallest scale features. Resolution of mesoscale features Using it as a key-value pair and mesoscale feature The query performs the first multi-head cross-attention calculation, and the result is compared with the original mesoscale features. Feature fusion; then the enhanced mesoscale features Upsampling to the resolution of large-scale features As new key-value pairs and large-scale features The query performs a second multi-head cross-attention calculation, and the final result is compared with the original large-scale features. Fusion, output static symmetric features The entire process, through two upsampling operations and multi-head cross-attention computation, achieves a gradual transfer of small-scale semantic information to large-scale spatial details.
[0109] The above process can be represented as follows:
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116] Among them, MHCA(·) is a bullish crossover signal. K and V come from different sequences. It is a scaling factor used for numerical normalization to prevent extremely small gradients. Deconv() represents the deconvolution operation. , , These represent large-scale features, medium-scale features, and small-scale features, respectively.
[0117] Clinical assessment of facial paralysis requires simultaneous attention to both static and dynamic manifestations, from static facial symmetry to the degree of functional completion during voluntary movement, all of which directly affect the final assessment results. Based on this clinical logic, this embodiment designs a Transformer-based cross-modal fusion module, such as... Figure 6 As shown, by constructing a bidirectional information interaction channel between dynamic facial expression features and static symmetry features, deep complementarity and synergistic enhancement of the two modalities are achieved. This is represented as:
[0118] Dynamic features and static features are transformed independently into query vector (Q), key vector (K), and value vector (V), respectively:
[0119]
[0120]
[0121] in, , is a learnable weight matrix.
[0122] Then the query vector of the static modality Key vectors of dynamic modes Perform similarity matching, calculate attention weights, and correlate them with the dynamic modality value vector. The feature maps from the h attention heads are then weighted and summed, and then concatenated to generate new feature representations, resulting in the final output of the cross-attention mechanism.
[0123]
[0124]
[0125] in, h is a scaling factor used to scale the dot product results, with h=8 attention heads capturing multi-granularity interactions in parallel.
[0126] Similarly, for static modal flow:
[0127]
[0128]
[0129] in, For dynamic modal query vectors, The key vector is the one used in the static mode. This is a static modal value vector.
[0130] Each feature stream, after passing through the cross-attention module, is residually connected to the original features, and then undergoes a nonlinear transformation via a normalization layer (LN) and a feedforward neural network (MLP) to enhance features and prevent gradient vanishing. The specific operations are as follows:
[0131]
[0132]
[0133]
[0134]
[0135] Finally, the features from the two modalities are fused and concatenated, then passed to the MLP to obtain a predicted score for the severity of facial paralysis.
[0136]
[0137] in, , These are dynamic facial features and static facial features, respectively, after being complemented by a cross-attention mechanism.
[0138] The following experiment was conducted to verify the effectiveness of the solution in this embodiment:
[0139] A. Data
[0140] The effectiveness of the facial paralysis assessment method proposed in this embodiment was evaluated using the MEEI and AFLFP datasets. The MEEI dataset consisted of 479 facial videos from 60 subjects, each subject independently performing 8 facial movements. The AFLFP dataset used 692 facial videos from 120 subjects, each subject performing 6 facial movements. Keyframe selection was performed on all subjects' video data, selecting 32 keyframes for each facial movement video, and each frame was marked with facial landmarks. All data, after standardization and data augmentation, were used in the experiments; 80% of the data was used for training, and 20% for testing.
[0141] B. Implementation Details
[0142] The model in this embodiment was trained on a single NVIDIA RTX 3090 GPU with 24GB of memory. The entire network was trained using PyTorch and a stochastic gradient descent (SGD) optimizer with a momentum of 0.9 and a weight decay of 0.0001. The batch size during training was 16, the learning rate (lr) was set to 0.01, and the maximum number of training epochs was set to 100.
[0143] C. Evaluation Indicators
[0144] Accuracy, Precision, Recall, and F1 score are used as metrics to evaluate model performance, and their mathematical expressions are as follows:
[0145]
[0146]
[0147]
[0148]
[0149] TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative, respectively.
[0150] D. Experimental verification
[0151] Table 6: Ablation Experiment Results on AFLFP Dataset
[0152]
[0153] Table 7: Ablation Experiment Results on MEEI Dataset
[0154]
[0155] To systematically verify the contributions of different modal features and cross-modal fusion modules, an ablation experiment was designed. By gradually stripping away key modules of the model, the impact of dynamic temporal features, static appearance features, and their fusion strategy on the performance of facial paralysis assessment was quantitatively analyzed. Specifically, four comparison conditions were set: (1) using only static features to simulate traditional image diagnosis methods; (2) using only dynamic features to represent mainstream video analysis methods; (3) simply splicing dynamic and static features as a multimodal baseline method; and (4) using the cross-modal fusion module proposed in this invention to fuse dynamic and static features.
[0156] Tables 6 and 7 present the ablation experiment results on the AFLFP and MEEI datasets. Compared to single-modal methods, when relying solely on static features, the model achieves 61.70% and 66.67% accuracy on AFLFP and MEEI, respectively. However, when using only dynamic features, the accuracy is 70.92% and 69.79%, demonstrating the crucial role of dynamic temporal information in capturing facial motor function abnormalities in patients with facial paralysis. When using a simple dynamic and static feature concatenation strategy, the accuracy on the AFLFP dataset slightly improves to 71.63%, but the precision is lower than when using only dynamic features, indicating that mechanical concatenation is insufficient to effectively mine complementary information between modalities. In contrast, after introducing a cross-modal cross-fusion module, the model accuracy on the AFLFP dataset improves to 73.05%, while the precision, recall, and F1 score reach 73.07%, 72.69%, and 71.52%, respectively, all of which are optimal. On the MEEI dataset, the cross-modal fusion model significantly outperformed the simple splicing strategy's 73.96% accuracy with an accuracy of 78.12%, fully validating the core value of the cross-modal fusion module.
[0157] The experimental results above demonstrate that the model proposed in this embodiment achieves higher robustness and discrimination accuracy in complex pathological scenarios by simulating the clinical diagnostic logic of doctors comprehensively observing static symmetry and dynamic functional integrity, and by utilizing cross-modal cross-attention mechanisms to achieve deep synergy between abnormal muscle movement patterns and facial structural asymmetry features.
[0158] Example 2
[0159] The purpose of this embodiment is to provide a facial paralysis grading system that integrates multimodal data, including:
[0160] The extraction unit is configured to extract depth visual features containing spatiotemporal information from dynamic videos of a target individual performing standardized facial movements.
[0161] The dynamic facial feature unit is configured to: obtain a bimodal representation by using a heterogeneous feature interaction attention mechanism based on the extracted deep visual features containing spatiotemporal information and combining them with the manual facial paralysis features based on prior knowledge for correlation modeling; and fuse the bimodal representation to obtain dynamic facial features.
[0162] The static symmetric feature unit is configured to: extract multi-scale features from the static image characteristics and key point features of a single frame facial image of the target individual, gradually transfer the small-scale features representing local asymmetry to the medium-scale and large-scale features, and perform feature fusion adaptively through dynamic weight allocation to obtain static symmetric features;
[0163] The grading unit is configured to: construct a two-way information interaction channel between dynamic facial features and static symmetry features, realize the complementarity of dynamic facial features and static symmetry features, fuse the complementary dynamic facial features and static symmetry features, and obtain the facial paralysis grading result of the target individual based on the fused features.
[0164] In further embodiments, the following is also provided:
[0165] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0166] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0167] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0168] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0169] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0170] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0171] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method of facial paralysis grading fusing multi-modal data, characterized in that, The method comprises the following steps: based on the target individual performing a standardized facial action dynamic video, extracting a deep visual feature containing spatiotemporal information; based on the extracted deep visual feature containing spatiotemporal information, using a heterogeneous feature interaction attention mechanism, combining prior knowledge-based manual facial paralysis features for correlation modeling to obtain a dual-modal representation, and fusing the dual-modal representation to obtain a dynamic facial feature, specifically: splicing the deep visual feature containing spatiotemporal information and the prior knowledge-based manual facial paralysis feature to obtain a joint feature representation; feed the joint feature representation into a joint interactive attention framework respectively, so that the high-dimensional information in the deep visual feature and the clinical quantitative indicators in the manual facial paralysis feature can guide and verify each other to obtain a dual-modal representation, and calculate the cross-correlation matrix of the dual-modal representation with the video feature and the manual feature respectively, and generate attention weights based on the cross-correlation matrix, and obtain enhanced features by combining residual connections; fuse the dual-modal representation through a weight calculation network to obtain a dynamic facial feature; extract multi-scale features from the static image characteristics and key point features of the single frame facial image of the target individual, gradually transfer small-scale features representing local asymmetry to medium-scale and large-scale features, and adaptively fuse the features through dynamic weight distribution to obtain a static symmetric feature, specifically: based on the static image characteristics and key point features of the single frame facial image of the target individual, use a cross-fusion Transformer encoder network to obtain large-scale, medium-scale and small-scale features; up-sample the smallest scale feature to the resolution of the medium-scale feature, and use the up-sampled smallest scale feature as a key-value pair and the medium-scale feature as a query for the first multi-head cross-attention calculation; fuse the first multi-head cross-attention calculation result with the medium-scale feature, up-sample the enhanced medium-scale feature to the resolution of the large-scale feature, and use the up-sampled enhanced medium-scale feature as a new key-value pair and the largest scale feature as a query for the second multi-head cross-attention calculation; fuse the second multi-head cross-attention calculation result with the large-scale feature to obtain a static symmetric feature; 2. The method of claim 1, wherein the fusion of multimodal data for facial paralysis grading is characterized by, build a bidirectional information interaction channel between the dynamic facial feature and the static symmetric feature to realize the complementarity of the dynamic facial feature and the static symmetric feature, fuse the complementary dynamic facial feature and the static symmetric feature, and obtain the facial paralysis grading result of the target individual according to the fused feature. feed the joint feature representation into a joint interactive attention framework respectively, so that the high-dimensional information in the deep visual feature and the clinical quantitative indicators in the manual facial paralysis feature can guide and verify each other to obtain a dual-modal representation, specifically: calculate the cross-correlation matrix of the joint feature representation with the deep visual feature containing spatiotemporal information and the prior knowledge-based manual facial paralysis feature respectively; generate corresponding attention weights based on the cross-correlation matrix, and obtain the attention feature of the deep visual feature containing spatiotemporal information and the attention feature of the prior knowledge-based manual facial paralysis feature by combining residual connections according to the attention weights; The heterogeneous feature interaction attention mechanism is used to process attention features of the deep visual features containing the space-time information and the attention features of the manual facial paralysis features based on the prior knowledge, to obtain the dual-modal representation.
3. The method of claim 1, wherein the fusion of multimodal data for facial paralysis grading is characterized by, A bidirectional information interaction channel between the dynamic facial features and the static symmetry features is constructed, the dynamic facial features and the static symmetry features are complemented, the complemented dynamic facial features and the static symmetry features are fused, and a facial paralysis grading result of the target individual is obtained according to the fused features, specifically as follows: The dynamic facial features and the static symmetry features are respectively transformed by independent linear transformation to generate query vectors, key vectors and value vectors; The cross-attention mechanism is used to obtain the complemented dynamic facial features based on the query vectors of the static symmetry features and the key vectors of the dynamic facial features; The cross-attention mechanism is used to obtain the complemented static facial features based on the query vectors of the dynamic facial features and the key vectors of the static facial features; The complemented dynamic facial features and the complemented static facial features are fused to obtain the fused features. The multi-layer perception is used to obtain the facial paralysis grading result of the target individual based on the fused features.
4. The method of claim 1, wherein the fusion of multimodal data for facial paralysis grading is characterized by, The dual-modal representation is fused through a weight calculation network to obtain the dynamic facial features, specifically as follows: the multi-modal representation is spliced along the feature dimension to obtain a combined feature; the dynamic weight parameters are generated based on the combined feature by using nonlinear transformation, and the dual-modal representation is fused and calculated according to the dynamic weight parameters to obtain the dynamic facial features.
5. The method of claim 1, wherein the fusion of multimodal data for facial paralysis grading is characterized by, The manual facial paralysis features based on the prior knowledge are obtained, specifically as follows: for the eyebrow area, eye area, nose area and mouth area of the face, the non-semantic numerical values of the specific facial organs and the linkage of different organs are quantified by mathematical modeling to obtain the manual facial paralysis features based on the prior knowledge.
6. A system for fusion of multi-modal data for facial paralysis grading, the system comprising: The method comprises the following steps: The extraction unit is configured to extract the deep visual features containing the space-time information based on the dynamic video of the target individual performing the standardized facial action; The dynamic facial feature unit is configured to obtain the dual-modal representation by using the heterogeneous feature interaction attention mechanism to associate and model the deep visual features containing the space-time information and the manual facial paralysis features based on the prior knowledge based on the extracted deep visual features containing the space-time information, and fuse the dual-modal representation to obtain the dynamic facial features, specifically as follows: the deep visual features containing the space-time information and the manual facial paralysis features based on the prior knowledge are spliced to obtain the joint feature representation; The joint feature representation is fed into the joint interaction attention framework to guide and verify the high-dimensional information in the deep visual features and the clinical quantitative indicators in the manual facial paralysis features, to obtain the dual-modal representation, and the mutual correlation matrix is calculated based on the dual-modal representation and the video features and the manual features, the attention weight is generated based on the mutual correlation matrix, and the enhanced features are obtained by combining the residual connection; The dual-modal representation is fused through the weight calculation network to obtain the dynamic facial features. The static symmetry feature unit is configured to extract multi-scale features from the static image characteristics and key point features of the single-frame face image of the target individual, gradually transfer small-scale features representing local asymmetry to medium-scale features and large-scale features, and adaptively perform feature fusion through dynamic weight distribution to obtain static symmetry features. Specifically, based on the static image characteristics and key point features of the single-frame face image of the target individual, a cross-fusion Transformer encoder network is used to obtain large-scale features, medium-scale features and small-scale features. The smallest scale features are up-sampled to the resolution of the medium-scale features, and the up-sampled smallest scale features are used as key-value pairs and the medium-scale features are used as queries to perform first multi-head cross attention calculation; The first multi-head cross attention calculation result is fused with the medium-scale features, the enhanced medium-scale features are up-sampled to the resolution of the largest scale features, and the up-sampled enhanced medium-scale features are used as new key-value pairs and the largest scale features are used as queries to perform second multi-head cross attention calculation; The second multi-head cross attention calculation result is fused with the largest scale features to obtain static symmetry features. The hierarchical unit is configured to construct a bidirectional information interaction channel between the dynamic face features and the static symmetry features, realize the complementation of the dynamic face features and the static symmetry features, fuse the complementary dynamic face features and the static symmetry features, and obtain the facial paralysis grading result of the target individual according to the fused features.
7. An electronic device, comprising: A memory and a processor are included, and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, complete the method of any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, A computer instruction is stored, and when the computer instruction is executed by the processor, the method of any one of claims 1-5 is completed.
Citation Information
Patent Citations
Dynamic and static two-way interactive and collaborative micro-expression recognition method
CN120388410A