Bimodal facial paralysis identification optimization method and system

By combining facial images and electromyography signals, using CNN and LSTM to extract facial features and aligning them through the DTW algorithm, the problem that existing methods cannot effectively capture dynamic symmetry deviation and static features is solved, and more accurate facial paralysis recognition is achieved.

CN119992618AActive Publication Date: 2025-05-13GUILIN INST OF INFORMATION TECH

Patent Information

Application Number
CN202411947043.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing facial paralysis recognition method cannot effectively capture the dynamic symmetry deviation in expression changes, the static feature analysis is single, the recognition is inaccurate, and there is a lack of comprehensive judgment of multiple data.

Method used

The dual-modal recognition method is used to combine facial images and facial EMG signals, and facial key points and muscle texture features are extracted through convolutional neural network (CNN) and long-term recording network (LSTM), and embossing expression motion trajectories are captured, and the time and amplitude are aligned through dynamic time regularization (DTW) algorithm, and feature fusion is performed by combining the characteristics of the EMG signal.

Benefits of technology

The fusion of dynamic and static characteristics of facial paralysis is achieved, the accuracy and consistency of recognition is improved, and the slight facial paralysis and dynamic expression changes can be more accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992618A_ABST
    Figure CN119992618A_ABST
Patent Text Reader

Abstract

The invention discloses a bimodal facial paralysis recognition optimization method and system, and the method comprises the steps: obtaining a facial image and an electromyographic signal, and obtaining a first facial feature and a first electromyographic feature; inputting the facial image into a first model, outputting facial key points and muscle texture features, obtaining an expression action track according to a time sequence, and outputting a dynamic symmetry deviation; extracting a second electromyographic feature of the facial electromyographic signal; inputting the output features of the first model and the electromyographic signal features into a second model to obtain a total loss function and outputting a classification category; and inputting the obtained features and the total loss function into the second model to obtain a classification probability for updating the output of the second model. According to the method, an optimized convolutional neural network model is adopted to extract facial key points, static deviation and dynamic change are combined, the problem of fusion of dynamic and static features is solved, individual differences are eliminated through DTW, and the effect of precisely recognizing the facial paralysis degree through bimodal data fusion is achieved based on fusion of an attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dual-modal recognition, and in particular to a dual-modal facial paralysis recognition optimization method and system. Background Art

[0002] Facial paralysis is a common neurological disease characterized by partial or complete loss of facial muscle function, leading to symptoms such as facial asymmetry and expression disorders. The traditional House-Brackmmann facial nerve paralysis grading and assessment method mainly relies on the doctor's clinical experience and naked eye observation, which has certain subjectivity and limitations, prompting neurosurgeons to seek facial paralysis identification tools with the help of computer vision and artificial intelligence technology to assist clinical diagnosis and efficacy evaluation.

[0003] However, existing facial paralysis recognition methods have the following technical deficiencies and shortcomings: (1) They are unable to capture dynamic symmetry deviations during expression changes, such as the upward range of the corners of the mouth when smiling, and lack the fusion of dynamic and static features; (2) They lack the ability to recognize mild facial paralysis, and the static feature analysis is single, such as insufficient muscle movement but no significant static asymmetry; (3) The analysis of dynamic expressions depends on the patient's movement coordination, and the standardization of expression movements is insufficient. However, the movement amplitude and speed of different patients may vary greatly, affecting the consistency of the results. (4) Existing methods usually rely only on visual data, and do not combine other biometric features for comprehensive evaluation, such as electromyography (EMG), and lack the fusion of dual-modal data. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a dual-modal facial paralysis recognition optimization method and system to solve the current problems of being unable to capture dynamic symmetry deviations in the process of facial expression changes, single static feature analysis, inaccurate recognition of facial paralysis expressions, and lack of comprehensive judgment of multiple data.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a dual-modal facial paralysis recognition optimization method, comprising: acquiring a facial image and a facial electromyographic signal, performing preprocessing, and acquiring a first facial feature and a first electromyographic feature;

[0008] The preprocessed facial image is input into the first model to extract features, output facial key points and muscle texture features, obtain the expression movement trajectory according to the time series, compare the movement trajectory with the standard template, and output the dynamic symmetry deviation;

[0009] Extracting a second electromyographic feature of the preprocessed facial electromyographic signal;

[0010] The output features of the first model and the electromyographic signal features are input into the second model to perform feature fusion, obtain the total loss function and output the classification category;

[0011] The acquired features and the total loss function are input into the second model to obtain the classification probability, which is used to update the output of the second model.

[0012] As a preferred solution of the dual-modal facial paralysis recognition optimization method of the present invention, obtaining the first facial feature includes:

[0013] Define the left and right coordinates of a pair of facial corner mirror points, and calculate the Euclidean distance of the left and right coordinates;

[0014] Define the vector angles of the left and right coordinates about the symmetry axis respectively, which are used to calculate the symmetry angle deviation;

[0015] Define an arc formed by fitting discrete coordinates within a preset range of the mouth corner or eyebrow arch, and calculate the curvature of the arc;

[0016] Define the initial key points of the face and the key points of the expression movements when smiling or closing the eyes, and calculate the amplitude of expression changes to obtain the left-right change rate of symmetry;

[0017] The curvature of the arc and the left-right change rate of the symmetry are used as constraints, and the facial symmetry features are calculated based on the Euclidean distance of the left and right coordinates and the symmetry angle deviation.

[0018] As a preferred solution of the dual-modal facial paralysis recognition optimization method of the present invention, wherein: the pre-processed facial image is input into the first model, feature extraction is performed, and facial key points and muscle texture features are output, including: the facial image is input into a CNN model including an LSTM architecture and a DTW algorithm;

[0019] After multiple layers of convolution, pooling, and fully connected layers, the model outputs the predicted coordinates of n facial key points.

[0020] The contrast and homogeneity features of muscle texture features are extracted based on the gray-level co-occurrence matrix.

[0021] As a preferred solution of the dual-modal facial paralysis recognition optimization method of the present invention, wherein: the expression action trajectory is obtained according to the time series, the action trajectory is compared with the standard template, and the dynamic symmetry deviation is output, including: the extracted facial key points and muscle texture features are input as time series data into the LSTM architecture in the first model to obtain the expression action trajectory;

[0022] Compare the acquired expression movement trajectory with the expression standard template, and use the DTW algorithm to align the time and amplitude, including:

[0023] Assume that the set of m standard templates is U = {U1, U2, ..., U m};

[0024] Suppose the set of expression trajectories of n patients is X = {X1, X2, ..., X n};

[0025] Assume that the goal of DTW is the optimal matching path P = {(i1, j1), (i2, j2), ..., (i k , j k )};

[0026] Aligning U and X in time and amplitude, we get the DTW distance metric, expressed as:

[0027]

[0028] Among them, i∈m, j∈n;

[0029] Calculate and output the dynamic symmetry deviation between the patient and the standard template, including:

[0030] Combined with the DTW distance metric, the similarity based on the synchronization of time series and the change of feature amplitude is calculated, which is expressed as:

[0031]

[0032] Where T represents the total number of facial paralysis patients X and standard templates U at time step t;

[0033] If the similarity exceeds the set threshold, it is considered that there is a dynamic symmetry deviation of facial paralysis, and the similarity is output as a dynamic symmetry deviation value.

[0034] As a preferred solution of the dual-modal facial paralysis recognition optimization method of the present invention, wherein: extracting the second electromyographic feature of the preprocessed facial electromyographic signal comprises:

[0035] The pre-processed facial electromyographic signal is processed by short-time Fourier transform to obtain the time-frequency characteristic energy spectrum of facial paralysis;

[0036] Based on the time-frequency characteristic energy spectrum of facial paralysis, the dynamic spectrum difference is calculated as the second electromyographic feature.

[0037] As a preferred solution of the dual-modal facial paralysis recognition optimization method described in the present invention, the output features of the first model and the electromyographic signal features are input into the second model, feature fusion is performed, a classification loss function is obtained and a classification category is output, specifically including:

[0038] The output features of the first model are defined as visual features, expressed as:

[0039] Z={Z1,Z,…,Z i} i=1,2,…,n

[0040] Among them, Z i represents the i-th vector output by the first model, including facial key points, muscle texture features, and dynamic symmetry deviation;

[0041] The myoelectric features are defined as:

[0042] E={E1,E2,…,E i} i=1,2,…,n

[0043] Among them, E i Represents the dynamic spectrum difference feature extracted by short-time Fourier;

[0044] The visual features and electromyographic features are input into the Transformer model based on the attention mechanism. Through two layers of self-attention layers with different features and feedforward network processing, feature weighted fusion is performed and the classification category label is output.

[0045] The total loss function is obtained based on the similarity measurement function and the classification loss function in the second model.

[0046] As a preferred solution of the dual-modal facial paralysis recognition optimization method described in the present invention, the acquired features and the total loss function are input into the second model to obtain the classification probability, which is used to update the output of the second model, including: the acquired features and the total loss function are input into the second model for fusion, which is expressed as:

[0047] F(t)=[S, RMS, ΔD, ΔE(f), L]

[0048] Where S represents the first facial feature, RMS represents the first electromyographic feature, ΔD represents the dynamic symmetry deviation, ΔE(f) represents the dynamic spectrum difference, and L represents the total loss function;

[0049] The classification probability is obtained by iterating the second model, which is used to update the output classification of the second model, expressed as:

[0050]

[0051] In a second aspect, the present invention provides a dual-modal facial paralysis recognition optimization system, comprising:

[0052] An acquisition module, used for acquiring a facial image and a facial electromyographic signal, performing preprocessing, and acquiring a first facial feature and a first electromyographic feature;

[0053] A first feature extraction module is used to input the preprocessed facial image into the first model, perform feature extraction, output facial key points and muscle texture features, obtain expression movement trajectory according to the time series, compare the movement trajectory with the standard template, and output dynamic symmetry deviation;

[0054] A second feature extraction module, used for extracting a second electromyographic feature of the preprocessed facial electromyographic signal;

[0055] A feature fusion module, used for inputting the output features of the first model and the electromyographic signal features into the second model, performing feature fusion, obtaining a total loss function and outputting a classification category;

[0056] The updating module is used to input the acquired features and the total loss function into the second model to obtain the classification probability, which is used to update the output of the second model.

[0057] In a third aspect, the present invention provides an electronic device, comprising:

[0058] Memory and processor;

[0059] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the dual-modal facial paralysis recognition optimization method are implemented.

[0060] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the dual-modal facial paralysis recognition optimization method.

[0061] Compared with the prior art, the present invention has the following beneficial effects: the present invention combines the Euclidean distance and angle difference symmetry indexes, adopts an optimized convolutional neural network model to extract facial key points, captures the dynamic pattern in the expression change process through the time series LSTM model, combines the static key point deviation with the dynamic change trend, thereby constructing a hybrid feature to solve the fusion problem of dynamic and static features. A muscle texture module is introduced into the CNN model to extract more delicate changes from the facial image, and comparative learning is performed through the grayscale co-occurrence matrix texture features to enhance the sensitivity to subtle changes and solve the recognition ability of mild facial paralysis. Dynamic time warping (DTW) is used to align the expression movement trajectory in time and amplitude, so that the movement features are regularized, individual differences are eliminated, and the problem of movement standardization is solved. The fusion model based on the attention mechanism is used to process heterogeneous data, and feature matrices of different modes are constructed, so as to achieve the effect of accurately identifying the degree of facial paralysis by dual-modal data fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0063] Figure 1 A schematic diagram of the overall process of a dual-modal facial paralysis recognition optimization method according to an embodiment of the present invention;

[0064] Figure 2 A schematic diagram of the overall framework structure of a dual-modal facial paralysis recognition optimization method according to an embodiment of the present invention;

[0065] Figure 3 A dynamic wavelet transform spectrum diagram of an electromyographic signal in a dual-modal facial paralysis recognition optimization method according to an embodiment of the present invention;

[0066] Figure 4 This is a diagram showing the training effect of the prediction loss function in the dual-modal facial paralysis recognition optimization method described in one embodiment of the present invention. DETAILED DESCRIPTION

[0067] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0068] Example 1

[0069] Reference Figure 1-Figure 2 , as an embodiment of the present invention, provides a dual-modal facial paralysis recognition optimization method, comprising:

[0070] S100: Acquire a facial image and a facial electromyographic signal, perform preprocessing, and acquire a first facial feature and a first electromyographic feature;

[0071] In the implementation manner of the present application, the preprocessing in step S100 includes: denoising and labeling key points of the facial image;

[0072] Specifically, denoising is performed by filtering the facial image I through a Gaussian filter with a mean of 0 and a variance of σ. Remove noise components.

[0073] Specifically, the annotation includes: marking the position of the key points of the face P X ∈{P eye , P nose,…,P mouth};

[0074] Among them, X represents the eyes, nose tip, corners of the mouth, etc.;

[0075] Let each key point form a two-dimensional vector, expressed as:

[0076] P X = {P1, P2, ..., P i} i=1,2,…,n (1)

[0077] Among them, P i =(x i ,y i ) represents the plane coordinates of the two-dimensional key point of any point on the corresponding facial part.

[0078] Based on the reference center of the nose tip and the reference standard scale of the two eyes, the two-dimensional vector coordinates of each key point are normalized and can be expressed as:

[0079]

[0080] In the implementation manner of the present application, obtaining the first facial feature in step S100 includes the following steps A1-A5:

[0081] A1: Define the left and right coordinates of a pair of facial corner mirror points:

[0082] P L,m =(x L,m ,y L,m ) R,m =(x R,m ,y R,m ),

[0083] And calculate the Euclidean distance between the left and right coordinates, expressed as:

[0084]

[0085] A2: Define the vector angles of the left and right coordinates about the axis of symmetry as θ L ,θ R , used to calculate the symmetry angle deviation, expressed as:

[0086] D θ =||θ L -θ R || (4)

[0087] Among them, ||·|| represents the Euclidean norm, θ L With θ R Both (c x , c y) represents the coordinates of the reference center point c of the nose tip.

[0088] A3: Define the arc s formed by fitting the discrete coordinates within the preset range of the mouth corner or eyebrow arch, and calculate the curvature of the arc, which is expressed as:

[0089]

[0090] Among them, x′(s) and y′(s) are the first-order derivatives of the arc s in the x and y directions, respectively, and x″(s) and y″(s) are the second-order derivatives of the arc s in the x and y directions, respectively.

[0091] A4: Define the initial key points of the face when smiling or closing eyes And the key points of facial expressions Calculate the expression change amplitude AP to obtain the left-right change rate υ of symmetry, expressed as:

[0092]

[0093] v=|ΔP L -ΔP R | (7)

[0094] Among them, |·| represents the absolute value, ΔP L , ΔP R Respectively represent the degree of change of left and right expressions.

[0095] A5: The curvature of the arc and the left-right change rate of symmetry are used as constraints, and the facial symmetry features are calculated based on the Euclidean distance of the left and right coordinates and the symmetry angle deviation, which can be expressed as:

[0096] S=αD m +βD θ

[0097]

[0098] Among them, α and β are the weighted deviation coefficients of Euclidean distance and angle, respectively, and ε1 and ε2 represent the curvature and expression change rate thresholds for clinical judgment of facial paralysis, respectively.

[0099] In the embodiment of the present application, the preprocessing of the electromyographic data in step S100 includes:

[0100] The facial muscle EMG signal collected by the wearable sensor is x(t), and the signal obtained by the bandpass filter is:

[0101]

[0102] Here, h(·) represents a bandpass filter.

[0103] To further extract EMG signal features, x f (t) The signal can be decomposed into different frequency bands, and the noise is removed by wavelet coefficients:

[0104] x d (t) = ∑ k ∑ j ω j,k ·ψ j,k (t) (10)

[0105] Among them, j and k represent the number of signal frequency components and the number of basis functions respectively, ω j,k represents the corresponding wavelet coefficient, ψ j,k (t) represents the corresponding wavelet basis function.

[0106] In the embodiment of the present application, the first electromyographic feature in step S100 is specifically:

[0107] The signal characteristics are extracted from the root mean square, and the RMS expression is:

[0108]

[0109] Where T represents the upper limit of time integration.

[0110] S200: inputting the preprocessed facial image into the first model, performing feature extraction, outputting facial key points and muscle texture features, obtaining expression movement trajectory according to the time series, comparing the movement trajectory with the standard template, and outputting dynamic symmetry deviation;

[0111] In the embodiment of the present application, in step S200, the preprocessed facial image is input into the first model, feature extraction is performed, and facial key points and muscle texture features are output, including: inputting the facial image into a CNN model including an LSTM architecture and a DTW algorithm;

[0112] After the model's multi-layer convolution, pooling, and fully connected layers, the predicted coordinates of n facial key points are output, which can be expressed as:

[0113]

[0114] The contrast and homogeneity features in the muscle texture features are extracted according to the gray level co-occurrence matrix calculation, where the gray level co-occurrence matrix calculation can be expressed as:

[0115]

[0116] Where p(i, j, d, θ) represents the co-occurrence probability of gray levels i and j at distance d and angle θ, M and N are the total number of pixels in the image, δ(·) is the indicator function, and I(·) is the gray pixel value;

[0117] The muscle texture features extracted from contrast CON and homogeneity HOMO are expressed as:

[0118] CON=∑ i,j (ij) 2 p(i, j) (14)

[0119]

[0120] In the embodiment of the present application, in step S200, the expression movement trajectory is obtained according to the time series, the movement trajectory is compared with the standard template, and the dynamic symmetry deviation is output, which includes the following steps B1-B3:

[0121] B1: The extracted facial key points and muscle texture features are input into the LSTM architecture in the first model as time series data to obtain the expression movement trajectory;

[0122] Specifically, step B1 can be implemented by assuming that at time step t, is the predicted value of the key point position, f t is the corresponding muscle texture feature, then the input of the model is the facial feature vector

[0123] Furthermore, in order to capture the dynamic patterns of facial expression changes, such as the changing patterns of facial key points and muscle texture features when smiling or frowning, the extracted facial key points and muscle texture features are regarded as time series data and input into the LSTM model. The basic operation units are:

[0124] Input gate: i t =σ(W i ·[h t-1 , X t ]+b i ) (16)

[0125] Forget gate: f t =σ(W f ·[h t-1 , X t ]+b f ) (17)

[0126] Output gate: o t =σ(W o ·[h t-1 , X t ]+b o ) (18)

[0127] Update status: c t =f t ·c t-1 +i t tanh(Wc ·[h t-1 , X t ]+b c ) (19)

[0128] Hidden state: h t =o t ·tanh(c t ) (20)

[0129] B2: Compare the acquired expression movement trajectory with the expression standard template and align the time and amplitude through the DTW algorithm, including:

[0130] Assume that the set of m standard templates is U = {U1, U2, ..., U m};

[0131] Suppose the set of expression trajectories of n patients is X = {X1, X2, ..., X n};

[0132] Assume that the goal of DTW is the optimal matching path P = {(i1, j1), (i2, j2), ..., (i k , j k )};

[0133] Aligning U and X in time and amplitude, we get the DTW distance metric, expressed as:

[0134]

[0135] Among them, i∈m, j∈n;

[0136] B3: Calculate and output the dynamic symmetry deviation between the patient and the standard template, including:

[0137] Combined with the DTW distance metric, the similarity based on the synchronization of time series and the change of feature amplitude is calculated, which is expressed as:

[0138]

[0139] Where T represents the total number of facial paralysis patients X and standard templates U at time step t;

[0140] If the similarity ΔD exceeds the set threshold, it is considered that there is a dynamic symmetry deviation of facial paralysis, and the similarity is output as a dynamic symmetry deviation value.

[0141] Exemplarily, the threshold value set for the similarity ΔD may be, for example, 0.1 or 0.05. If the threshold value is exceeded, it is considered that there is a dynamic symmetry deviation of facial paralysis.

[0142] It should be noted that step S200 can effectively extract rich local and global features, such as the location of facial key points, from the preprocessed facial images by using a CNN model containing multiple layers of convolution and pooling; the facial key points and muscle texture features are input into the LSTM architecture as time series data, so that the model can understand the temporal dimension of expression changes and capture continuous expression movements (such as the development process of a smile). The DTW algorithm can accurately compare two expression sequences even when the speed or intensity is different, thereby improving the flexibility and accuracy of matching. The above method can help detect potential facial nerve damage at an early stage, which is of great significance for clinical diagnosis.

[0143] S300: extracting a second electromyographic feature of the preprocessed facial electromyographic signal;

[0144] In the embodiment of the present application, extracting the second electromyographic feature of the preprocessed facial electromyographic signal in step S300 includes:

[0145] The pre-processed facial electromyographic signal is processed by short-time Fourier transform to obtain the time-frequency characteristic energy spectrum of facial paralysis;

[0146] Specifically, it can be expressed as:

[0147] Assume the original signal is x(t), the mathematical expression of STFT is:

[0148]

[0149] Where X(t, f) represents the local time-frequency spectrum of the original signal x(t) at time t and frequency f, ω(·) represents the window function, and e -j2πfτ Fourier basis functions representing frequency components.

[0150] The time-frequency characteristic energy spectrum of facial paralysis is further obtained:

[0151]

[0152] Here, |·| represents an absolute value.

[0153] Because the activities of some muscles of patients with facial paralysis may be unbalanced or slow to react when they smile or make other movements, it is necessary to calculate the dynamic spectrum difference;

[0154] Based on the time-frequency characteristic energy spectrum of facial paralysis, the dynamic spectrum difference is calculated as the second electromyographic feature;

[0155] Specifically, it can be expressed as:

[0156]

[0157] Among them, E n(f) represents the time-frequency characteristic energy spectrum of a normal healthy person, and f1 and f2 represent the lower and upper limits of the spectrum respectively.

[0158] S400: Input the output features of the first model and the electromyographic signal features into the second model, perform feature fusion, obtain a total loss function and output a classification category;

[0159] In the embodiment of the present application, in step S400, the output features of the first model and the electromyographic signal features are input into the second model, feature fusion is performed, a classification loss function is obtained and a classification category is output, which specifically includes:

[0160] The output features of the first model are defined as visual features, expressed as:

[0161] Z={Z1,Z,…,Z i} i=1,2,…,n (26)

[0162] Among them, Z i represents the i-th vector output by the first model, including facial key points, muscle texture features, and dynamic symmetry deviation;

[0163] The myoelectric features are defined as:

[0164] E={E1,E2,…,E i} i=1,2,…,n (27)

[0165] Among them, E i Represents the dynamic spectrum difference feature extracted by short-time Fourier;

[0166] The visual features and electromyographic features are input into the Transformer model based on the attention mechanism. Through two layers of self-attention layers with different features and feedforward network processing, feature weighted fusion is performed and the classification category label is output.

[0167] Specifically, the self-attention mechanism (Multi-Head Self-Attention) of the two layers (visual features and EMG features) of the Transformer model can be set as:

[0168]

[0169] Among them, Q, K, V are query, key and value matrices respectively.

[0170] Set the feed-forward neural network of the Transformer model to:

[0171] FFN(X)=max(0,XW1+b1)W2+b2 (29)

[0172] In each Transformer encoder layer, visual and EMG features are processed by self-attention layers and feed-forward networks to learn higher-level feature representations.

[0173] Assuming that the feature representation of Transformer output is H, the final classification result is:

[0174]

[0175] Among them, W c , b c represents the weights and biases of the classification layer, Represents the predicted class label.

[0176] Exemplarily, the classification label can be expressed as y∈{0, 1, 2, 3}, where 0 represents normal, 1 represents mild facial paralysis, 2 represents moderate facial paralysis, and 3 represents severe facial paralysis.

[0177] The total loss function is obtained based on the similarity measurement function and the classification loss function in the second model.

[0178] Specifically, by comparing the learning model, we can learn the detailed differences between different facial paralysis expressions;

[0179] Define a similarity measurement function to determine whether two feature representations belong to the same category, expressed as:

[0180]

[0181] Among them, y represents the unpredicted category label, D(x i , x j ) represents the feature vector x i 、x j Euclidean distance, · is a hyperparameter indicating the minimum distance boundary, and max(·) indicates the maximum value.

[0182] The classification loss function is:

[0183]

[0184] The total loss function is:

[0185] L=L classi +λL contra (33)

[0186] S500: Input the acquired features and the total loss function into the second model to obtain the classification probability, which is used to update the second model output.

[0187] In the embodiment of the present application, in step S500, the acquired features and the total loss function are input into the second model to obtain the classification probability for updating the second model output, including: inputting the acquired features and the total loss function into the second model for fusion, which is expressed as:

[0188] F(t)=[S, RMS, ΔD, ΔE(f), L]

[0189] Where S represents the first facial feature, RMS represents the first electromyographic feature, ΔD represents the dynamic symmetry deviation, ΔE(f) represents the dynamic spectrum difference, and L represents the total loss function;

[0190] It should be noted that integrating information from different sources (such as facial image features, electromyographic signal features, dynamic symmetry deviation, etc.) into a unified feature vector can enable the model to understand the input data from a more comprehensive perspective and capture more subtle changes in expression, thereby improving classification accuracy.

[0191] The classification probability is obtained by iterating the second model, which is used to update the output classification of the second model, expressed as:

[0192]

[0193] It should be noted that the comprehensive information is used to optimize the total loss function, which is an indicator that measures the difference between the predicted results and the actual categories. Finally, the model outputs the classification category according to the principle of minimizing the loss function, that is, determining the most likely expression type or other related labels.

[0194] Example 2

[0195] The above is a schematic scheme of a dual-mode facial paralysis recognition optimization method of the present embodiment. It should be noted that the technical scheme of the dual-mode facial paralysis recognition optimization system and the technical scheme of the above-mentioned dual-mode facial paralysis recognition optimization method belong to the same conception, and the details of the technical scheme of the dual-mode facial paralysis recognition optimization system not described in detail in the present embodiment can all be referred to the description of the technical scheme of the above-mentioned dual-mode facial paralysis recognition optimization method.

[0196] This embodiment also provides a system for a dual-modal facial paralysis recognition optimization method, including:

[0197] An acquisition module, used for acquiring a facial image and a facial electromyographic signal, performing preprocessing, and acquiring a first facial feature and a first electromyographic feature;

[0198] A first feature extraction module is used to input the preprocessed facial image into the first model, perform feature extraction, output facial key points and muscle texture features, obtain expression movement trajectory according to the time series, compare the movement trajectory with the standard template, and output dynamic symmetry deviation;

[0199] A second feature extraction module, used for extracting a second electromyographic feature of the preprocessed facial electromyographic signal;

[0200] A feature fusion module, used for inputting the output features of the first model and the electromyographic signal features into the second model, performing feature fusion, obtaining a total loss function and outputting a classification category;

[0201] The updating module is used to input the acquired features and the total loss function into the second model to obtain the classification probability, which is used to update the output of the second model.

[0202] This embodiment also provides an electronic device suitable for the case of dual-modal facial paralysis recognition optimization, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the dual-modal facial paralysis recognition optimization method proposed in the above embodiment.

[0203] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the dual-modal facial paralysis recognition optimization method proposed in the above embodiment is implemented.

[0204] The storage medium proposed in this embodiment and the method for realizing dual-modal facial paralysis recognition optimization proposed in the above embodiment belong to the same inventive concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0205] Example 3

[0206] Reference Figure 3-Figure 4 Based on the previous embodiment, this embodiment provides an application case of a dual-modal facial paralysis recognition optimization method to illustrate the feasibility and beneficial effects of our solution.

[0207] like Figure 3 The figure shows the dynamic wavelet transform spectrum energy diagram of the electromyographic signal. Figure 2 The dynamic spectrum energy results were calculated to obtain the following characteristics: ① Root mean square (RMS) value: 0.0404; ② Average value of dynamic spectrum difference: 0.1028; ③ Root mean square (RMS) of healthy people: 0.0367; ④ Root mean square (RMS) of patients with facial paralysis: 0.0613; ⑤ Average value of dynamic spectrum difference of healthy people: 0.0527; ⑥ Average value of dynamic spectrum difference of patients with facial paralysis: 0.2604;

[0208] Based on the above features, we input them into the second model and get Figure 3 The prediction loss function training effect;

[0209] from Figure 4From the loss function training results, it can be seen that the optimization method designed by the present invention obtains a lower final loss value, indicating that the model can accurately predict four types of people: normal, mild facial paralysis, moderate facial paralysis, and severe facial paralysis.

[0210] Through the above description of the implementation methods, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course can also be implemented by hardware. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ReadOnly, Memory, ROM), random access memory (Random Access Memory, RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform the methods of various embodiments of the present invention.

[0211] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A dual-modal facial paralysis recognition optimization method, characterized in that: include: Acquire a facial image and a facial electromyographic signal, perform preprocessing, and acquire a first facial feature and a first electromyographic feature; The preprocessed facial image is input into the first model to extract features, output facial key points and muscle texture features, obtain the expression movement trajectory according to the time series, compare the movement trajectory with the standard template, and output the dynamic symmetry deviation; Extracting a second electromyographic feature of the preprocessed facial electromyographic signal; The output features of the first model and the electromyographic signal features are input into the second model to perform feature fusion, obtain the total loss function and output the classification category; The acquired features and the total loss function are input into the second model to obtain the classification probability, which is used to update the output of the second model.

2. The dual-modal facial paralysis recognition optimization method according to claim 1, characterized in that: Get the first facial feature, including: Define the left and right coordinates of a pair of facial corner mirror points, and calculate the Euclidean distance of the left and right coordinates; Define the vector angles of the left and right coordinates about the symmetry axis respectively, which are used to calculate the symmetry angle deviation; Define an arc formed by fitting discrete coordinates within a preset range of the mouth corner or eyebrow arch, and calculate the curvature of the arc; Define the initial key points of the face and the key points of the expression movements when smiling or closing the eyes, and calculate the amplitude of expression changes to obtain the left-right change rate of symmetry; The curvature of the arc and the left-right change rate of the symmetry are used as constraints, and the facial symmetry features are calculated based on the Euclidean distance of the left and right coordinates and the symmetry angle deviation.

3. The dual-modal facial paralysis recognition optimization method according to claim 2, characterized in that: Inputting the preprocessed facial image into the first model, performing feature extraction, and outputting facial key points and muscle texture features, including: inputting the facial image into a CNN model including an LSTM architecture and a DTW algorithm; After multiple layers of convolution, pooling, and fully connected layers, the model outputs the predicted coordinates of n facial key points. The contrast and homogeneity features of muscle texture features are extracted based on the gray-level co-occurrence matrix.

4. The dual-modal facial paralysis recognition optimization method as claimed in claim 3, characterized in that: Acquire the expression movement trajectory according to the time series, compare the movement trajectory with the standard template, and output the dynamic symmetry deviation, including: input the extracted facial key points and muscle texture features as time series data into the LSTM architecture in the first model to obtain the expression movement trajectory; Compare the acquired expression movement trajectory with the expression standard template, and use the DTW algorithm to align the time and amplitude, including: Suppose the set of m standard templates is U = {U1,U2,…,U m }; Suppose the set of expression trajectories of n patients is X = {X1, X2, …, X n }; Assume that the goal of DTW is the optimal matching path P = {(i1, j1), (i2, j2), ..., (i k ,j k )}; Aligning U and X in time and amplitude, we get the DTW distance metric, expressed as: Among them, i∈m,j∈n; Calculate and output the dynamic symmetry deviation between the patient and the standard template, including: Combined with the DTW distance metric, the similarity based on the synchronization of time series and the change of feature amplitude is calculated, which is expressed as: Where T represents the total number of facial paralysis patients X and standard templates U at time step t; If the similarity exceeds the set threshold, it is considered that there is a dynamic symmetry deviation of facial paralysis, and the similarity is output as a dynamic symmetry deviation value.

5. The dual-modal facial paralysis recognition optimization method according to claim 4, characterized in that: Extracting the second electromyographic feature of the preprocessed facial electromyographic signal includes: The pre-processed facial electromyographic signal is processed by short-time Fourier transform to obtain the time-frequency characteristic energy spectrum of facial paralysis; Based on the time-frequency characteristic energy spectrum of facial paralysis, the dynamic spectrum difference is calculated as the second electromyographic feature.

6. The dual-modal facial paralysis recognition optimization method according to claim 5, characterized in that: The output features of the first model and the electromyographic signal features are input into the second model for feature fusion to obtain the classification loss function and output the classification category, which specifically includes: The output features of the first model are defined as visual features, expressed as: Z={Z1,Z,…,Z i } i=1,2,…,n Among them, Z i represents the i-th vector output by the first model, including facial key points, muscle texture features, and dynamic symmetry deviation; The myoelectric features are defined as: E={E1,E2,…,E i } i=1,2,…,n Among them, E i Represents the dynamic spectrum difference feature extracted by short-time Fourier; The visual features and electromyographic features are input into the Transformer model based on the attention mechanism. Through two layers of self-attention layers with different features and feedforward network processing, feature weighted fusion is performed and the classification category label is output. The total loss function is obtained based on the similarity measurement function and the classification loss function in the second model.

7. The dual-modal facial paralysis recognition optimization method according to claim 6, characterized in that: Input the acquired features and the total loss function into the second model to obtain the classification probability, which is used to update the output of the second model, including: inputting the acquired features and the total loss function into the second model for fusion, which is expressed as: F(t)=[S,RMS,ΔD,ΔE(f),L] Among them, S represents the first facial feature, RMS represents the first electromyographic feature, ΔD represents the dynamic symmetry deviation, ΔE(f) represents the dynamic spectrum difference, and L represents the total loss function; The classification probability is obtained by iterating the second model, which is used to update the output classification of the second model, expressed as:

8. A system applied to the dual-modal facial paralysis recognition optimization method according to any one of claims 1 to 7, characterized in that: include: An acquisition module, used for acquiring a facial image and a facial electromyographic signal, performing preprocessing, and acquiring a first facial feature and a first electromyographic feature; A first feature extraction module is used to input the preprocessed facial image into the first model, perform feature extraction, output facial key points and muscle texture features, obtain expression movement trajectory according to the time series, compare the movement trajectory with the standard template, and output dynamic symmetry deviation; A second feature extraction module, used for extracting a second electromyographic feature of the preprocessed facial electromyographic signal; A feature fusion module, used for inputting the output features of the first model and the electromyographic signal features into the second model, performing feature fusion, obtaining a total loss function and outputting a classification category; The updating module is used to input the acquired features and the total loss function into the second model to obtain the classification probability, which is used to update the output of the second model.

9. An electronic device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the dual-modal facial paralysis recognition optimization method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the dual-modal facial paralysis recognition optimization method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Facial paralysis grading diagnosis method and device based on artificial intelligence

    CN112768065A

  • Facial paralysis level evaluation method based on dynamic region quantitative index

    CN113053517A

  • Multi-modal emotion recognition method based on cross attention mechanism

    CN116311423A

  • Multi-modal emotion recognition method and system based on regularization fusion

    CN118656745A

Cited By

  • Facial paralysis grading method, system and equipment fusing multi-modal data and medium

    CN120876479A