AD screening method based on voice and expression feature multi-modal fusion
By aligning cross-modal features and optimizing the neural network architecture, the alignment and computational complexity issues in the fusion of speech and facial expression data are resolved, enabling comprehensive capture of Alzheimer's disease patient characteristics, improving the accuracy and robustness of screening, and making it suitable for practical medical scenarios.
Patent Information
- Application Number
- CN202510827419.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, single-modal data is difficult to fully reflect the complex characteristics of Alzheimer's patients. Alignment of voice and facial expression data is difficult, computational complexity is high, and model robustness is insufficient, resulting in limited accuracy and reliability of screening results.
By introducing cross-modal feature alignment technology, bottleneck structure, and spatial-channel joint attention module, combined with deep belief network and optimized neural network architecture, multimodal fusion of speech and facial expression features is achieved, dynamically solving data alignment and fusion problems, reducing computational complexity, and enhancing the robustness and generalization ability of the model.
It achieves comprehensive capture of the complex characteristics of Alzheimer's disease patients, improves the accuracy and reliability of early screening, reduces computational complexity, is suitable for application in real medical scenarios, and has the accuracy to identify AD in the early stages.
Smart Images

Figure CN120954741A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and medical diagnostics, and in particular to an AD screening method based on multimodal fusion of voice and facial expression features. Background Technology
[0002] Alzheimer's disease (AD) is a common neurodegenerative disease, and early diagnosis is crucial for slowing disease progression and improving patients' quality of life. In recent years, with the development of artificial intelligence technology, AD screening methods based on speech and behavioral analysis have gradually gained attention. These methods identify early signs of AD by analyzing patients' speech characteristics (such as speech rate, tone, and pause frequency) and behavioral characteristics (such as facial expressions and body movements). However, existing technologies have the following limitations:
[0003] Limitations of single-modal data: While speech analysis can capture language impairment, it cannot reflect nonverbal behavioral characteristics such as facial expressions and emotional responses; while facial expression image analysis can capture emotional and behavioral abnormalities, it cannot reflect changes in language ability. Therefore, single-modal data cannot comprehensively reflect the complex characteristics of AD patients, resulting in limited comprehensiveness and accuracy of screening results.
[0004] Data alignment and fusion issues: Speech data is a high-sampling-rate time-series signal, while face image data is a low-sampling-rate spatial feature, making direct alignment difficult. Furthermore, multimodal data contains numerous irrelevant features, increasing computational complexity and noise interference, thus impacting model performance and efficiency.
[0005] High computational complexity and insufficient model robustness: The high-dimensional features in multimodal data increase computational complexity, leading to inefficient model training and inference, making it difficult to widely apply in real-world medical scenarios. Furthermore, during multimodal fusion, information conflicts may exist between data from different modalities, resulting in insufficient model robustness and affecting the reliability of screening results. In addition, the imbalance in the number of samples from different categories (such as AD, MCI, SCD, and normal individuals) in existing datasets affects the model's generalization ability and classification performance. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides an AD screening method based on multimodal fusion of speech and facial expression features. The aim is to comprehensively capture the complex characteristics of AD patients by combining multimodal information from speech and facial expression images, thereby improving the accuracy and reliability of early screening. By proposing a nonlinear transformation layer based on a bottleneck structure and a spatial-channel joint attention (SCJA) module, the invention effectively solves the noise problem in multimodal data alignment and fusion, improving the model's robustness. Through an optimized neural network architecture and fusion strategy, the invention reduces parameter computation and computational complexity, improving the model's training and inference efficiency, making it suitable for widespread application in real-world medical scenarios. Finally, by using a probabilistic generative model based on Deep Belief Networks (DBN) and decision consistency constraints, the invention dynamically resolves intermodal conflict problems, enhancing the model's adaptability and robustness across multiple scenarios.
[0007] The technical solution of this invention is implemented as follows:
[0008] An AD screening method based on multimodal fusion of speech and facial expression features includes the following steps:
[0009] S1. Perform keyframe extraction, face detection and screening, and image standardization on face image data; simultaneously, extract Fbank features from speech data; introduce data augmentation techniques, including time-frequency masking of speech and random cropping and flipping of face images;
[0010] S2. For the speech modality, the speech features are extracted using the local and global feature fusion module. For the face modality, the temporal information is captured through a multi-level dynamic feature interaction network, and important features are highlighted by the adaptive hierarchical attention fusion module. The discriminativeness of the features is improved by the discriminative embedding generation module.
[0011] Before multimodal fusion, cross-modal feature alignment technology is introduced. By learning a shared feature space, speech and facial features are ensured to have better consistency in the feature space, thereby providing higher quality input for subsequent fusion.
[0012] S3. Perform feature extraction and classification label prediction independently for each modality of data, and dynamically optimize the feature extraction path through deep reinforcement learning;
[0013] S4. Based on the confidence scores of each modality, use fusion strategies such as weighted average, Bayesian fusion, or deep belief network to integrate the feature information of speech and facial expression images to generate the final classification decision.
[0014] S5. Use the AdamW optimizer in conjunction with WarmupCosineSchedulerLR to train the model and dynamically adjust the learning rate to improve training efficiency.
[0015] S6. Introduce the ArcFace loss function enhanced by contrastive learning. Optimize feature embedding through positive and negative sample contrast learning to further improve the model's classification performance and generalization ability.
[0016] Step S1 includes:
[0017] Isochronous sampling was performed using a computer vision library;
[0018] Preliminary face localization is performed using a CNN face detector, achieved through a pre-trained facial feature point detection model;
[0019] Image processing techniques are applied for geometric correction, uniform scaling to a specified pixel size, grayscale normalization, and filtering of non-subject face data.
[0020] Speech features are extracted using a filter bank Fbank, simulating human auditory perception, and spectral information reflecting speech features is extracted.
[0021] Step S2 includes:
[0022] An attention feature fusion module is used to fuse adjacent feature maps within the residual block;
[0023] A bottom-up approach is used to modulate features at different time scales;
[0024] The geometric deformation characteristics of the facial features are captured by deformable convolution kernel groups, and the micro-feature response is enhanced by combining depthwise separable convolution and coordinate attention modules.
[0025] Design a recursive attention propagation mechanism to solve the cross-layer feature mismatch problem through bidirectional feature flow;
[0026] By employing a multi-level compression excitation module and a bottleneck structure nonlinear transformation layer, feature embedding is optimized, improving intra-class compactness and inter-class separability.
[0027] Step S3 includes:
[0028] The confidence calculation method based on Softmax output performs independent feature extraction and classification prediction on the data of each modality, and obtains the classification label of each modality and its corresponding confidence score.
[0029] Step S4 includes:
[0030] By dynamically adjusting the weights of each modality by incorporating the confidence scores of the classification results of each modality, the fusion effect can be improved.
[0031] Based on Bayes' theorem, by constructing a probabilistic model of the classification results of each modality, multimodal information is integrated to make accurate classification decisions;
[0032] By capturing the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and combining decision consistency constraints to dynamically resolve modality conflicts, classification performance is improved.
[0033] Step S5 includes:
[0034] By using an optimizer, more efficient regularization can be achieved by separating weight decay from gradient updates;
[0035] A learning rate scheduler is used to set a minimum learning rate for learning rate warm-up and training over multiple epochs.
[0036] A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries;
[0037] By using fully connected networks (MLP) and activation functions, computational complexity is reduced and the discriminative ability of features is enhanced.
[0038] By generating a joint attention map, the spatial distribution and channel weights of features are adaptively adjusted to improve the discriminative power of the features;
[0039] By integrating multiple pooling techniques, comprehensive characterization of multi-scale features can be achieved.
[0040] Step S6 includes:
[0041] A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries;
[0042] By comparing positive and negative samples, we can optimize feature embedding.
[0043] Compared with the prior art, the present invention has the following advantages:
[0044] 1. This invention combines multimodal information from speech and facial expression images, using multimodal fusion technology to comprehensively capture the complex characteristics of AD patients. The speech modality can capture the decline in language ability, while the facial expression image modality can capture emotional and behavioral abnormalities. The combination of the two effectively utilizes the complementarity of multimodal information, significantly improving the comprehensiveness and accuracy of screening;
[0045] 2. A cross-modal feature alignment technique is introduced. By learning a shared feature space, better consistency between speech and facial features is ensured in the feature space. This not only solves the alignment problem between speech and facial data, but also optimizes feature embedding through multi-level compression excitation modules and bottleneck structure nonlinear transformation layers, further improving the discriminative power of features.
[0046] 3. A deep belief network (DBN) fusion strategy is adopted to capture the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and to dynamically resolve modality conflicts by combining decision consistency constraints. Furthermore, a stratified sampling method is used to effectively alleviate class imbalance, ensuring the statistical reliability of experimental results and improving the model's generalization ability.
[0047] 4. By introducing a bottleneck structure and a spatial-channel joint attention module, the amount of parameter computation is reduced, computational complexity is lowered, and the training and inference efficiency of the model is improved. Furthermore, a filter bank (Fbank) is used to extract speech features, and depthwise separable convolution and a coordinate attention module are combined to enhance micro-feature responses, further improving the efficiency of feature extraction.
[0048] 5. This invention can accurately identify AD in the early clinical stages (such as SCD and MCI), providing a basis for early intervention and having important clinical significance. Through multimodal fusion technology and optimized neural network architecture, this invention can effectively improve the comprehensiveness and accuracy of screening, and truly achieve accurate identification in the early clinical stages of AD. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the AD screening method based on multimodal fusion of voice and facial expression features according to Embodiment 1 of this application;
[0050] Figure 2 This is a flowchart illustrating the AD screening method based on multimodal fusion of voice and facial expression features in Embodiment 2 of this application. Detailed Implementation
[0051] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0052] Example 1
[0053] like Figure 1As shown, this embodiment provides an AD screening method based on multimodal fusion of voice and facial expression features, characterized by the following steps:
[0054] S1. Perform keyframe extraction, face detection and screening, and image standardization on face image data; simultaneously, extract Fbank features from speech data; introduce data augmentation techniques, including time-frequency masking of speech and random cropping and flipping of face images;
[0055] S2. For the speech modality, the speech features are extracted using the local and global feature fusion module. For the face modality, the temporal information is captured through a multi-level dynamic feature interaction network, and important features are highlighted by the adaptive hierarchical attention fusion module. The discriminativeness of the features is improved by the discriminative embedding generation module.
[0056] Before multimodal fusion, cross-modal feature alignment technology is introduced. By learning a shared feature space, speech and facial features are ensured to have better consistency in the feature space, thereby providing higher quality input for subsequent fusion.
[0057] S3. Perform feature extraction and classification label prediction independently for each modality of data, and dynamically optimize the feature extraction path through deep reinforcement learning;
[0058] S4. Based on the confidence scores of each modality, use fusion strategies such as weighted average, Bayesian fusion, or deep belief network to integrate the feature information of speech and facial expression images to generate the final classification decision.
[0059] S5. Use the AdamW optimizer in conjunction with WarmupCosineSchedulerLR to train the model and dynamically adjust the learning rate to improve training efficiency.
[0060] S6. Introduce the ArcFace loss function enhanced by contrastive learning. Optimize feature embedding through positive and negative sample contrast learning to further improve the model's classification performance and generalization ability.
[0061] Further, step S1 includes:
[0062] Use a computer vision library for equidistant sampling to reduce data redundancy;
[0063] Preliminary face localization is performed using a CNN face detector, achieved through a pre-trained facial feature point detection model;
[0064] Image processing techniques are applied for geometric correction, uniform scaling to a specified pixel size, grayscale normalization, and filtering of non-subject face data.
[0065] Speech features are extracted using a filter bank Fbank, simulating human auditory perception, and spectral information reflecting speech features is extracted.
[0066] Specifically, the DLib computer vision library is used to perform equidistant sampling at intervals of 30 frames to reduce data redundancy;
[0067] Initial face localization was performed using Dlib's CNN face detector, achieved through a pre-trained 68-point facial feature point detection model.
[0068] Affine transformations from OpenCV were used for geometric correction, uniformly scaled to (47,55) pixel size, grayscale normalized, and non-subject face data were filtered out.
[0069] Speech features are extracted using a filter bank Fbank (Filter Bank Features), simulating human auditory perception, and spectral information reflecting speech features is extracted.
[0070] For example, input: a video file (e.g., a 300-frame video), extract keyframes at 30-frame intervals to reduce data redundancy:
[0071] frame_indices={i∣i=0,30,60,…,video_length}
[0072] Where video_length is the total number of frames in the video; by extracting one frame every 30 frames, the amount of data can be effectively reduced while retaining the key information of the video;
[0073] The keyframes extracted from the input are used to detect faces and extract 68 facial feature points. A Dlib pre-trained model is used. Dlib's CNN face detector can quickly locate face regions, while the 68-point feature point detection model can accurately extract facial key points (such as eyes, nose, mouth, etc.), providing a foundation for subsequent image processing.
[0074] Input the detected face region and feature points, perform geometric correction on the face, uniformly scale it to a specified size (47×55 pixels), and perform grayscale normalization;
[0075] Affine transformation, by calculating the positions of facial key points (such as the eyes and nose), determines rotation and translation parameters to adjust the face to a standard position:
[0076]
[0077] Gray-level normalization normalizes pixel values to the [0,1] range, reducing the impact of lighting conditions on the image.
[0078]
[0079] The purity of the data is ensured by identifying and filtering out facial data that does not belong to the subjects through a pre-trained model.
[0080] Input the speech signal and extract 26-dimensional Fbank features:
[0081] Mel filter banks convert frequencies to the Mel scale, simulating the human ear's perception of different frequencies:
[0082]
[0083] Calculate the power spectrum of the speech signal using Short Time Fourier Transform (STFT);
[0084] Fbank feature extraction uses a Mel filter bank to perform weighted summation of the power spectrum to obtain spectral information reflecting speech features, and then takes a logarithm to enhance the dynamic range of the features.
[0085] Fbank=log(Mel_filter_bank×Power_spectrum).
[0086] Further, step S2 includes:
[0087] An attention feature fusion module is used to fuse adjacent feature maps within the residual block to enhance the audio's discriminative ability;
[0088] A bottom-up approach is used to modulate features at different time scales, thereby improving the robustness of audio embedding from a global perspective.
[0089] The geometric deformation characteristics of the facial features are captured by deformable convolution kernel groups, and the micro-feature response is enhanced by combining depthwise separable convolution and coordinate attention modules.
[0090] Design a recursive attention propagation mechanism to solve the cross-layer feature mismatch problem through bidirectional feature flow;
[0091] By employing a multi-level compression excitation module and a bottleneck structure nonlinear transformation layer, feature embedding is optimized, improving intra-class compactness and inter-class separability.
[0092] Specifically, an attention feature fusion module is used to fuse adjacent feature maps within the residual block to enhance the audio's discriminative ability;
[0093] A bottom-up approach is used to modulate features at different time scales, thereby improving the robustness of audio embedding from a global perspective.
[0094] The geometric deformation characteristics of the facial features are captured by deformable convolution kernel groups, and the micro-feature response is enhanced by combining depthwise separable convolution and coordinate attention modules.
[0095] A recursive attention propagation mechanism is designed to solve the cross-layer feature mismatch problem through bidirectional feature flow (bottom-up Transformer modeling of topological relationships, top-down channel recalibration gate modulation of low-level responses);
[0096] By employing a multi-level compression excitation (MCSE) module and a bottleneck structure nonlinear transformation layer, feature embedding is optimized to improve intra-class compactness and inter-class separability.
[0097] For example, input speech feature maps (e.g., outputs from intermediate layers of a convolutional neural network) are weighted and fused with adjacent feature maps within the residual block through an attention mechanism to enhance the discriminative power of audio features.
[0098] Attention weight calculation:
[0099]
[0100] Among them, F i is the i-th feature map, N is the total number of feature maps, and score is a learnable scoring function, usually a simple linear transformation or MLP;
[0101] Weighted feature fusion:
[0102]
[0103] By dynamically adjusting the weight of each feature map through an attention mechanism, more important features are given greater weight in the fusion process, thereby enhancing the discriminative power of audio features.
[0104] Input speech features at different time scales (e.g., short-time features and long-time features), and modulate the features at different time scales in a bottom-up manner to improve the robustness of global features;
[0105] Feature modulation:
[0106] F modulated =UpSample(F short )☉DownSample(F long )
[0107] Among them, F short It is a short-term characteristic, F long These are long-term features. UpSample and DownSample are upsampling and downsampling operations, respectively. ⊙ represents element-wise multiplication.
[0108] By upsampling short-term features and multiplying them with the result of downsampling long-term features, the modulation of features at different time scales is achieved, thereby improving the robustness of audio embedding from a global perspective.
[0109] Input facial image feature mapping, capture the geometric deformation characteristics of facial features through deformable convolution kernels, and enhance the micro-feature response by combining depthwise separable convolution and coordinate attention module;
[0110] Deformable convolution:
[0111]
[0112] Among them, w i The weights of the deformable convolution kernel, Δx i and Δy i I is the offset, and I is the input feature map.
[0113] Depthwise separable convolution:
[0114] F sep =DepthwiseConv(F deform PointwiseConv(F) deform )
[0115] Coordinate attention:
[0116] α(x,y)=Sigmoid(W·[x,y])
[0117] F att =F sep ☉α(x,y)
[0118] Deformable convolution adapts to the geometric changes of facial features by dynamically adjusting the shape of the convolution kernel; depthwise separable convolution reduces computational cost; coordinate attention modules further enhance feature responses by introducing positional information.
[0119] The input multi-layer feature map is used to solve the cross-layer feature mismatch problem through a recursive attention propagation mechanism combined with bidirectional feature flow;
[0120] Bottom-up feature propagation:
[0121]
[0122] Top-down feature modulation:
[0123]
[0124] Recursive attention propagation:
[0125]
[0126] By modeling topological relationships using a bottom-up Transformer and modulating the underlying response using a top-down channel recalibration gate, bidirectional feature flow can effectively solve the problem of cross-layer feature mismatch and improve the hierarchical nature of features.
[0127] The optimized feature map is input, and the feature embedding is optimized through a multi-level compression excitation module and a bottleneck structure nonlinear transformation layer to improve intra-class compactness and inter-class separability.
[0128] Multi-level compression excitation (MCSE):
[0129]
[0130] Among them, GAP is global mean pooling, GMP is global maximum pooling, and RSP is regional saliency pooling.
[0131] Nonlinear transformation of bottleneck structure:
[0132] F bottleneck =MLP(ReLU(FC(F) MCSE )))
[0133] Where FC stands for fully connected layer and ReLU stands for activation function;
[0134] The multi-level compression excitation module enhances the representation ability of features by integrating global and local information; the bottleneck structure nonlinear transformation layer improves the discriminative ability of features by reducing the feature dimension, thereby optimizing feature embedding.
[0135] Further, step S3 includes:
[0136] The confidence calculation method based on Softmax output performs independent feature extraction and classification prediction on the data of each modality, and obtains the classification label of each modality and its corresponding confidence score.
[0137] For example, input the feature representations of each modality (e.g., voice features and face features), perform independent classification predictions on the data of each modality, and calculate the classification label and its corresponding confidence score;
[0138] The softmax function performs classification prediction on the features of each modality, and is typically implemented using a fully connected layer and the softmax function.
[0139]
[0140] Among them, z i is the score of the i-th category, and K is the total number of categories; the Softmax output is a probability distribution representing the confidence level of each category;
[0141] Confidence score: Confidence = max(Softmax(z)); The final confidence score is the maximum value in the Softmax output, and the corresponding category is the predicted class label.
[0142] Further, step S4 includes:
[0143] By dynamically adjusting the weights of each modality by incorporating the confidence scores of the classification results of each modality, the fusion effect can be improved.
[0144] Based on Bayes' theorem, by constructing a probabilistic model of the classification results of each modality, multimodal information is integrated to make accurate classification decisions;
[0145] By capturing the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and combining decision consistency constraints to dynamically resolve modality conflicts, classification performance is improved.
[0146] Specifically, the weights of each modality are dynamically adjusted by incorporating the confidence scores of the classification results of each modality, thereby improving the fusion effect;
[0147] Based on Bayes' theorem, by constructing a probabilistic model of the classification results of each modality, multimodal information is integrated to make accurate classification decisions;
[0148] By capturing the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and combining decision consistency constraints to dynamically resolve modality conflicts, classification performance is improved.
[0149] For example, input the classification labels of each modality and their corresponding confidence scores, and dynamically adjust the weights of each modality based on the confidence scores to improve the fusion effect;
[0150] Weighting is performed such that the weight of each modality is proportional to its confidence score, with higher-confidence modalities receiving greater weight during the fusion process.
[0151]
[0152] Among them, Confidence i is the confidence score of the i-th mode, and M is the total number of modes;
[0153] Weighted fusion combines the classification scores from different modalities using a weighted average method to obtain the final fusion score.
[0154]
[0155] Among them, Score i It is the classification score of the i-th modality;
[0156] Input the classification labels of each modality and their corresponding confidence scores, construct a probability model using Bayes' theorem, integrate multimodal information, and make accurate classification decisions;
[0157] By using Bayes' theorem and combining the likelihood and prior probabilities of each modality, the posterior probability of each class is calculated:
[0158]
[0159] Among them, C k Let X be the k-th category, and P(C) be the input data. k |X) is the posterior probability, P(X|C) is the posterior probability. k P(C) is the likelihood probability. k P(X) is the prior probability, and P(X) is the evidential probability.
[0160] Likelihood probability:
[0161]
[0162] Among them, X i It is a feature of the i-th mode, w i It is the weight of the i-th mode;
[0163] The posterior probability is the final classification decision, which is the category with the highest posterior probability.
[0164]
[0165] Input the feature representations of each modality, capture the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and dynamically resolve the modality conflict problem by combining decision consistency constraints to improve classification performance;
[0166] Unsupervised pre-training learns the intrinsic structure of data through an autoencoder, capturing the dependencies between modalities:
[0167] Pretrained_Model = Autoencoder(X); where X is the input data and Autoencoder is the autoencoder model;
[0168] Supervised fine-tuning improves classification performance by fine-tuning the pre-trained model using a classifier.
[0169] FineTuned_Model = Classifier(Pretrained_Model); where Classifier is the classifier model;
[0170] Decision consistency constraints dynamically resolve modality conflicts by minimizing the differences in classification probability distributions between different modalities.
[0171]
[0172] Here, KL is the Kullback-Leibler divergence, which measures the difference between two probability distributions.
[0173] Further, step S5 includes:
[0174] By using an optimizer, more efficient regularization can be achieved by separating weight decay from gradient updates;
[0175] A learning rate scheduler is used to set a minimum learning rate for learning rate warm-up and training over multiple epochs.
[0176] A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries;
[0177] By using fully connected networks (MLP) and activation functions, computational complexity is reduced and the discriminative ability of features is enhanced.
[0178] By generating a joint attention map, the spatial distribution and channel weights of features are adaptively adjusted to improve the discriminative power of the features;
[0179] By integrating multiple pooling techniques, comprehensive characterization of multi-scale features can be achieved.
[0180] Specifically, the AdamW optimizer is used to achieve more efficient regularization by separating weight decay from gradient updates;
[0181] WarmupCosineSchedulerLR was used, with a minimum learning rate of 1e-5, a warm-up epoch of 10 learning rate, and training for 60-160 epochs.
[0182] The ArcFace loss function is adopted, and by introducing additional angular margins at the classification boundaries, the compactness of intra-class features and the separability of inter-class features are enhanced.
[0183] By using a two-layer fully connected network (MLP) and a Gaussian error linear unit (GELU) activation function, the computational complexity is reduced and the discriminative ability of features is enhanced.
[0184] By generating a joint attention map of spatial and channel dimensions using the Sigmoid function, the spatial distribution and channel weights of features are adaptively adjusted to improve the discriminative power of the features.
[0185] By integrating Global Mean Pooling (GAP), Global Maximum Pooling (GMP), and Regional Significance Pooling (RSP), a comprehensive representation of multi-scale features can be achieved.
[0186] For example, input model parameters can be used to achieve more efficient regularization by separating weight decay from gradient updates;
[0187] The AdamW optimizer achieves more efficient regularization by applying weight decay directly to parameter updates, rather than gradient updates.
[0188]
[0189] Where, m t and v t These are the first and second moments of the gradient, respectively; α is the learning rate; λ is the weight decay coefficient; and ∈ is the numerical stability term.
[0190] This method can avoid the impact of weight decay on the learning rate and improve the training effect of the model;
[0191] Input the learning rate during the training process, dynamically adjust the learning rate through the learning rate scheduler, set the minimum learning rate, perform learning rate warm-up, and train for multiple epochs.
[0192] WarmupCosineSchedulerLR:
[0193]
[0194] Where, α t α is the learning rate for the t-th epoch. init It is the initial learning rate, α min It is the minimum learning rate, T warmup T is the number of warm-up epochs, and T is the total number of training epochs.
[0195] In the early stages of training, the learning rate increases linearly (warm-up phase) to avoid training instability caused by an excessively high initial learning rate.
[0196] After the warm-up phase, the learning rate gradually decreases according to the cosine curve until it reaches the minimum learning rate.
[0197] Specific settings: Minimum learning rate is 10. -5 The number of warm-up epochs is 10, and the total number of training epochs is 60-160.
[0198] The input model's feature embeddings and classification labels are used to enhance the compactness of intra-class features and the separability of inter-class features by introducing additional angular spacing at the classification boundaries.
[0199] The ArcFace loss function, by introducing an additional angular margin m at the classification boundary, makes the features of different classes more distinct in the feature space:
[0200]
[0201] Where N is the total number of samples and K is the total number of categories. is the angle between the i-th sample and its class center, m is the additional angular interval, and s is the scaling factor; the scaling factor s is used to amplify the differences in features and enhance the classification effect;
[0202] Input feature embedding, through fully connected networks (MLP) and activation functions, reduces computational complexity and enhances the discriminative ability of features;
[0203] MLP performs nonlinear transformations on features through two fully connected layers, reducing feature dimensionality while enhancing the discriminative power of the features:
[0204] h = ReLU(W1x + b1)
[0205] y = W2h + b2
[0206] Where x is the input feature, W1 and W2 are weight matrices, b1 and b2 are bias terms, and ReLU is the activation function;
[0207] GELU activation function is a smooth activation function with better gradient propagation properties than ReLU, making it suitable for deep neural networks.
[0208]
[0209] Input feature maps are used to generate joint attention maps, which adaptively adjust the spatial distribution and channel weights of features to improve the discriminative power of features.
[0210] By using joint attention maps, the model can adaptively adjust the spatial distribution and channel weights of features, highlighting important feature regions and improving the discriminative power of features.
[0211] α = Sigmoid(W s [x,y])
[0212] F att =F☉α
[0213] Where x and y are the features of the spatial dimension and the channel dimension, respectively, and W s It is the weight matrix, Sigmoid is the activation function, and ⊙ is element-wise multiplication;
[0214] Input feature maps, through the integration of multiple pooling techniques, achieve comprehensive representation of multi-scale features;
[0215] Global Mean Pooling (GAP):
[0216]
[0217] Global Maximum Pooling (GMP):
[0218]
[0219] Regional Significant Pooling (RSP):
[0220]
[0221] Where H and W are the height and width of the feature map, and α(i,j) is the attention weight;
[0222] Feature fusion:
[0223]
[0224] By integrating global mean pooling, global maximum pooling, and regional saliency pooling, the model can capture multi-scale feature information and enhance the feature representation capability.
[0225] Further, step S6 includes:
[0226] A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries;
[0227] By comparing positive and negative samples, we can optimize feature embedding and improve the classification performance and generalization ability of the model.
[0228] For example, the feature embeddings and classification labels of the input model enhance the compactness of intra-class features and the separability of inter-class features by introducing additional angular spacing at the classification boundaries;
[0229] Contrastive learning loss functions enhance the discriminative power of features by bringing positive samples closer together and widening the distance between negative samples.
[0230]
[0231] Among them, f + It is the feature embedding of positive samples, f i is the feature embedding of the i-th sample, τ is the temperature parameter, and N is the total number of samples;
[0232] Additional angular margin introduces an extra angle at the classification boundary, making features of different classes more separable in the feature space, enhancing the compactness of intra-class features and the separability of inter-class features.
[0233]
[0234] in, is the angle between the i-th sample and its class center, and m is the additional angular interval;
[0235] Input the model's feature embeddings and classification labels, and learn by comparing positive and negative samples to optimize the feature embeddings and improve the model's classification performance and generalization ability;
[0236] By learning through comparison of positive and negative samples, the model can better learn the discriminative power of features, thereby optimizing feature embedding:
[0237]
[0238] Among them, f + It is the feature embedding of positive samples, f i is the feature embedding of the i-th sample, τ is the temperature parameter, and N is the total number of samples;
[0239] Feature optimization improves the model's classification performance and generalization ability by updating feature embeddings based on the gradient of the loss function through gradient descent.
[0240]
[0241] Where f is the feature embedding and α is the learning rate. It is the gradient of the loss function with respect to the features.
[0242] Example 2
[0243] like Figure 2 As shown, this embodiment provides an AD screening method based on multimodal fusion of voice and facial expression features, characterized by the following steps:
[0244] S1. Perform keyframe extraction, face detection and screening, and image standardization on face image data; simultaneously, extract Fbank features from speech data.
[0245] Data augmentation techniques are introduced, including time-frequency masking of speech and random cropping and flipping of face images;
[0246] S2. For the speech modality, the speech features are extracted using the local and global feature fusion module. For the face modality, the temporal information is captured through a multi-level dynamic feature interaction network, and important features are highlighted by the adaptive hierarchical attention fusion module. The discriminativeness of the features is improved by the discriminative embedding generation module.
[0247] Before multimodal fusion, cross-modal feature alignment technology is introduced. By learning a shared feature space, speech and facial features are ensured to have better consistency in the feature space, thereby providing higher quality input for subsequent fusion.
[0248] S3. Perform feature extraction and classification label prediction independently for each modality of data, and dynamically optimize the feature extraction path through deep reinforcement learning;
[0249] S4. Based on the confidence scores of each modality, use fusion strategies such as weighted average, Bayesian fusion, or deep belief network to integrate the feature information of speech and facial expression images to generate the final classification decision.
[0250] S5. Use the AdamW optimizer in conjunction with WarmupCosineSchedulerLR to train the model and dynamically adjust the learning rate to improve training efficiency.
[0251] S6. Introduce the ArcFace loss function enhanced by contrastive learning. Optimize feature embedding through positive and negative sample contrast learning to further improve the model's classification performance and generalization ability.
[0252] Further, step S1 includes:
[0253] Use a computer vision library for equidistant sampling to reduce data redundancy;
[0254] Preliminary face localization is performed using a CNN face detector, achieved through a pre-trained facial feature point detection model;
[0255] Image processing techniques are applied for geometric correction, uniform scaling to a specified pixel size, grayscale normalization, and filtering of non-subject face data.
[0256] Speech features are extracted using a filter bank Fbank, simulating human auditory perception, and spectral information reflecting speech features is extracted.
[0257] Specifically, the DLib computer vision library is used to perform equidistant sampling at intervals of 30 frames to reduce data redundancy;
[0258] Initial face localization was performed using Dlib's CNN face detector, achieved through a pre-trained 68-point facial feature point detection model.
[0259] Affine transformations from OpenCV were used for geometric correction, uniformly scaled to (47,55) pixel size, grayscale normalized, and non-subject face data were filtered out.
[0260] Speech features are extracted using a filter bank Fbank (Filter Bank Features), simulating human auditory perception, and spectral information reflecting speech features is extracted.
[0261] Furthermore, the introduced data augmentation techniques include time-frequency masking of speech and random cropping and flipping of face images, specifically:
[0262] Enhance the robustness of voice data by simulating noise and interference that voice may encounter in real-world environments;
[0263] In the time-domain representation of a speech signal, a continuous time interval is randomly selected and either set to zero or replaced with noise:
[0264]
[0265] t start and t end The masking start and end times are randomly selected. The noise can be Gaussian noise or other types of background noise. For example, 10% of the speech signal's duration can be masked, with the start point randomly selected.
[0266] In the frequency domain representation of a speech signal, a continuous frequency range is randomly selected and either set to zero or replaced with noise:
[0267]
[0268] f start and f end The masking start and end frequencies are randomly selected, and the noise can be Gaussian noise or other types of background noise.
[0269] For example, a 10% frequency range of the speech signal can be selected for masking, with the starting point randomly chosen;
[0270] Random cropping aims to enhance the diversity of image data and improve the model's adaptability to different image regions.
[0271] Randomly select a sub-region from the original image as the cropped image:
[0272] I crop =I[x:x+h,y:y+w]
[0273] (x,y) is the randomly selected cropping starting point, and h and w are the cropping height and width, which are usually adjusted according to the target image size. For example, you can select 80% of the image area for cropping and randomly select the cropping starting point.
[0274] Randomly flipped targets increase data diversity and simulate image changes in different directions;
[0275] The image is flipped horizontally or vertically with a certain probability:
[0276]
[0277] p is the probability of flipping, usually set to 0.5. Horizontal flipping is achieved by reversing the column index of the image, and vertical flipping is achieved by reversing the row index of the image. For example, you can choose a horizontal flip probability of 0.5 and a vertical flip probability of 0.5.
[0278] Further, step S2 includes:
[0279] An attention feature fusion module is used to fuse adjacent feature maps within the residual block to enhance the audio's discriminative ability;
[0280] A bottom-up approach is used to modulate features at different time scales, thereby improving the robustness of audio embedding from a global perspective.
[0281] The geometric deformation characteristics of the facial features are captured by deformable convolution kernel groups, and the micro-feature response is enhanced by combining depthwise separable convolution and coordinate attention modules.
[0282] Design a recursive attention propagation mechanism to solve the cross-layer feature mismatch problem through bidirectional feature flow;
[0283] By employing a multi-level compression excitation module and a bottleneck structure nonlinear transformation layer, feature embedding is optimized, improving intra-class compactness and inter-class separability.
[0284] Specifically, an attention feature fusion module is used to fuse adjacent feature maps within the residual block to enhance the audio's discriminative ability;
[0285] A bottom-up approach is used to modulate features at different time scales, thereby improving the robustness of audio embedding from a global perspective.
[0286] The geometric deformation characteristics of the facial features are captured by deformable convolution kernel groups, and the micro-feature response is enhanced by combining depthwise separable convolution and coordinate attention modules.
[0287] A recursive attention propagation mechanism is designed to solve the cross-layer feature mismatch problem through bidirectional feature flow (bottom-up Transformer modeling of topological relationships, top-down channel recalibration gate modulation of low-level responses);
[0288] By employing a multi-level compression excitation (MCSE) module and a bottleneck structure nonlinear transformation layer, feature embedding is optimized to improve intra-class compactness and inter-class separability.
[0289] Prior to multimodal fusion, a cross-modal feature alignment technique is introduced. By learning a shared feature space, this ensures better consistency between speech and facial features in the feature space, thereby providing higher-quality input for subsequent fusion. Specifically:
[0290] Extract features from speech data, for example, using the Fbank feature extraction method;
[0291] Extracting features from face images, for example, using convolutional neural networks (CNNs);
[0292] Use a multi-modal autoencoder (MMAE) or a shared feature embedding network (SFEN);
[0293] Speech and facial features are encoded separately and mapped to a shared feature space.
[0294] The speech features F voice Mapping to the shared space Z, facial features F face Mapped to shared space Z:
[0295] Z voice =Encoder voice (F voice )
[0296] Z face =Encoder face (F face )
[0297] Use a shared fully connected layer or Transformer network to map speech and facial features to the same feature space:
[0298] Z shared =SFEN(F voice F face )
[0299] F voice and F face These are the feature representations of voice and face, Z. shared It is a shared feature space representation;
[0300] Optimize feature representations in a shared feature space using contrastive learning or a consistency loss function:
[0301]
[0302] Z voice,i and Z face,i It is the representation of the speech and face features of the i-th sample in the shared space, and the contrastive learning loss function (such as InfoNCE loss) is used to optimize the consistency of the features;
[0303] The effectiveness of feature alignment is verified by calculating the similarity or distance between features:
[0304]
[0305] Cosine similarity is used to calculate the similarity between speech and facial features in a shared space. A high similarity value indicates good feature alignment.
[0306] Further, step S3 includes:
[0307] The confidence calculation method based on Softmax output performs independent feature extraction and classification prediction on the data of each modality, and obtains the classification label of each modality and its corresponding confidence score.
[0308] Further, step S4 includes:
[0309] By dynamically adjusting the weights of each modality by incorporating the confidence scores of the classification results of each modality, the fusion effect can be improved.
[0310] Based on Bayes' theorem, by constructing a probabilistic model of the classification results of each modality, multimodal information is integrated to make accurate classification decisions;
[0311] By capturing the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and combining decision consistency constraints to dynamically resolve modality conflicts, classification performance is improved.
[0312] Specifically, the weights of each modality are dynamically adjusted by incorporating the confidence scores of the classification results of each modality, thereby improving the fusion effect;
[0313] Based on Bayes' theorem, by constructing a probabilistic model of the classification results of each modality, multimodal information is integrated to make accurate classification decisions;
[0314] By capturing the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and combining decision consistency constraints to dynamically resolve modality conflicts, classification performance is improved.
[0315] Further, step S5 includes:
[0316] By using an optimizer, more efficient regularization can be achieved by separating weight decay from gradient updates;
[0317] A learning rate scheduler is used to set a minimum learning rate for learning rate warm-up and training over multiple epochs.
[0318] A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries;
[0319] By using fully connected networks (MLP) and activation functions, computational complexity is reduced and the discriminative ability of features is enhanced.
[0320] By generating a joint attention map, the spatial distribution and channel weights of features are adaptively adjusted to improve the discriminative power of the features;
[0321] By integrating multiple pooling techniques, comprehensive characterization of multi-scale features can be achieved.
[0322] Specifically, the AdamW optimizer is used to achieve more efficient regularization by separating weight decay from gradient updates;
[0323] WarmupCosineSchedulerLR was used, with a minimum learning rate of 1e-5, a warm-up epoch of 10 learning rate, and training for 60-160 epochs.
[0324] The ArcFace loss function is adopted, and by introducing additional angular margins at the classification boundaries, the compactness of intra-class features and the separability of inter-class features are enhanced.
[0325] By using a two-layer fully connected network (MLP) and a Gaussian error linear unit (GELU) activation function, the computational complexity is reduced and the discriminative ability of features is enhanced.
[0326] By generating a joint attention map of spatial and channel dimensions using the Sigmoid function, the spatial distribution and channel weights of features are adaptively adjusted to improve the discriminative power of the features.
[0327] By integrating Global Mean Pooling (GAP), Global Maximum Pooling (GMP), and Regional Significance Pooling (RSP), a comprehensive representation of multi-scale features can be achieved.
[0328] Further, step S6 includes:
[0329] A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries;
[0330] By comparing positive and negative samples, we can optimize feature embedding and improve the classification performance and generalization ability of the model.
[0331] The specific embodiments of the invention have been described in detail above, but these are merely examples. The invention is not limited to the specific embodiments described above. Those skilled in the art should understand that the embodiments and descriptions in the specification are only illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. An AD screening method based on multimodal fusion of voice and facial expression features, characterized in that, Includes the following steps: S1. Perform keyframe extraction, face detection and screening, and image standardization on face image data; simultaneously, extract Fbank features from speech data. S2. For the speech modality, the speech features are extracted using the local and global feature fusion module. For the face modality, the temporal information is captured through a multi-level dynamic feature interaction network, and important features are highlighted by the adaptive hierarchical attention fusion module. The discriminativeness of the features is improved by the discriminative embedding generation module. S3. Perform feature extraction and classification label prediction independently for each modality of data, and dynamically optimize the feature extraction path through deep reinforcement learning; S4. Based on the confidence scores of each modality, use fusion strategies such as weighted average, Bayesian fusion, or deep belief network to integrate the feature information of speech and facial expression images to generate the final classification decision. S5. Use the AdamW optimizer in conjunction with WarmupCosineSchedulerLR to train the model and dynamically adjust the learning rate to improve training efficiency. S6. Introduce the ArcFace loss function enhanced by contrastive learning. Optimize feature embedding through positive and negative sample contrast learning to further improve the model's classification performance and generalization ability.
2. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 1, characterized in that: Step S1 includes: Isochronous sampling was performed using a computer vision library; Preliminary face localization is performed using a CNN face detector, achieved through a pre-trained facial feature point detection model; Image processing techniques are applied for geometric correction, uniform scaling to a specified pixel size, grayscale normalization, and filtering of non-subject face data. Speech features are extracted using a filter bank Fbank, simulating human auditory perception, and spectral information reflecting speech features is extracted.
3. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 1, characterized in that: Step S2 includes: An attention feature fusion module is used to fuse adjacent feature maps within the residual block; A bottom-up approach is used to modulate features at different time scales; The geometric deformation characteristics of the facial features are captured by deformable convolution kernel groups, and the micro-feature response is enhanced by combining depthwise separable convolution and coordinate attention modules. Design a recursive attention propagation mechanism to solve the cross-layer feature mismatch problem through bidirectional feature flow; By employing a multi-level compression excitation module and a bottleneck structure nonlinear transformation layer, feature embedding is optimized, improving intra-class compactness and inter-class separability.
4. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 1, characterized in that: Step S3 includes: The confidence calculation method based on Softmax output performs independent feature extraction and classification prediction on the data of each modality, and obtains the classification label of each modality and its corresponding confidence score.
5. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 1, characterized in that: Step S4 includes: By dynamically adjusting the weights of each modality by incorporating the confidence scores of the classification results of each modality, the fusion effect can be improved. Based on Bayes' theorem, by constructing a probabilistic model of the classification results of each modality, multimodal information is integrated to make accurate classification decisions; By capturing the dependencies between modalities through unsupervised pre-training and supervised fine-tuning, and combining decision consistency constraints to dynamically resolve modality conflicts, classification performance is improved.
6. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 1, characterized in that: Step S5 includes: By using an optimizer, more efficient regularization can be achieved by separating weight decay from gradient updates; A learning rate scheduler is used to set a minimum learning rate for learning rate warm-up and training over multiple epochs. A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries; By using fully connected networks (MLP) and activation functions, computational complexity is reduced and the discriminative ability of features is enhanced. By generating a joint attention map, the spatial distribution and channel weights of features are adaptively adjusted to improve the discriminative power of the features; By integrating multiple pooling techniques, comprehensive characterization of multi-scale features can be achieved.
7. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 1, characterized in that: Step S6 includes: A contrastive learning-enhanced loss function is adopted, which enhances the compactness of intra-class features and the separability of inter-class features by introducing additional angular margins at the classification boundaries; By comparing positive and negative samples, we can optimize feature embedding.
8. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 2, characterized in that: Step S1 further includes: Data augmentation techniques are introduced, including time-frequency masking of speech and random cropping and flipping of face images.
9. The AD screening method based on multimodal fusion of speech and facial expression features according to claim 3, characterized in that: In step S2, before multimodal fusion, cross-modal feature alignment technology is introduced. By learning a shared feature space, speech and facial features are ensured to have better consistency in the feature space, thereby providing higher quality input for subsequent fusion.