A mental health detection method and system based on heart rate audio fusion

By aligning heart rate and audio signals in a timely manner and performing cross-modal fusion, the reliability and adaptability issues of mental health testing in home settings have been resolved, achieving high-precision personalized mental health testing.

CN122440189APending Publication Date: 2026-07-24ZHEJIANG UNIVERSITY OF MEDIA AND COMMUNICATIONS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIVERSITY OF MEDIA AND COMMUNICATIONS
Filing Date
2026-02-27
Publication Date
2026-07-24

Smart Images

  • Figure CN122440189A_ABST
    Figure CN122440189A_ABST
Patent Text Reader

Abstract

The application provides a mental health detection method and system based on heart rate audio fusion, comprising: time sequence alignment of synchronously collected original audio signals and original heart rate signals to obtain an original audio-heart rate signal group, calculation of corresponding emotion correlation scores, cross-modal mutual verification and weighted noise suppression of the original audio-heart rate signal group to obtain a denoised audio-heart rate signal group; parallel feature extraction of the denoised audio-heart rate signal group, emotion mutual information correlation and feature enhancement to obtain an enhanced audio-heart rate feature group; phase synchronization of the enhanced audio-heart rate feature group to obtain a synchronous audio-heart rate feature group, cross-modal fusion based on a reconstructed attention matrix to obtain cross-modal emotion features; construction of a mental health model according to a pre-collected bimodal data set, membership calculation and emotion analysis of the cross-modal emotion features to obtain a detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mental health technology, and in particular to a mental health detection method and system based on heart rate audio fusion. Background Technology

[0002] Mental health testing is a crucial link in preventing and intervening in mental health problems and improving the public's mental health level, and it has significant value in areas such as home-based elderly care and clinical auxiliary diagnosis. However, current mainstream mental health testing methods still face significant technical challenges and reliability bottlenecks when applied to specific scenarios such as home-based elderly care.

[0003] First, traditional methods primarily rely on self-reporting scales or subjective interviews. These methods are not only susceptible to the influence of the subjects' subjective will, cognitive level, and social approval effect, leading to distorted results, but also struggle to achieve continuous and dynamic monitoring. Their effectiveness is particularly insufficient for elderly individuals living at home who express themselves implicitly or experience cognitive decline. Second, while current technologies have begun to attempt objective analysis using single physiological signals (such as heart rate or audio signals), they often overlook the inherent correlations and asynchronicities between multimodal information. For example, there is a significant physiological delay between audio changes triggered by emotional stimuli and the physiological response of heart rate; simple feature splicing fails to capture this deep correlation, leading to misjudgments of emotional states. Furthermore, various background noises introduced by complex home environments (such as television sounds and conversations) can severely interfere with audio emotional features, and occasional physical activity by the elderly can cause abnormal fluctuations in heart rate signals. Existing methods lack effective cross-modal verification mechanisms, cannot adaptively distinguish between emotion-related signals and irrelevant noise, and have weak anti-interference capabilities. Finally, the expression of psychological states exhibits high individual variability and ambiguous boundaries. General-purpose emotion recognition models are difficult to adapt to the unique physiological and behavioral patterns of different elderly people. They have limited accuracy in recognizing complex and subtle emotional states such as depression and anxiety, and cannot meet the high reliability requirements of personalized health monitoring. Summary of the Invention

[0004] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a mental health detection method and system based on heart rate audio fusion, which has advantages such as strong anti-interference ability, high recognition accuracy, and good personalization adaptation. It solves the problems of low reliability and poor adaptability of mental health status detection in real home scenarios due to environmental noise interference, asynchronous physiological response, and significant individual differences.

[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: This invention provides a mental health detection method based on heart rate audio fusion, comprising the following steps: The original audio signal and the original heart rate signal were time-aligned to obtain the original audio-heart rate signal group, and the emotion relevance score of the noise signal group in the original audio-heart rate signal group was calculated. Based on the emotion relevance score, cross-modal cross-verification and weighted noise suppression are performed on the original audio-heart rate signal set to obtain a noise-reduced audio-heart rate signal set. Parallel feature extraction and mutual information feature enhancement were performed on the noise-reduced audio-heart rate signal group using a general emotion model to obtain an enhanced audio-heart rate feature group. The phase synchronization coefficient of the original audio-heart rate signal group is calculated. Based on the phase synchronization coefficient, the reconstruction attention matrix of the enhanced audio-heart rate feature group is constructed using a matrix factorization mechanism. Cross-modal fusion is then performed based on the reconstruction attention matrix to obtain cross-modal emotional features. The general sentiment model is pre-trained and meta-learning adapted based on the pre-collected bimodal dataset to obtain a mental health model. The cross-modal sentiment features are then updated using the mental health model, and membership degree calculation and sentiment analysis are performed on the updated cross-modal sentiment features to obtain the detection results.

[0006] According to a preferred embodiment of the present invention, calculating the emotion relevance score of the noise signal group in the original audio-heart rate signal group includes: Audio noise candidate signals and heart rate interference candidate signals are extracted from the original audio-heart rate signal group respectively, and a noise signal group is formed. The original audio emotional features and the original heart rate emotional features are extracted from the original audio-heart rate signal set, and an emotional feature set is constructed. Based on a pre-acquired real elderly bimodal dataset, the emotional feature set and the noise signal set are jointly modeled to obtain a spatial distribution model of emotional features; The emotional relevance score between the emotional feature set and the noise signal set is calculated based on the spatial distribution model of the emotional features.

[0007] According to another preferred embodiment of the present invention, calculating the emotional correlation score between the emotional feature set and the noise signal group based on the emotional feature spatial distribution model includes: Calculate the emotional information entropy of the emotional feature set, the noise information entropy of the noise signal group, and the joint information entropy of the emotional feature set and the noise signal group; Emotion-noise mutual information is calculated based on the emotion information entropy, noise information entropy, and joint information entropy. The emotion relevance score is calculated based on the emotion-noise mutual information, emotion information entropy, and noise information entropy.

[0008] According to another preferred embodiment of the present invention, cross-modal cross-validation and weighted noise suppression are performed on the original audio-heart rate signal set based on the emotion relevance score to obtain a noise-reduced audio-heart rate signal set, including: Based on the emotion relevance score, noise signal groups in the original audio-heart rate signal group are labeled with noise to obtain primary noise type labels; Based on the stability of the original audio-heart rate signal group, the primary noise type label is bidirectionally verified to obtain the standard noise type label; Based on the standard noise type labels, audio noise suppression weights and heart rate interference suppression weights are constructed respectively, and the audio noise suppression weights and heart rate interference suppression weights are optimized based on the optimization objective of minimizing conditional entropy. The original audio-heart rate signal group is subjected to weighted noise suppression using optimized audio noise suppression weights and heart rate interference suppression weights to obtain a noise-reduced audio-heart rate signal group.

[0009] According to another preferred embodiment of the present invention, a general emotion model is used to perform parallel feature extraction and mutual information feature enhancement on the noise-reduced audio-heart rate signal group to obtain an enhanced audio-heart rate feature group, including: A general sentiment model is used to perform depthwise separable temporal convolution on the noise-reduced frequency signal in the noise-heart rate signal group to obtain noise-reduced frequency sentiment features. Long-term and short-term time-series feature modeling is performed on the denoised heart rate signal in the denoised frequency-heart rate signal group to obtain the denoised heart rate emotional features. The proportion of audio mutual information of the noise-reduced frequency emotion features and the proportion of heart rate mutual information of the noise-reduced heart rate emotion features are calculated respectively. Based on the audio mutual information ratio, the noise-reduced audio emotional features are enhanced to obtain enhanced audio emotional features. Based on the heart rate mutual information ratio, the noise-reduced heart rate emotional features are enhanced to obtain enhanced heart rate emotional features. The enhanced audio emotional features and the enhanced heart rate emotional features are then combined into an enhanced audio-heart rate feature group.

[0010] According to another preferred embodiment of the present invention, the phase synchronization coefficient of the original audio-heart rate signal group is calculated, including: The audio timestamp sequence of the original audio signal and the heart rate timestamp sequence of the original heart rate signal in the original audio-heart rate signal group are extracted respectively. The phase synchronization coefficients corresponding to the audio timestamp sequence and the heart rate timestamp sequence are calculated based on a preset time delay constant.

[0011] According to another preferred embodiment of the present invention, in conjunction with the phase synchronization coefficient, a reconstructed attention matrix for the enhanced audio-heart rate feature group is constructed based on a matrix factorization mechanism, including: Combining the phase synchronization coefficients, an initial attention matrix for the enhanced audio-heart rate feature group is constructed using a preset learnable weight matrix; The initial attention matrix is ​​decomposed into a nonnegative matrix to obtain the basis matrix and the coefficient matrix; Based on the basis matrix and coefficient matrix, matrix reconstruction is performed to obtain the reconstructed attention matrix.

[0012] According to another preferred embodiment of the present invention, a mental health model is obtained by pre-training the general emotion model and performing meta-learning adaptation based on a pre-collected bimodal dataset, including: The pre-collected bimodal dataset is split into a general dataset, a support set, and a query set, and the general sentiment model is pre-trained based on the general dataset. During pre-training, the model parameters of the general sentiment model are fine-tuned in an inner loop based on the support set to obtain temporary parameters; The meta-loss of the temporary parameters is calculated based on the query set, and the model parameters of the general model are updated according to the meta-loss to obtain the mental health model.

[0013] According to another preferred embodiment of the present invention, the cross-modal emotional features are updated using the mental health model, and membership degree calculation and sentiment analysis are performed on the updated cross-modal emotional features to obtain detection results, including: Based on the support set, individual emotion calibration is performed to obtain an individual-specific emotion feature center set; Calculate the membership degree of the updated cross-modal sentiment features relative to each individual-specific sentiment feature center in the set of individual-specific sentiment feature centers to obtain the membership degree set; The detection results are generated based on the individual-specific emotional feature center set and the corresponding membership degree set.

[0014] To achieve at least one of the above-mentioned objectives, the present invention further provides a mental health detection system based on heart rate audio fusion, the system comprising a time-series alignment module, a noise suppression module, a feature extraction module, a cross-modal fusion module, and an emotion analysis module, wherein: The timing alignment module performs timing alignment on the synchronously acquired raw audio signal and raw heart rate signal to obtain the raw audio-heart rate signal group, and calculates the emotion relevance score of the noise signal group in the raw audio-heart rate signal group. The noise suppression module performs cross-modal cross-validation and weighted noise suppression on the original audio-heart rate signal group based on the emotion relevance score to obtain a noise-reduced audio-heart rate signal group. The feature extraction module uses a general emotion model to perform parallel feature extraction and mutual information feature enhancement on the noise-reduced audio-heart rate signal group to obtain an enhanced audio-heart rate feature group. The cross-modal fusion module calculates the phase synchronization coefficient of the original audio-heart rate signal group, and constructs the reconstruction attention matrix of the enhanced audio-heart rate feature group based on the matrix factorization mechanism in combination with the phase synchronization coefficient. Then, cross-modal fusion is performed based on the reconstruction attention matrix to obtain cross-modal emotional features. The sentiment analysis module performs model pre-training and meta-learning adaptation on the general sentiment model based on the pre-collected bimodal dataset to obtain a mental health model. The mental health model is then used to update the cross-modal sentiment features, and the membership degree of the updated cross-modal sentiment features is calculated and sentiment analysis is performed to obtain the detection results.

[0015] The present invention further provides a computer-readable storage medium storing a computer program, which is executed by a processor to implement the above-described method for mental health detection based on heart rate audio fusion.

[0016] (III) Beneficial Effects Compared with existing technologies, the present invention provides a mental health detection method and system based on heart rate audio fusion, which has the following beneficial effects: This heart rate-audio fusion-based mental health detection method utilizes the temporal alignment results of the original audio signal and the original heart rate signal. Without relying on single-modal judgment, it introduces information theory methods to quantitatively evaluate the statistical correlation between noise signals and emotional features, thereby achieving accurate differentiation between emotion-related noise and emotion-independent noise. Through the construction of an emotional feature spatial distribution model based on a real elderly bimodal dataset, the emotional relevance score has a stable statistical prior, avoiding misjudgments caused by instantaneous noise or occasional heart rate fluctuations. It provides a reliable quantitative basis for subsequent cross-modal mutual verification and noise suppression weight modeling, effectively ensuring that emotional information is not weakened during the denoising process, and improving the accuracy and robustness of subsequent emotional feature purification, cross-modal fusion, and mental health analysis from the source.

[0017] This heart rate-audio fusion-based mental health detection method avoids misclassifying breathing sounds and sudden heart rate changes caused by real emotions as noise by using emotion relevance scoring and cross-modal verification. Simultaneously, information theory-driven weight optimization suppresses non-emotional noise while preserving emotional perturbations. This provides highly reliable, low-biased basic signal input for subsequent emotion feature purification, temporal modeling, and personalized emotion boundary learning, significantly improving the overall robustness and interpretability of emotion recognition. Through parallel heterogeneous modeling and mutual information-driven enhancement, it captures local details of emotional sound quality using deep separable temporal convolution, targeting the high-frequency instantaneous characteristics of audio signals. For the long-term physiological evolution trend of heart rate signals, it captures temporal dependencies using long- and short-term temporal feature modeling, thereby extracting more abstract and discriminative deep emotional features. Supervised feature enhancement is achieved by weighting the feature dimensions most important for the final emotion classification task, ensuring greater weight for these dimensions. This improves feature quality and representativeness before subsequent fusion, effectively highlighting the most emotion-related information and significantly enhancing the stability and generalization ability of mental health detection.

[0018] This mental health detection method based on heart rate-audio fusion, through a phase synchronization modeling mechanism based on the time delay patterns of elderly people's emotional physiological responses, introduces physiological prior constraints on the asynchronous nature of audio-heart rate-emotional triggering, building upon traditional bimodal fusion methods that rely solely on feature splicing or simple temporal alignment. This enables time delay-perceived weighting of cross-modal feature associations. By mapping the temporal offset between audio timestamps and heart rate timestamps to phase synchronization coefficients and introducing them as explicit priors into the attention construction process, the allocation of attention weights aligns with the true physiological response patterns of elderly people's emotional stimuli. This suppresses non-emotional associations under temporal mismatch conditions and strengthens cross-modal feature interactions with genuine emotional causal relationships. Furthermore, by performing low-rank reconstruction of the high-dimensional initial attention matrix through non-negative matrix factorization, the main emotional collaboration patterns are preserved while reducing computational complexity. This not only improves the stability and discriminativeness of cross-modal emotional feature fusion but also provides a structurally clear and physically interpretable feature foundation for subsequent personalized mental health models in emotional analysis under conditions of low-sampling-rate heart rate and high-temporal-resolution audio. Attached Figure Description

[0019] Figure 1 The diagram shows a flowchart of a mental health detection method based on heart rate audio fusion according to the present invention. Figure 2 The diagram shown is the network structure diagram corresponding to the extraction of noise-reduced audio emotional features in this invention. Figure 3 The diagram shown is the network structure diagram corresponding to the noise-reduced heart rate and emotional features extracted in this invention. Detailed Implementation

[0020] The following description is intended to disclose the present invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious modifications will occur to those skilled in the art. The basic principles of the invention defined in the following description can be applied to other embodiments, modifications, improvements, equivalents, and other technical solutions that do not depart from the spirit and scope of the invention.

[0021] It is understood that the term "a" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element can be one, while in another embodiment, the number of the element can be multiple, and the term "a" should not be understood as a limitation on the number.

[0022] Example 1: Please combine Figure 1 This invention discloses a mental health detection method based on heart rate audio fusion, the method comprising the following steps: The original audio signal and original heart rate signal were time-aligned to obtain the original audio-heart rate signal group, and the emotion relevance score of the noise signal group in the original audio-heart rate signal group was calculated.

[0023] The original audio signal is an audio signal collected from elderly people living at home (sampling rate 16kHz, 16-bit quantization, mono), and the original heart rate signal is a heart rate signal collected from elderly people living at home (sampling rate 1Hz, effective range 50-130 beats / min). Both the original audio signal and the original heart rate signal carry a collection timestamp, and the corresponding collection accuracy is 10ms. Timing alignment means that the continuous audio signal is segmented into 5-second windows using the heart rate signal timestamp as the anchor point. Each 5-second audio segment corresponds to a heart rate data point within the same time interval, ensuring that each audio signal is time-synchronized with the heart rate signal of the corresponding time period. The original audio-heart rate signal group contains the aligned original audio signal and the original heart rate signal.

[0024] Specifically, the emotional relevance score of the noise signal group in the original audio-heart rate signal group is calculated, including: Audio noise candidate signals and heart rate interference candidate signals are extracted from the original audio-heart rate signal group respectively, and a noise signal group is formed. The original audio emotional features and the original heart rate emotional features are extracted from the original audio-heart rate signal set, and an emotional feature set is constructed. Based on a pre-acquired real elderly bimodal dataset, the emotional feature set and the noise signal set are jointly modeled to obtain a spatial distribution model of emotional features; The emotional relevance score between the emotional feature set and the noise signal set is calculated based on the spatial distribution model of the emotional features.

[0025] Wherein, when the original audio signal is The original heart rate signal is At that time, the corresponding audio noise candidate signal is Heart rate interference candidate signal is The noise signal group is Candidate audio noise and heart rate interference signals can be extracted using adaptive filters or time-domain analysis. For example, candidate audio noise signals include background noise from the television in a home environment, the sound of a door closing, and external ambient noise, as well as emotion-related audio interference, such as rapid breathing during anxiety. Candidate heart rate interference signals include abnormal heart rate fluctuations caused by accidental movement in the elderly, equipment acquisition errors, and emotion-related heart rate interference, such as a sudden increase in heart rate during emotional excitement. The extraction of original audio emotional features and original heart rate emotional features from the original audio-heart rate signal set includes extracting original audio emotional features such as fundamental frequency F0, spectral entropy, energy, and speech rate from the original audio signal, and original heart rate emotional features such as mean, fluctuation variance, rise / fall slope, and difference from baseline from the original heart rate signal, thus constructing an emotional feature set. .

[0026] In detail, the aforementioned real-world bimodal dataset of elderly people refers to audio signals and heart rate signals of elderly people living at home that have been genuinely collected and labeled with emotional and noise tags. Based on the statistical characteristics of the real-world bimodal dataset of elderly people, a spatial distribution model of emotional features can be obtained through kernel density estimation. , , ,in, This is the noise signal corresponding to a real bimodal dataset of elderly people. This is the emotional feature set corresponding to a real bimodal dataset of elderly people. Let be the probability density function. represent and The joint probability density is also used to characterize and The joint probability density, represent The probability density is also used to characterize The probability density, represent The probability density is also used to characterize The probability density.

[0027] Specifically, the emotional relevance score between the emotional feature set and the noise signal group is calculated based on the emotional feature spatial distribution model, including: Calculate the emotional information entropy of the emotional feature set, the noise information entropy of the noise signal group, and the joint information entropy of the emotional feature set and the noise signal group; Emotion-noise mutual information is calculated based on the emotion information entropy, noise information entropy, and joint information entropy. The emotion relevance score is calculated based on the emotion-noise mutual information, emotion information entropy, and noise information entropy.

[0028] The sentiment relevance score can be calculated using the following formula: in, For information entropy function, For noise information entropy, For emotional information entropy, For joint information entropy, For emotion-noise mutual information, Score the sentiment relevance.

[0029] By utilizing the temporal alignment results of the original audio signal and the original heart rate signal, and without relying on single-modal judgment, information theory methods are introduced to quantify the statistical correlation between noise signals and emotional features. This enables accurate differentiation between emotion-related noise and emotion-independent noise. Through the construction of an emotional feature spatial distribution model based on a real bimodal dataset of elderly people, the emotional relevance score has a stable statistical prior, avoiding misjudgments caused by instantaneous noise or occasional heart rate fluctuations. This provides a reliable quantitative basis for subsequent cross-modal mutual verification and noise suppression weight modeling, effectively ensuring that emotional information is not weakened during the denoising process. This fundamentally improves the accuracy and robustness of subsequent emotional feature purification, cross-modal fusion, and mental health analysis.

[0030] Based on the emotion relevance score, cross-modal cross-verification and weighted noise suppression are performed on the original audio-heart rate signal set to obtain a noise-reduced audio-heart rate signal set.

[0031] The noise-reduced frequency-heart rate signal group includes noise-reduced frequency signals and noise-reduced heart rate signals.

[0032] Specifically, based on the emotion relevance score, cross-modal cross-validation and weighted noise suppression are performed on the original audio-heart rate signal set to obtain a noise-reduced audio-heart rate signal set, including: Based on the emotion relevance score, noise signal groups in the original audio-heart rate signal group are labeled with noise to obtain primary noise type labels; Based on the stability of the original audio-heart rate signal group, the primary noise type label is bidirectionally verified to obtain the standard noise type label; Based on the standard noise type labels, audio noise suppression weights and heart rate interference suppression weights are constructed respectively, and the audio noise suppression weights and heart rate interference suppression weights are optimized based on the optimization objective of minimizing conditional entropy. The original audio-heart rate signal group is subjected to weighted noise suppression using optimized audio noise suppression weights and heart rate interference suppression weights to obtain a noise-reduced audio-heart rate signal group.

[0033] The noise labeling refers to judging the noise of each noise signal group according to a preset correlation threshold, such as 0.6. When the emotional correlation score of the corresponding segment is greater than the correlation threshold, it is marked as emotionally relevant noise and retained; otherwise, it is marked as emotionally irrelevant noise. The bidirectional verification refers to using the synchronicity of dual modes for logical verification. When a segment of a signal in a noise signal group is emotionally irrelevant noise, the stationarity of another signal in the same time period is verified. If it is stationary, it is determined to be environmental noise and is not modified. If it is not stationary, it is updated to be emotionally relevant noise. For example, if a certain noise is detected in an audio signal, its status as emotionally irrelevant noise is verified by the stability of the heart rate signal in the corresponding time period. That is, if the heart rate signal is stable (fluctuation ≤ 5 beats / min within 5 seconds), it is likely to be emotionally irrelevant noise; if the heart rate signal fluctuates synchronously, it is likely to be emotionally relevant noise.

[0034] In detail, the formulas for constructing the audio noise suppression weights and heart rate interference suppression weights are as follows: The formulas for optimizing the audio noise suppression weights and heart rate interference suppression weights based on the optimization objective of minimizing conditional entropy are as follows: The formula for weighted noise suppression is as follows: in, For audio noise suppression weights, For heart rate interference suppression weights, It is the sigmoid activation function. Given an original audio signal, the conditional entropy of the corresponding original heart rate signal is used to measure the audio's ability to verify heart rate noise. Given a raw heart rate signal, the conditional entropy of the corresponding raw audio signal is used to measure the heart rate's ability to validate audio noise. To remove noise from frequency signals, For noise-reduced heart rate signals, and This constitutes a noise-reduced frequency-heart rate signal group.

[0035] By using emotion relevance scoring and cross-modal verification, we avoid misjudging breathing sounds and sudden changes in heart rate caused by real emotions as noise. At the same time, we suppress non-emotional noise and preserve emotional perturbations through information theory-driven weight optimization. This provides a highly reliable and low-biased basic signal input for subsequent emotion feature purification, temporal modeling, and personalized emotion boundary learning, significantly improving the overall robustness and interpretability of emotion recognition.

[0036] Parallel feature extraction and mutual information feature enhancement are performed on the noise-reduced audio-heart rate signal group using a general emotion model to obtain an enhanced audio-heart rate feature group.

[0037] The enhanced audio-heart rate feature group includes enhanced audio emotional features and enhanced heart rate emotional features.

[0038] Specifically, a general emotion model is used to perform parallel feature extraction and mutual information feature enhancement on the noise-reduced audio-heart rate signal group to obtain an enhanced audio-heart rate feature group, including: A general sentiment model is used to perform depthwise separable temporal convolution on the noise-reduced frequency signal in the noise-heart rate signal group to obtain noise-reduced frequency sentiment features. Long-term and short-term time-series feature modeling is performed on the denoised heart rate signal in the denoised frequency-heart rate signal group to obtain the denoised heart rate emotional features. The proportion of audio mutual information of the noise-reduced frequency emotion features and the proportion of heart rate mutual information of the noise-reduced heart rate emotion features are calculated respectively. Based on the audio mutual information ratio, the noise-reduced audio emotional features are enhanced to obtain enhanced audio emotional features. Based on the heart rate mutual information ratio, the noise-reduced heart rate emotional features are enhanced to obtain enhanced heart rate emotional features. The enhanced audio emotional features and the enhanced heart rate emotional features are then combined into an enhanced audio-heart rate feature group.

[0039] The general sentiment model refers to a general sentiment classification neural network model with gradient update capability that outputs sentiment labels based on sentiment features. This general sentiment model includes CNN layers, LSTM layers, an attention fusion layer for cross-modal fusion, and a fully connected classification layer. (Please refer to...) Figure 2 This is the network structure diagram for extracting the emotional features of noise-reduced audio signals. It corresponds to the CNN layer of a general sentiment model. During parallel extraction, the noise-reduced audio signal is input through the input layer and then subjected to one-dimensional temporal convolution through a Conv 1d layer to extract local features. This is followed by non-linear activation through a ReLU layer and one-dimensional max pooling through a MaxPool 1d layer. This process of one-dimensional temporal convolution, non-linear activation, and one-dimensional max pooling is repeated. Finally, a third one-dimensional temporal convolution is performed, and the output layer outputs the noise-reduced audio signal's emotional features. Please refer to... Figure 3 This is the network structure diagram for extracting denoised heart rate emotional features, corresponding to the LSTM layer of a general emotional model. On the other hand, the denoised heart rate signal is input through the input layer, then fully connected through a Linear layer, non-linearly activated through a ReLU layer, and its long short-term memory temporal features are extracted through an LSTM layer and fully connected through another Linear layer. Finally, the output is passed through the output layer to obtain the denoised heart rate emotional features. The formula for feature enhancement is as follows: in, It enhances the emotional features of audio. This refers to noise reduction of frequency signals. Perform depthwise separable temporal convolution, i.e., the noise-removed frequency sentiment features. The dot product symbol. It is the Softmax normalization function. This refers to the noise-reduced heart rate signal Long-term and short-term time-series feature modeling is performed, namely the noise-reduced heart rate-emotional features. For the set of real numbers, To enhance the number of audio frames corresponding to audio emotional features, each frame is 20ms long with a step size of 10ms. To increase the number of heart rate data points corresponding to heart rate emotional characteristics, a 5-second window corresponds to 5 data points.

[0040] By employing parallel heterogeneous modeling and mutual information-driven enhancement, this study captures local details of emotional sound quality using deep separable temporal convolution, targeting the high-frequency instantaneous characteristics of audio signals. Furthermore, it captures temporal dependencies by modeling long- and short-term temporal features, thereby extracting more abstract and discriminative deep emotional features. Supervised feature enhancement is achieved by weighting the mutual information ratio with the emotional feature set, ensuring that the most important feature dimensions for the final emotion classification task receive greater weight. This improves the quality and representativeness of features before subsequent fusion, effectively highlighting the information most relevant to emotions and significantly enhancing the stability and generalization ability of mental health testing.

[0041] The phase synchronization coefficient of the original audio-heart rate signal group is calculated. Based on the phase synchronization coefficient, the reconstruction attention matrix of the enhanced audio-heart rate feature group is constructed using a matrix factorization mechanism. Cross-modal fusion is then performed based on the reconstruction attention matrix to obtain cross-modal emotional features.

[0042] Since the emotional and physiological responses of the elderly are not instantaneous, an emotional stimulus will be immediately reflected in the audio signal, but the change in heart rate needs to be regulated by nerves and body fluids and will appear after a lag of about 1.2 seconds. Therefore, it is necessary to calculate the phase synchronization coefficient to facilitate the cross-modal fusion of audio-heart rate feature groups.

[0043] Specifically, the phase synchronization coefficient of the original audio-heart rate signal group is calculated, including: The audio timestamp sequence of the original audio signal and the heart rate timestamp sequence of the original heart rate signal in the original audio-heart rate signal group are extracted respectively. The phase synchronization coefficients corresponding to the audio timestamp sequence and the heart rate timestamp sequence are calculated based on a preset time delay constant.

[0044] The audio timestamp sequence refers to the sequence of timestamps corresponding to each audio frame in the original audio signal, while the heart rate timestamp sequence is the sequence of timestamps corresponding to each heart rate data point in the original heart rate signal. The formula for calculating the phase synchronization coefficient is as follows: in, It refers to the first time stamp in the audio timestamp sequence. The first audio timestamp and heart rate timestamp sequence Phase synchronization coefficient between heart rate timestamps For the audio timestamp sequence An audio timestamp The first in the heart rate timestamp sequence Heart rate timestamp , representing the time delay constant of heart rate lag in audio emotion triggering based on statistics from an elderly emotion response dataset. , refers to the smoothing coefficient; The closer the value is to 1, the more synchronized the emotional trigger points of the enhanced audio-heart rate feature group are.

[0045] In detail, based on the phase synchronization coefficients, a reconstruction attention matrix for the enhanced audio-heart rate feature group is constructed using a matrix factorization mechanism, including: Combining the phase synchronization coefficients, an initial attention matrix for the enhanced audio-heart rate feature group is constructed using a preset learnable weight matrix; The initial attention matrix is ​​decomposed into a nonnegative matrix to obtain the basis matrix and the coefficient matrix; Based on the basis matrix and coefficient matrix, matrix reconstruction is performed to obtain the reconstructed attention matrix.

[0046] The formula for constructing the initial attention matrix is ​​as follows: in, And it is a learnable weight matrix. To enhance audio features The learnable weight matrix, To enhance heart rate characteristics The learnable weight matrix, Phase synchronization coefficient The learnable weight matrix, The first one in the initial attention matrix Line number The column corresponding to the weight element; Non-negative matrix factorization (NMF) refers to decomposing a matrix into the product of two smaller matrices with non-negative elements, i.e., decomposing it into a base matrix. sum coefficient matrix ,in , is the decomposition dimension used to balance accuracy and complexity, and seeks a globally optimal approximate solution that satisfies This yields the basis matrix and coefficient matrix. Matrix reconstruction refers to obtaining the reconstructed attention matrix by multiplying the decomposed basis matrix and coefficient matrix. , To reconstruct the attention matrix, decomposition and reconstruction can reduce the computational complexity of the initial attention matrix, thus reducing the computational load. ,when At that time, the computational workload was reduced by 68%.

[0047] Specifically, cross-modal fusion is performed based on the reconstructed attention matrix to obtain cross-modal sentiment features, including: Obtain the number of audio frames corresponding to the enhanced audio emotional features and the number of heart rate data points corresponding to the enhanced heart rate emotional features; The enhanced audio emotional features and enhanced heart rate emotional features are weighted and fused based on the reconstructed attention matrix, and the weighted fusion result is normalized based on the number of audio frames and the number of heart rate data points to obtain cross-modal emotional features. This step corresponds to the attention fusion layer of the general emotional model.

[0048] The mathematical formula for cross-modal fusion is as follows: in, For cross-modal sentiment features, To reconstruct the attention matrix of the th Line 1 The weight element corresponding to the column.

[0049] By employing a phase synchronization modeling mechanism based on the time delay patterns of elderly emotional physiological responses, this approach builds upon traditional bimodal fusion methods that rely solely on feature splicing or simple temporal alignment. It introduces a physiological prior constraint on the asynchronicity of audio-heart rate-emotional triggering, enabling time delay-perceived weighting of cross-modal feature associations. By mapping the temporal offset between audio and heart rate timestamps to phase synchronization coefficients and introducing them as explicit priors into the attention construction process, the allocation of attention weights aligns with the actual physiological response patterns of elderly individuals to emotional stimuli. This suppresses non-emotional associations under temporal mismatch conditions and strengthens cross-modal feature interactions with genuine emotional causal relationships. Furthermore, by performing low-rank reconstruction of the high-dimensional initial attention matrix through non-negative matrix factorization, the computational complexity is reduced while preserving the main emotional collaboration patterns. This not only improves the stability and discriminative power of cross-modal emotional feature fusion but also provides a structurally clear and physically interpretable feature foundation for subsequent personalized mental health models in emotional analysis under conditions of low-sampling-rate heart rate and high-temporal-resolution audio.

[0050] The general sentiment model is pre-trained and meta-learning adapted based on the pre-collected bimodal dataset to obtain a mental health model. The cross-modal sentiment features are then updated using the mental health model, and membership degree calculation and sentiment analysis are performed on the updated cross-modal sentiment features to obtain the detection results.

[0051] The bimodal dataset includes a general dataset, a support set, and a query set. The general dataset consists of bimodal data collected in advance from 35 target elderly people (60-85 years old) living at home, labeled with four emotion tags including calm, depressed, anxious, and happy, with a total sample size of more than 5,000, used for model pre-training. The support set refers to a small amount of individual data (≤50 data points, including the four emotion categories) collected for each target elderly person for meta-learning. The query set refers to 10 additional data points collected for individual adaptation fine-tuning.

[0052] Specifically, based on a pre-collected bimodal dataset, the general emotion model is pre-trained and meta-learning adapted to obtain a mental health model, including: The pre-collected bimodal dataset is split into a general dataset, a support set, and a query set, and the general sentiment model is pre-trained based on the general dataset. During pre-training, the model parameters of the general sentiment model are fine-tuned in an inner loop based on the support set to obtain temporary parameters; The meta-loss of the temporary parameters is calculated based on the query set, and the model parameters of the general model are updated according to the meta-loss to obtain the mental health model.

[0053] Specifically, pre-training the general sentiment model based on the general dataset means using the general dataset as input and the corresponding sentiment labels as labels for training. The initial model parameters for pre-training are... With a learning rate of 0.001 and 100 training epochs, the formula for calculating temporary parameters is as follows: The formula for calculating the element loss is as follows: The formula for updating model parameters using meta-loss is as follows: in, These are temporary parameters. This is a fine-tuning of the learning rate, set to 0.01. This refers to the gradient corresponding to each update of the model parameters. Let cross-entropy be the classification loss function. For the model parameters in each update round, For the meta-loss function, This represents the number of target elderly people corresponding to the general dataset, which is 35. It refers to the first A target elderly person, For the first A support group for elderly people with specific goals. For the first A query set targeting elderly individuals. To complete the final updated model parameters, the model parameters include all convolutional kernel parameters and bias parameters of the CNN, the gating weights and bias parameters of the LSTM, all learnable weight parameters of the initial attention matrix, and the final fully connected weights; updating the cross-modal emotional features using the mental health model means recalculating the cross-modal emotional features corresponding to the target elderly person for whom emotional analysis needs to be performed using the mental health model, thereby completing the corresponding update.

[0054] In detail, the cross-modal emotional features are updated using the aforementioned mental health model, and membership degree calculation and sentiment analysis are performed on the updated cross-modal emotional features to obtain detection results, including: Based on the support set, individual emotion calibration is performed to obtain an individual-specific emotion feature center set; Calculate the membership degree of the updated cross-modal sentiment features relative to each individual-specific sentiment feature center in the set of individual-specific sentiment feature centers to obtain the membership degree set; The detection results are generated based on the individual-specific emotional feature center set and the corresponding membership degree set.

[0055] The support set refers to the support set data of the target elderly person for whom sentiment analysis is required. Individual sentiment calibration is performed once a week. That is, for the data in the support set labeled with various sentiment types, cross-modal sentiment feature clusters corresponding to each sentiment type are extracted using cross-modal sentiment features, and the cluster center of each cross-modal sentiment feature cluster is used as the corresponding individual-specific sentiment feature center. , These correspond to four different emotional categories: calm, depressed, anxious, and joyful. The formula for calculating membership degree is as follows: in, Cross-modal sentiment features Compared to the first Membership degree of each individual's unique emotional characteristic center For fuzzy bandwidth, empirical values ​​based on individual data statistics are obtained. The symbol is Euclidean distance. Generating detection results based on the individual-specific emotional feature center set and the corresponding membership degree set means taking the emotional type corresponding to the individual-specific emotional feature center with the maximum membership degree as the emotional recognition result, and simultaneously outputting the corresponding membership degree. When displayed graphically, the system matches a preset warning color for visualization output based on the emotional type and membership degree level of the recognition result. For example, the membership degree of anxiety is 80, and the corresponding color is red. When the membership degree of anxiety is 40, the corresponding color is orange. When the corresponding voice recognition module detects high-risk words such as "help," "danger," and "fire," the system ignores the membership degree of the emotion and directly triggers the highest level red alert.

[0056] By introducing a meta-learning-based emotion model adaptation mechanism, the system can quickly model and accurately adapt to the differences in emotional expression among elderly people living at home under limited individual sample conditions, avoiding the problem of insufficient generalization ability of traditional unified models in individual emotion recognition. At the same time, in the model inference stage, an individual-specific emotion feature center and fuzzy membership degree calculation mechanism based on support set construction are introduced to transform the originally discrete emotion classification results into continuous and interpretable membership degree expressions. This not only improves the stability and robustness of emotion recognition in boundary states, but also supports the coexistence of multiple emotions and gradual discrimination. Combined with the rule-based fallback trigger strategy for high-risk semantics, the system can balance intelligent judgment and safety reliability in complex home environments, thereby significantly improving the practical value and engineering deployability of the overall mental health monitoring solution.

[0057] Example 2: This invention discloses a mental health detection system based on heart rate audio fusion. The system includes a time-series alignment module, a noise suppression module, a feature extraction module, a cross-modal fusion module, and a sentiment analysis module, wherein: The timing alignment module performs timing alignment on the synchronously acquired raw audio signal and raw heart rate signal to obtain the raw audio-heart rate signal group, and calculates the emotion relevance score of the noise signal group in the raw audio-heart rate signal group. The noise suppression module performs cross-modal cross-validation and weighted noise suppression on the original audio-heart rate signal group based on the emotion relevance score to obtain a noise-reduced audio-heart rate signal group. The feature extraction module uses a general emotion model to perform parallel feature extraction and mutual information feature enhancement on the noise-reduced audio-heart rate signal group to obtain an enhanced audio-heart rate feature group. The cross-modal fusion module calculates the phase synchronization coefficient of the original audio-heart rate signal group, and constructs the reconstruction attention matrix of the enhanced audio-heart rate feature group based on the matrix factorization mechanism in combination with the phase synchronization coefficient. Then, cross-modal fusion is performed based on the reconstruction attention matrix to obtain cross-modal emotional features. The sentiment analysis module performs model pre-training and meta-learning adaptation on the general sentiment model based on the pre-collected bimodal dataset to obtain a mental health model. The mental health model is then used to update the cross-modal sentiment features, and the membership degree of the updated cross-modal sentiment features is calculated and sentiment analysis is performed to obtain the detection results.

[0058] The processes described above with reference to the flowcharts in the embodiments disclosed in this invention can be implemented as computer software programs. The embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs the functions defined in the methods of this application. It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wire segments, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless segments, wire segments, optical fibers, RF, etc., or any suitable combination thereof.

[0059] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0060] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The purpose of the present invention has been fully and effectively achieved. The functions and structural principles of the present invention have been shown and explained in the embodiments. Without departing from the stated principles, the implementation of the present invention may have any variations or modifications.

Claims

1. A method for detecting mental health based on heart rate-audio fusion, characterized in that, The method includes: The original audio signal and the original heart rate signal were time-aligned to obtain the original audio-heart rate signal group, and the emotion relevance score of the noise signal group in the original audio-heart rate signal group was calculated. Based on the emotion relevance score, cross-modal cross-verification and weighted noise suppression are performed on the original audio-heart rate signal set to obtain a noise-reduced audio-heart rate signal set. Parallel feature extraction and mutual information feature enhancement were performed on the noise-reduced audio-heart rate signal group using a general emotion model to obtain an enhanced audio-heart rate feature group. The phase synchronization coefficient of the original audio-heart rate signal group is calculated. Based on the phase synchronization coefficient, the reconstruction attention matrix of the enhanced audio-heart rate feature group is constructed using a matrix factorization mechanism. Cross-modal fusion is then performed based on the reconstruction attention matrix to obtain cross-modal emotional features. The general sentiment model is pre-trained and meta-learning adapted based on the pre-collected bimodal dataset to obtain a mental health model. The cross-modal sentiment features are then updated using the mental health model, and membership degree calculation and sentiment analysis are performed on the updated cross-modal sentiment features to obtain the detection results.

2. The mental health detection method based on heart rate audio fusion according to claim 1, characterized in that, The emotional relevance score of the noise signal group in the original audio-heart rate signal group was calculated, including: Audio noise candidate signals and heart rate interference candidate signals are extracted from the original audio-heart rate signal group respectively, and a noise signal group is formed. The original audio emotional features and the original heart rate emotional features are extracted from the original audio-heart rate signal set, and an emotional feature set is constructed. Based on a pre-acquired real elderly bimodal dataset, the emotional feature set and the noise signal set are jointly modeled to obtain a spatial distribution model of emotional features; The emotional relevance score between the emotional feature set and the noise signal set is calculated based on the spatial distribution model of the emotional features.

3. The mental health detection method based on heart rate audio fusion according to claim 2, characterized in that, The emotional relevance score between the emotional feature set and the noise signal group is calculated based on the aforementioned emotional feature spatial distribution model, including: Calculate the emotional information entropy of the emotional feature set, the noise information entropy of the noise signal group, and the joint information entropy of the emotional feature set and the noise signal group; Emotion-noise mutual information is calculated based on the emotion information entropy, noise information entropy, and joint information entropy. The emotion relevance score is calculated based on the emotion-noise mutual information, emotion information entropy, and noise information entropy.

4. The mental health detection method based on heart rate audio fusion according to claim 3, characterized in that, Based on the emotion relevance score, cross-modal cross-validation and weighted noise suppression are performed on the original audio-heart rate signal set to obtain a noise-reduced audio-heart rate signal set, including: Based on the emotion relevance score, noise signal groups in the original audio-heart rate signal group are labeled with noise to obtain primary noise type labels; Based on the stability of the original audio-heart rate signal group, the primary noise type label is bidirectionally verified to obtain the standard noise type label; Based on the standard noise type labels, audio noise suppression weights and heart rate interference suppression weights are constructed respectively, and the audio noise suppression weights and heart rate interference suppression weights are optimized based on the optimization objective of minimizing conditional entropy. The original audio-heart rate signal group is subjected to weighted noise suppression using optimized audio noise suppression weights and heart rate interference suppression weights to obtain a noise-reduced audio-heart rate signal group.

5. A method for mental health detection based on heart rate audio fusion according to claim 1, characterized in that, Parallel feature extraction and mutual information feature enhancement are performed on the noise-reduced audio-heart rate signal group using a general emotion model to obtain an enhanced audio-heart rate feature group, including: A general sentiment model is used to perform depthwise separable temporal convolution on the noise-reduced frequency signal in the noise-heart rate signal group to obtain noise-reduced frequency sentiment features. Long-term and short-term time-series feature modeling is performed on the denoised heart rate signal in the denoised frequency-heart rate signal group to obtain the denoised heart rate emotional features. The proportion of audio mutual information of the noise-reduced frequency emotion features and the proportion of heart rate mutual information of the noise-reduced heart rate emotion features are calculated respectively. Based on the audio mutual information ratio, the noise-reduced audio emotional features are enhanced to obtain enhanced audio emotional features. Based on the heart rate mutual information ratio, the noise-reduced heart rate emotional features are enhanced to obtain enhanced heart rate emotional features. The enhanced audio emotional features and the enhanced heart rate emotional features are then combined into an enhanced audio-heart rate feature group.

6. The mental health detection method based on heart rate audio fusion according to claim 1, characterized in that, The phase synchronization coefficient of the original audio-heart rate signal group is calculated, including: The audio timestamp sequence of the original audio signal and the heart rate timestamp sequence of the original heart rate signal in the original audio-heart rate signal group are extracted respectively. The phase synchronization coefficients corresponding to the audio timestamp sequence and the heart rate timestamp sequence are calculated based on a preset time delay constant.

7. A method for mental health detection based on heart rate audio fusion according to claim 6, characterized in that, Based on the phase synchronization coefficients, a reconstruction attention matrix for the enhanced audio-heart rate feature group is constructed using a matrix factorization mechanism, including: Combining the phase synchronization coefficients, an initial attention matrix for the enhanced audio-heart rate feature group is constructed using a preset learnable weight matrix; The initial attention matrix is ​​decomposed into a nonnegative matrix to obtain the basis matrix and the coefficient matrix; Based on the basis matrix and coefficient matrix, matrix reconstruction is performed to obtain the reconstructed attention matrix.

8. A method for mental health detection based on heart rate audio fusion according to claim 1, characterized in that, Based on a pre-collected bimodal dataset, the general emotion model is pre-trained and meta-learning adapted to obtain a mental health model, including: The pre-collected bimodal dataset is split into a general dataset, a support set, and a query set, and the general sentiment model is pre-trained based on the general dataset. During pre-training, the model parameters of the general sentiment model are fine-tuned in an inner loop based on the support set to obtain temporary parameters; The meta-loss of the temporary parameters is calculated based on the query set, and the model parameters of the general model are updated according to the meta-loss to obtain the mental health model.

9. A method for mental health detection based on heart rate audio fusion according to claim 8, characterized in that, Membership degrees and sentiment analysis were performed on the updated cross-modal sentiment features to obtain detection results, including: Based on the support set, individual emotion calibration is performed to obtain an individual-specific emotion feature center set; Calculate the membership degree of the updated cross-modal sentiment features relative to each individual-specific sentiment feature center in the set of individual-specific sentiment feature centers to obtain the membership degree set; The detection results are generated based on the individual-specific emotional feature center set and the corresponding membership degree set.

10. A mental health detection system based on heart rate audio fusion, characterized in that, The system includes a temporal alignment module, a noise suppression module, a feature extraction module, a cross-modal fusion module, and a sentiment analysis module, wherein: The timing alignment module performs timing alignment on the synchronously acquired raw audio signal and raw heart rate signal to obtain the raw audio-heart rate signal group, and calculates the emotion relevance score of the noise signal group in the raw audio-heart rate signal group. The noise suppression module performs cross-modal cross-validation and weighted noise suppression on the original audio-heart rate signal group based on the emotion relevance score to obtain a noise-reduced audio-heart rate signal group. The feature extraction module uses a general emotion model to perform parallel feature extraction and mutual information feature enhancement on the noise-reduced audio-heart rate signal group to obtain an enhanced audio-heart rate feature group. The cross-modal fusion module calculates the phase synchronization coefficient of the original audio-heart rate signal group, and constructs the reconstruction attention matrix of the enhanced audio-heart rate feature group based on the matrix factorization mechanism in combination with the phase synchronization coefficient. Then, cross-modal fusion is performed based on the reconstruction attention matrix to obtain cross-modal emotional features. The sentiment analysis module performs model pre-training and meta-learning adaptation on the general sentiment model based on the pre-collected bimodal dataset to obtain a mental health model. The mental health model is then used to update the cross-modal sentiment features, and the membership degree of the updated cross-modal sentiment features is calculated and sentiment analysis is performed to obtain the detection results.