Mental health early warning method and system for mental health education

By collecting and preprocessing multimodal data in the mental health warning system, extracting and fusing feature vectors, calculating trajectory complexity and performing weighted sum-scoring scores, the problem of low accuracy in the identification of high-risk emotional states in the prior art is solved, and a more efficient and accurate mental health warning is achieved.

CN120148769AActive Publication Date: 2025-06-13TIBET XINYAN SCIENCE & TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510220736.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art relies on a single model or a simple feature averaging method in the process of feature extraction and fusion, and lacks effective control of data redundancy and noise, resulting in bias in emotional recognition and mental health scores, and is unable to effectively capture the fluctuations of individual psychological states, resulting in low accuracy in identifying high-risk emotional states.

Method used

By collecting multimodal data and preprocessing, extracting and fusing feature vectors, using delayed embedding to construct phase space trajectory points to calculate trajectory complexity, using weighted sum to calculate mental health scores and early warnings, and building a visual interface to display mental health scores in real time.

Benefits of technology

It improves sensitivity to emotional state fluctuations, improves the accuracy and timeliness of mental health warnings, and avoids the lag of users and monitors when monitoring psychological states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148769A_ABST
    Figure CN120148769A_ABST
Patent Text Reader

Abstract

The invention discloses a psychological health early warning method and system for psychological health education, and relates to the technical field of psychological health monitoring, and the method comprises the steps: collecting and preprocessing multi-modal data, and extracting feature vectors of the preprocessed multi-modal data for fusion; phase space trajectory points are constructed by using delay embedding to calculate trajectory complexity, and mental health scores and early warning are calculated by using weighted summation; and constructing a visual interface to display the psychological health score in real time, and storing, collecting and analyzing the generated multi-modal data. According to the method, the multi-modal data is collected and preprocessed, and the feature vectors of the preprocessed multi-modal data are extracted and fused; phase space trajectory points are constructed by using delay embedding to calculate trajectory complexity, and mental health scores and early warning are calculated by using weighted summation; sensitivity to emotional state fluctuation is improved, accuracy and timeliness of mental health early warning are improved, and hysteresis of a user and a monitor during mental state monitoring is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mental health monitoring, and in particular to a mental health early warning method and system for mental health education. Background Art

[0002] With the continuous improvement of the society's attention to mental health problems, the research on mental health education and early warning methods has gradually become a hot topic. Traditional mental health monitoring mainly relies on methods such as questionnaires and self-reports. Although these methods can reflect the mental health status to a certain extent, they are highly subjective and cannot dynamically capture the emotional changes of individuals in different situations, thus affecting the accuracy and timeliness of the evaluation results. With the rapid development of artificial intelligence and multi-modal data processing technologies, mental health monitoring systems based on emotion recognition have gradually come into the research field of vision. The multi-modal data fusion technology can more comprehensively depict the emotional state of individuals by combining physiological behavior characteristics such as voice and facial expressions. At present, there are still some deficiencies in common multi-modal emotion recognition technologies, such as the influence of noise interference on recognition accuracy, poor feature fusion effect, and lack of effective modeling of data time series characteristics, which limit their application effects in mental health early warning.

[0003] There are still some obvious deficiencies in the application of multi-modal data in mental health early warning. In the process of feature extraction and fusion, the existing technologies usually rely on a single model or simple feature averaging methods, lacking effective control of data redundancy and noise, resulting in deviations in emotion recognition and mental health scoring of the system. The existing mental health early warning systems mostly stay in the static analysis stage during data analysis, failing to fully utilize the dynamic emotion fluctuation characteristics contained in time series data and unable to effectively capture the fluctuation rules of individual mental states, resulting in low recognition accuracy for high-risk emotional states. Summary of the Invention

[0004] In view of the problems existing in the above-mentioned existing mental health early warning methods and systems for mental health education, the present invention is proposed.

[0005] Therefore, the problems to be solved by the present invention are that in the process of feature extraction and fusion, the existing technologies usually rely on a single model or simple feature averaging methods, lacking effective control of data redundancy and noise, resulting in deviations in emotion recognition and mental health scoring of the system. The existing mental health early warning systems mostly stay in the static analysis stage during data analysis, failing to fully utilize the dynamic emotion fluctuation characteristics contained in time series data and unable to effectively capture the fluctuation rules of individual mental states, resulting in low recognition accuracy for high-risk emotional states.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: A mental health early warning method for mental health education, which includes collecting multi-modal data and performing preprocessing, extracting feature vectors of the preprocessed multi-modal data for fusion; using delayed embedding to construct phase space trajectory points to calculate trajectory complexity, using weighted summation to calculate mental health scores and early warnings; constructing a visualization interface to display mental health scores in real time, and storing the multi-modal data generated by collection and analysis.

[0007] As a preferred solution of the mental health early warning method for mental health education according to the present invention, wherein: the collection of multi-modal data and preprocessing refers to collecting and preprocessing the user's multi-modal data using a high-definition camera and a microphone during the mental health education process;

[0008] The multi-modal data includes voice and facial expression data;

[0009] The preprocessing includes adding timestamps to the collected multi-modal data, using linear interpolation to correct the time synchronization of the multi-modal data, using a band-pass filter to remove the noise of the voice data, performing normalization processing on the denoised voice data, using Gaussian filtering to remove the noise of the facial expression data, using the Laplace operator to sharpen the facial expression data, and using the affine transformation method to perform normalization alignment on the facial expression data.

[0010] As a preferred solution of the mental health early warning method for mental health education according to the present invention, wherein: the extraction of feature vectors of the preprocessed multi-modal data for fusion refers to collecting historical multi-modal data and performing preprocessing and feature vector extraction to generate a training set;

[0011] Using a pre-trained FER+ model to classify the emotions of historical voice data and assign corresponding emotion labels, and using a pre-trained OpenSMILE model to classify the emotions of historical facial expression data and assign corresponding emotion labels;

[0012] Using a pre-trained Whisper model of OpenAI to extract the voice feature vectors of historical voice data, and using a pre-trained VGG-Face model to extract the facial feature vectors of historical facial expression data;

[0013] Introduce random noise ∈ from the standard normal distribution, use a fully connected network to convert the historical voice and facial feature vectors into means and variances respectively, use weighted average fusion to fuse the means and variances respectively, obtain the means and variances of historical fusion, and define a variational distribution q(Z│F′)=U(μ,σ based on the fused means and variances 2), where U is the symbol of the normal distribution, Z is the historical fusion feature vector, μ is the mean after fusion, σ is the variance after fusion, and F′ is the speech and facial feature vector;

[0014] Generate the historical fusion feature vector Z using the reparameterization formula;

[0015] Define the prior distribution p(Z) = U(0, I) using the standard normal distribution, where I is the identity variance matrix;

[0016] Calculate the KL divergence KL(q(Z│F′)||p(Z)) between the variational distribution q(Z│F′) and the prior distribution p(Z). The formula is:

[0017]

[0018] where K is the dimension of the historical fusion feature vector Z, and are the mean and variance of the variational distribution q(Z│F′) in the k-th dimension, respectively;

[0019] Define the KL divergence KL(q(Z│F′)||p(Z)) as the redundant information loss function L red , and the formula is:

[0020] L red = KL(q(Z│F′)||p(Z)),

[0021] Use the backpropagation algorithm to optimize the model parameters of the KL divergence and iteratively minimize the redundant information loss function L red ;

[0022] Set the number of principal components a using the cumulative variance contribution rate method. Calculate the covariance matrices of the speech and facial feature vectors respectively, decompose the covariance matrices to obtain the eigenvalues and eigenvectors of the speech and facial features, sort the eigenvalues in descending order, retain the eigenvectors corresponding to the first a largest eigenvalues to form the principal component matrix, and project the speech and facial feature vectors using the principal component matrix to obtain the latent representations of the speech and facial feature vectors;

[0023] Construct the energy function E(Z, Y) using the energy benchmark model. The formula is:

[0024] E(Z, Y) = -<Z, f(Y)> + b Z ,

[0025] where Z is the historical fusion feature vector, <Z, f(Y)> is the inner product of the historical fusion feature vector and the linear mapping function f(Y), f(Y) is the linear mapping function, and b Z is the energy benchmark bias term;

[0026] The linear mapping function f(Y) has the formula:

[0027] f(Y) = W y ·Y + b Y ,

[0028] where W y is the parameter matrix of the mapping, Y is the emotion label, and b Y is the bias term of the mapping;

[0029] Randomly extract historical fusion feature vectors and corresponding emotion labels from the training set, denoted as positive sample pairs, and randomly select another emotion label from the historical fusion feature vectors extracted in the same batch, denoted as negative sample pairs;

[0030] Define the energy difference loss function L info based on contrastive learning, with the formula:

[0031]

[0032] where N is the total number of samples, Z i ′ is the historical fusion feature vector of the i-th sample, Y i ′ is the emotion label of the i-th sample, Y j ′ is the incorrect emotion label paired with the historical fusion feature vector of the i-th sample, exp - E(Z i ′, Y i ′)) is the energy term of the positive sample, and ∑ j≠i exp(-E(Z i ′, Y i ′)) is the energy term of the negative sample;

[0033] Use the gradient descent method to optimize the parameters of the energy function, iteratively minimizing the energy difference loss function L info ;

[0034] Use multi-objective optimization to combine the redundant information loss function L red and the energy difference loss function L info to define the information bottleneck loss function L IB , with the formula:

[0035] L IB = L info + α·L red ,

[0036] where α is the trade-off coefficient;

[0037] Use the second-order optimization method for iterative optimization of the mean and variance after fusion;

[0038] Use the trained information bottleneck loss function L IB, obtain the real-time fusion feature vector z;

[0039] Sort the real-time fusion feature vector z in time order to form a time series z(t).

[0040] As a preferred solution of the mental health early warning method for mental health education according to the present invention, wherein: the calculation of the trajectory complexity by using delay embedding to construct phase space trajectory points refers to calculating the optimal delay time τ of the time series by using the mutual information method;

[0041] Calculate the embedding dimension m by using the FNN method;

[0042] Based on the optimal delay time τ and the embedding dimension m, use delay embedding to construct phase space trajectory points X(t);

[0043] Sort the phase space trajectory points X(t) in chronological order to form a trajectory sequence {X(t i )} =

[0044] {X(t 1 ), X(t 2 ), …, X(t A )}, A is the total number of phase space trajectory points, and t i is the i-th time point of the time series;

[0045] Connect each point in the phase space trajectory sequence in chronological order to form a closed polygon, and the vertices of the polygon are {X(t 1 ), X(t 2 ), …, X(t A )}, and use the polygon area formula to calculate the trajectory area S;

[0046] Use the Euclidean distance accumulation method to calculate the trajectory length L of adjacent phase space trajectory points;

[0047] Use the local curvature formula to calculate the local curvature of adjacent phase space trajectory points;

[0048] Calculate the average value of the local curvature C i to obtain the overall curvature C avg ;

[0049] Through the combination of the trajectory length L, the trajectory area S, and the overall curvature C avg , define the trajectory complexity D, and the formula is:

[0050]

[0051] where D is the trajectory complexity.

[0052] As a preferred solution of the mental health early warning method for mental health education according to the present invention, wherein: the number of clusters set for calculating the mental health score and early warning index by weighted summation is 3, and the K-means clustering algorithm is used to classify the emotional state of the trajectory complexity D, including stable, fluctuating, and severely fluctuating;

[0053] Use statistical analysis to calculate the proportion of emotional state time O of the classification result i ;

[0054] Use the information entropy formula to calculate the emotional state entropy value H;

[0055] Based on the trajectory complexity D and the emotional state entropy value H, use weighted summation to calculate the mental health score R;

[0056] Use the normal distribution method to set the health threshold ε, compare the mental health score R with the health threshold ε. If R>ε, it is determined as high risk, an early warning is issued and the monitor is notified by email. If R≤ε, it is determined as normal and the multimodal data is continuously monitored.

[0057] As a preferred solution of the mental health early warning method for mental health education according to the present invention, wherein: the construction of the visualization interface to display the mental health score in real time means using the front-end framework Vue.js to construct the visualization interface, displaying the mental health score in real time, and displaying a risk alarm logo above the visualization interface. When an early warning is issued, the alarm logo becomes red and flashes;

[0058] Allow users who have passed real-name verification to view.

[0059] As a preferred solution of the mental health early warning method for mental health education according to the present invention, wherein: the storage of the multimodal data generated by collection and analysis means storing the collected multimodal data and the mental health score generated by analysis in the central database. The central database is sorted in chronological order and marked with corresponding tags, and at the same time, the collected multimodal data and the mental health score generated by analysis are backed up to the cloud, and the integrity of the backup data is detected regularly.

[0060] Another object of the present invention is to provide a mental health early warning system for mental health education, which includes,

[0061] A collection and fusion module for collecting multimodal data and performing preprocessing, and extracting and fusing the feature vectors of the preprocessed multimodal data;

[0062] A calculation and early warning module for using delay embedding to construct phase space trajectory points to calculate the trajectory complexity, and using weighted summation to calculate the mental health score and early warning;

[0063] A visualization storage module for constructing a visualization interface to display mental health scores in real time and store the multi-modal data generated by collection, analysis.

[0064] A computer device includes: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, the steps of the mental health warning method for the above-mentioned mental health education are implemented.

[0065] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the mental health warning method for the above-mentioned mental health education are implemented.

[0066] The beneficial effects of the present invention are as follows: By collecting multi-modal data and performing preprocessing, extracting the feature vectors of the preprocessed multi-modal data for fusion; using delay embedding to construct phase space trajectory points to calculate the trajectory complexity, and using weighted summation to calculate mental health scores and warnings; improving the sensitivity to fluctuations in emotional states, enhancing the accuracy and timeliness of mental health warnings, and avoiding the lag of users and monitors in monitoring mental states. Description of the Drawings

[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0068] Figure 1 It is a flowchart of the mental health warning method for mental health education.

[0069] Figure 2 It is a structural diagram of the mental health warning system for mental health education. Detailed Embodiments

[0070] To make the above-mentioned objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings in the specification.

[0071] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention, but the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0072] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separate or selectively mutually exclusive with other embodiments.

[0073] Embodiment 1, referring to Figure 1 , is the first embodiment of the present invention. This embodiment provides a mental health early warning method for mental health education. The mental health early warning method for mental health education includes

[0074] S1. Collect multi-modal data and perform preprocessing, and extract the feature vectors of the preprocessed multi-modal data for fusion;

[0075] Specifically, collecting multi-modal data and performing preprocessing means that during the process of mental health education, a high-definition camera and a microphone are used to collect the user's multi-modal data and perform preprocessing;

[0076] The multi-modal data includes voice and facial expression data;

[0077] The preprocessing includes adding time stamps to the collected multi-modal data, using the linear interpolation method to perform time synchronization correction on the multi-modal data, using a band-pass filter to remove the noise of the voice data, performing normalization processing on the denoised voice data, using Gaussian filtering to remove the noise of the facial expression data, using the Laplace operator to sharpen the facial expression data, and using the affine transformation method to perform normalization alignment on the facial expression data.

[0078] By adding time stamps to the multi-modal data and using the linear interpolation method to perform time synchronization correction on the data, the present invention can ensure the time consistency of all data modalities, overcome the problem of differences in sampling frequencies among different modality data. By using a band-pass filter to remove the noise in the voice data, the present invention effectively reduces the environmental noise and interference signals outside the voice frequency band, thereby improving the purity of the voice data and enhancing the robustness of the emotion recognition model to low-frequency and high-frequency interference, providing a more reliable data source for voice emotion analysis. The combined application of Gaussian filtering and the Laplace operator can effectively remove noise and enhance edge information, making the details of facial expression features clearer. It not only helps to eliminate the noise in the facial image but also can improve the resolution of micro-expressions, providing richer image information for subsequent emotion recognition and micro-expression analysis. By using the affine transformation method to perform normalization alignment on the facial expression data, the present invention solves the problem of feature offset caused by differences in angle or position during the acquisition process. The aligned facial features provide a basis for the dynamic analysis of facial expression changes, helping to reveal the mental health fluctuations of individuals.

[0079] Furthermore, extracting the feature vectors of the preprocessed multi-modal data for fusion means collecting historical multi-modal data, preprocessing it, and extracting feature vectors to generate a training set;

[0080] Using the pre-trained FER+ model to classify the emotions of historical speech data and assign corresponding emotion labels, and using the pre-trained OpenSMILE model to classify the emotions of historical facial expression data and assign corresponding emotion labels;

[0081] Using the pre-trained Whisper model of OpenAI to extract the speech feature vectors of historical speech data, and using the pre-trained VGG-Face model to extract the facial feature vectors of historical facial expression data;

[0082] Introduce random noise ∈~U(0,1) from the standard normal distribution. Use a fully connected network to convert the historical speech and facial feature vectors into means and variances respectively, and use weighted average fusion to fuse the means and variances respectively to obtain the means and variances of historical fusion. Define the variational distribution q(Z│F′) = U(μ,σ 2 ) based on the fused means and variances, where U is the symbol of the normal distribution, Z is the historical fusion feature vector, μ is the fused mean, σ is the fused variance, and F′ is the speech and facial feature vectors;

[0083] Use the reparameterization formula to generate the historical fusion feature vector Z. The formula is:

[0084] Z = μ + σ·∈,

[0085] where ∈ is the random noise;

[0086] Define the prior distribution p(Z) = U(0,I) using the standard normal distribution, where I is the identity variance matrix to ensure that the variance of each dimension is 1 and avoid the accumulation of redundant variance information in the features;

[0087] Calculate the KL divergence KL(q(Z│F′)‖p(Z)) between the variational distribution q(Z│F′) and the prior distribution p(Z) to measure the redundancy of irrelevant information in the fusion features. The formula is:

[0088]

[0089] where K is the dimension of the historical fusion feature vector Z, and are the mean and variance of the variational distribution q(Z│F′) in the k-th dimension respectively;

[0090] The KL divergence has mathematical rigor and computability. The distribution optimized through the KL divergence is closer to the standard normal distribution. Although other measurement methods (such as the JS divergence) can also be used to compare distributions, their convergence characteristics and sensitivity to model training are relatively low, so the same effect cannot be achieved. By calculating the KL divergence, the distribution of historical comprehensive features is made close to the standard normal distribution, ensuring the regularization processing of features. This method improves the stability of the model in processing multi-modal emotion features and avoids the problem of poor model generalization caused by feature drift or high dimensionality. The application of the KL divergence not only improves the accuracy of comprehensive features but also ensures the sensitivity and adaptability of the model to emotional states, making the model perform more stably when dealing with complex emotion data.

[0091] Define the KL divergence KL(q(Z│F′)‖p(Z)) as the redundant information loss function L red , the formula is:

[0092] L red = KL(q(Z│F′)‖p(Z)),

[0093] Use the backpropagation algorithm to optimize the model parameters of the KL divergence and iteratively minimize the redundant information loss function L red ;

[0094] In the process of generating historical comprehensive features, reducing redundant information helps the model focus on key features. The redundant information loss function provides a means to directly optimize redundant information, ensuring the representativeness and conciseness of features. By defining the redundant information loss function, feature redundancy is quantified and minimized. In the prior art, the problem of feature redundancy is usually not considered. This improvement makes the generated historical comprehensive features reduce redundancy while maintaining information volume, providing a more concise and effective feature representation for subsequent emotion recognition. Compared with the prior art, the method of the present invention pays more attention to improving the effectiveness of features and the robustness of the model through redundant information loss.

[0095] Use the cumulative variance contribution rate method to set the number of principal components a. Calculate the covariance matrices of the speech and facial feature vectors respectively, decompose the covariance matrices to obtain the eigenvalues and eigenvectors of the speech and facial features, sort the eigenvalues in descending order, retain the eigenvectors corresponding to the first a largest eigenvalues, form the principal component matrix, and use the principal component matrix to project the speech and facial feature vectors to obtain the latent representations of the speech and facial feature vectors;

[0096] Use the energy benchmark model to construct the energy function, the formula is:

[0097] E(Z,Y)= -<Z,f(Y)>+b Z ,

[0098] where Z is the historical fusion feature vector, f(Y) is the linear mapping function, <Z, f(Y)> is the inner product of the historical fusion feature vector Z and the linear mapping function f(Y), which is used to measure the similarity between positive and negative samples, and b Z is the energy reference bias term, which is used to adjust the reference level of energy;

[0099] The energy function is the core mechanism for distinguishing positive and negative sample pairs. Through the energy reference model, the matching degree between different emotion labels and feature vectors can be effectively calculated, optimizing the accuracy and reliability of emotion classification. Using the inner product as the energy reference can retain the similarity information between emotion labels and feature vectors in the low-dimensional feature space, ensuring high accuracy in emotion recognition. Compared with the feature matching method that directly relies on distance metrics, the energy reference model constructs the energy term through the inner product, reducing the computational complexity and improving the processing speed of the model. By optimizing the energy term, the model has a higher anti-interference ability when processing the matching of emotion labels and feature vectors, enhancing the robustness of the model in multi-modal emotion recognition.

[0100] The linear mapping function f(Y) transforms the emotion label into the feature space, and the formula is:

[0101] f(Y) = W y ·Y + b Y ,

[0102] where W y is the parameter matrix of the mapping, Y is the emotion label, and b Y is the bias term of the mapping, which adjusts the reference of label mapping;

[0103] The introduction of the linear mapping function transforms the emotion label into the form of a feature vector. Through this mapping, the effective energy measurement between the emotion label and the historical fusion feature vector can be achieved, with low computational complexity and seamless integration with the energy reference model, ensuring the efficiency and robustness of the model. Different from the traditional processing method of discrete emotion labels, this method transforms the discrete label into a continuous feature representation through linear mapping, enabling the emotion label to better integrate into the feature space. This not only enhances the continuity of label expression but also improves the discrimination in the feature space, optimizing the accuracy of emotion classification.

[0104] Randomly extract the historical fusion feature vector and the corresponding emotion label from the training set, denoted as the positive sample pair, and randomly select another emotion label from the historical fusion feature vectors extracted in the same batch, denoted as the negative sample pair;

[0105] Define the energy difference loss function L info based on contrastive learning, so that the energy of the positive sample pair is as low as possible and the energy of the negative sample pair is as high as possible, thereby achieving an effective discrimination effect between positive and negative samples. The formula is:

[0106]

[0107] where N is the total number of samples, and Z i ′ is the historical fusion feature vector of the i-th sample, and Y i ′ is the emotion label of the i-th sample, and Y j ′ is the incorrect emotion label paired with the historical fusion feature vector of the i-th sample, and exp(-E(Z i ′, Y i ′)) is the energy term of the positive sample. By performing a negative exponential transformation to reduce the energy of the positive sample, the preference of the model for the positive sample is enhanced. ∑ j≠ i exp(-E(Z i ′, Y i ′)) is the energy term of the negative sample, and the energy terms of all negative sample pairs are summed to increase the energy of the negative samples;

[0108] The energy difference loss function is a core loss function based on contrastive learning, which is specifically used to optimize the energy difference between positive and negative samples. By optimizing the energy term of the positive sample and suppressing the energy term of the negative sample, the preference of the model for the positive sample can be effectively improved, thereby enhancing the accuracy and robustness of emotion recognition. Although other loss functions (such as the common cross-entropy loss) are applicable to classification problems, they cannot achieve energy optimization in contrastive learning as effectively. Therefore, in this case, only the energy difference loss function can be used to achieve the optimal effect. Compared with traditional loss functions, the present invention realizes the optimization of the energy difference between positive and negative samples through the energy difference loss function, strengthens the energy of positive samples and suppresses the energy of negative samples, enhances the discrimination ability of the model for positive samples, and thus improves the accuracy of emotion recognition. In the emotion recognition task, by optimizing the energy difference between positive and negative samples, the model can more accurately match emotion labels and fusion features, improve the adaptability of the model to complex emotion data, enhance the discrimination between different emotion labels and feature vectors, and increase the sensitivity and adaptability of the model to emotion states. This improvement enables the model to perform more robustly in the emotion recognition task of processing multi-modal data and better handle the diversity and complexity of emotion states.

[0109] Use the gradient descent method to optimize the parameters of the energy function and iteratively minimize the energy difference loss function L info ;

[0110] Use multi-objective optimization to combine the redundant information loss function L red and the energy difference loss function L info and define it as the information bottleneck loss function L IB , and the formula is:

[0111] LIB = L info + α·L red ,

[0112] where α is a trade-off coefficient that controls the intensity of redundant information compression;

[0113] Through multi-objective optimization, accurate extraction of effective features and effective compression of redundant information are achieved. The definition of this multi-objective loss can ensure that when extracting comprehensive features, both the expressiveness and simplicity of the features are taken into account. Compared with the prior art, the definition of the information bottleneck loss function combines the effectiveness and conciseness of the features, ensuring that the extracted comprehensive features reduce redundancy while maintaining information integrity. By optimizing the information bottleneck loss function, more expressive features can be obtained, thereby improving the accuracy of emotion recognition. By introducing a trade-off coefficient, controllable compression of redundant information is achieved, enabling the feature vector to maintain a concise and stable feature representation while meeting the requirements of emotion state expression. Through the multi-objective optimized information bottleneck loss function, the model can dynamically adjust the complexity and expressiveness of the features during the feature extraction process to adapt to different emotion states and diverse emotion data. Compared with the fixed feature extraction strategy of traditional methods, this flexibility makes the model more robust and reliable in complex scenarios.

[0114] Use the second-order optimization method to perform iterative optimization of the mean and variance after fusion;

[0115] Apply the trained information bottleneck loss function L to the multi-modal data collected in real time IB , to obtain the real-time fusion feature vector z;

[0116] Sort the real-time fusion feature vector z according to time to form a time series z(t).

[0117] The preprocessing steps improve the accuracy and consistency of the data through temporal calibration and noise reduction, laying a solid foundation for emotion feature extraction and subsequent emotion prediction. By improving the accuracy and consistency of the data through temporal calibration and noise reduction, a solid foundation is laid for emotion feature extraction and subsequent emotion prediction, effectively improving the quality of the dataset and the adaptability of the model. The weighted average fusion of speech and facial feature vectors helps to integrate multi-modal features, enabling the emotion recognition model to comprehensively evaluate the features of both modalities and improving the accuracy of emotion classification. The weighted average fusion can not only adapt to the contributions of speech and facial features in different emotional states, but also achieve a more flexible measure of emotional expression by controlling the weights, ensuring the multi-modal fusion effect of the model. Introducing random noise from the standard normal distribution and generating a variational distribution through reparameterization ensures the diversity of the data and the robustness of the model. Without distorting the data, it improves the generalization performance of the model in feature extraction, making the fused feature vector more representative and enhancing the sensitivity to fluctuations in emotional states, providing support for more accurate psychological state assessment. By using the KL divergence as the redundant information loss function, the model can minimize redundant information while retaining key features, optimizing the representation ability of the multi-modal fusion feature vector, reducing feature redundancy, avoiding over-reliance of the model on invalid data, improving the accuracy of emotion recognition, and making it more suitable for the assessment of mental health status. After retaining the principal components, it can effectively simplify the feature vector, improve the computational efficiency and prediction accuracy of the model, thus ensuring the real-time performance and applicability of the model in practical applications. The introduction of the energy difference loss function makes the model more sensitive in distinguishing positive and negative samples, reducing the interference of incorrect emotion labels and enhancing the ability to recognize real emotional states. The optimization process of contrastive learning can effectively improve the robustness and classification accuracy of the model, making it more stable in recognizing different psychological states. The information bottleneck loss function achieves a balance between information compression and emotion feature representation through multi-objective optimization by combining redundant information and energy difference loss functions. The optimized fused feature vector is more concise and accurate, avoiding the interference of redundant information on the model. The information bottleneck loss function enables the emotion feature vector to retain the core features of emotional fluctuations under limited information, providing a concise and reliable expression of emotional states for mental health assessment.

[0118] S2. Use delay embedding to construct phase space trajectory points to calculate the trajectory complexity, and use weighted summation to calculate the mental health score and early warning;

[0119] Specifically, using delay embedding to construct phase space trajectory points to calculate the trajectory complexity means using the mutual information method to calculate the optimal delay time τ of the time series, including

[0120] Calculate the τ mutual information value I(τ) of the calculation delay time, sort the τ mutual information values I(τ) of the delay time from small to large, and set the time corresponding to the τ mutual information value I(τ) with the smallest delay time as the optimal delay time τ;

[0121] Use the FNN method to calculate the embedding dimension m;

[0122] Based on the optimal delay time τ and the embedding dimension m, use delay embedding to construct the phase space trajectory point X(t), and the formula is:

[0123] X(t) = [z(t), z(t + τ), z(t + 2τ), …, z(t + (m - 1)τ)],

[0124] Sort the phase space trajectory points X(t) in chronological order to form a trajectory sequence {X(t i )} =

[0125] {X(t 1 ), X(t 2 ), …, X(t A )}, A is the total number of phase space trajectory points, t i is the i-th time point of the time series;

[0126] Set u as the y-axis, connect each point in the phase space trajectory sequence in chronological order to form a trajectory path that evolves over time. This path shows the dynamic behavior of the system in the phase space. Connect the last time point X(t A ) and the first time point X(t 1 ) in the time series to form a closed polygon. The vertices of the polygon are {X(t 1 ), X(t 2 ), …, X(t A )}. Use the polygon area formula to calculate the trajectory area S, and the formula is:

[0127]

[0128] where X x (t i ) is the coordinate of the phase space trajectory point X(t i ) on the x-axis, and U u (t i ) is the coordinate of the phase space trajectory point X(t i ) on the u-axis;

[0129] Use the Euclidean distance cumulative addition method to calculate the trajectory length L of adjacent phase space trajectory points, and the formula is

[0130]

[0131] where X j (t i ) is the coordinate value of the phase space trajectory point X(t i ) in the j-th dimension;

[0132] Calculate the local curvature C of adjacent phase space trajectory points using the local curvature formula i , and the formula is:

[0133]

[0134] where × is the vector cross product, representing the angle between two vectors and used to measure the degree of bending of the trajectory. ‖(X(t i+1 ) - X(t i ))‖ and ‖Xt i ) - X(t i-1 )‖ are the Euclidean distances between the phase space trajectory points X(t i+1 ) and X(t i ) and between the phase space trajectory points X(t i ) and X(t i-1 ), respectively;

[0135] Calculate the average value of the local curvature C i to obtain the overall curvature C avg , and the formula is:

[0136]

[0137] Define the trajectory complexity D by combining the trajectory length L, trajectory area S, and overall curvature C avg , and the formula is:

[0138]

[0139] where D is the trajectory complexity.

[0140] The trajectory complexity formula combines the three major indicators of trajectory length, area, and curvature. By comprehensively considering the fluctuation amplitude, spatial range, and change frequency, it can comprehensively and accurately quantify the complexity of emotional states. Other simple features (such as only using trajectory length or area) cannot simultaneously take into account the fluctuation intensity and change severity of emotional states, so they cannot achieve the same recognition effect. Curvature reflects the severity of trajectory changes and is of great significance in emotional fluctuation analysis. Relying solely on trajectory length or area will ignore the frequent fluctuations of emotions, while adding curvature can effectively make up for this deficiency, thus comprehensively reflecting the dynamic changes of emotional states, overcoming the limitations of traditional single-feature methods, making the quantification of emotional fluctuations more representative, improving the accuracy of the emotional recognition model, helping to identify potential mental health risks earlier and more accurately, and making the model more stable and reliable when dealing with complex emotional data.

[0141] The optimal delay time can avoid excessive redundant information while retaining the main features in the time series, making the trajectory of emotional fluctuations clearer. The invention can effectively capture the temporal characteristics of emotional fluctuations, thereby more accurately evaluating the changes in the user's mental state, providing in-depth data support in the time dimension for mental health assessment. After determining the appropriate embedding dimension, the present invention can better display the dynamic evolution of emotional states in the reconstructed phase space, avoiding the interference of false nearest neighbors and making the recognition of emotional states more accurate. The serialized trajectory structure intuitively reflects the amplitude and frequency of emotional state fluctuations, laying a foundation for the quantification of subsequent trajectory complexity and providing a more reliable data source for the real-time assessment of mental health status. The longer the trajectory length, the greater the fluctuation of the emotional state, which may reflect the instability of the mental state. As a measure of the intensity of emotional fluctuations, the trajectory length can help the mental health monitoring system accurately identify high-risk states and issue timely warnings. A larger trajectory area usually indicates a higher complexity of the emotional state, which helps to identify the user states with intense or diverse emotional fluctuations. The present invention quantifies the breadth of emotional fluctuations through area calculation, providing additional geometric features for mental health status assessment. Through the calculation of curvature, the present invention can analyze the characteristics of emotional fluctuations from a geometric perspective, providing an important reference basis for mental health early warning. As the core index of mental health assessment, trajectory complexity provides an innovative method to evaluate the complexity of the user's emotional state, with high practicality and scientific value.

[0142] Furthermore, weighted summation is used to calculate the mental health score and the early warning index. Based on expert opinions, the number of clusters is set to 3, and the K-means clustering algorithm is used to classify the emotional states of the trajectory complexity D, including stable, fluctuating, and violently fluctuating;

[0143] The statistical analysis method is used to calculate the proportion O of the emotional state time of the classification result i , and the formula is:

[0144]

[0145] where T i is the duration of the i-th type of emotional state, and T total is the total observation time;

[0146] The information entropy formula is used to calculate the emotional state entropy value H, and the formula is:

[0147] H = -∑ i O i log(O i ),

[0148] Based on the trajectory complexity D and the emotional state entropy value H, weighted summation is used to calculate the mental health score R;

[0149] Set the health threshold ε using the normal distribution method, compare the mental health score R with the health threshold ε. If R > ε, it is determined as high risk, an early warning is issued and the monitor is notified via email. If R ≤ ε, it is determined as normal, and the multimodal data continues to be monitored.

[0150] The classification method can intuitively reflect different patterns of emotional fluctuations, providing a basis for the pattern recognition and analysis of emotional fluctuations. By calculating the time proportion of each emotional state through statistical analysis methods, the duration and occurrence frequency of emotional states can be quantified. The indicators help to identify the stability and change trends of the user's emotional state. By calculating the entropy value of the emotional state using the information entropy formula, the complexity and uncertainty of the emotional state can be quantified. The entropy value calculation helps to accurately grasp the fluctuation characteristics of the emotional state. Especially in the case of frequent emotional fluctuations, the value of information entropy can reveal the fluctuation law of the emotional state, thus helping to more comprehensively understand the user's mental state. The emotional fluctuation amplitude, as a direct measure of the intensity of emotional fluctuations, provides key data support for the mental health score. The quantification of the fluctuation amplitude enables the system to objectively evaluate the intensity of emotional fluctuations, helping to identify abnormal emotional fluctuations in a timely manner. The mental health score is a quantitative representation of the mental state, which can provide an intuitive mental health assessment result for the user. The introduction of the mental health score simplifies the quantification process of the mental health state, enabling the system to quickly and accurately evaluate the user's mental state. This early warning mechanism can help the monitor intervene in a timely manner to avoid potential mental health crises. The setting of the health threshold provides a scientific judgment standard, providing an automated and efficient risk identification ability for mental health monitoring.

[0151] S3. Build a visual interface to display the mental health score in real time and store the multimodal data collected, analyzed.

[0152] Specifically, building a visual interface to display the mental health score in real time means using the front-end framework Vue.js to build a visual interface to display the mental health score in real time, and displaying a risk alert logo above the visual interface. When an early warning is issued, the alert logo turns red and flashes.

[0153] Allow users who have passed real-name verification to view.

[0154] The reactive feature of Vue.js ensures that data changes can be immediately reflected on the interface, eliminating the need for users to manually refresh the page when viewing mental health scores, thus enhancing the user experience. Compared with simple numerical displays, graphical interfaces can more vividly show the trends of mood fluctuations and changes in mental health scores, enabling non-professional users to easily understand the significance of mental health scores and thereby increasing their attention to their own mental health. The design of the red flashing indicator not only visually highlights the risk status but also has a psychological hint effect, enabling users to more intuitively feel the urgency of their mental state, greatly enhancing the warning effect of the system, providing an immediate feedback mechanism for mental health monitoring and intervention, allowing changes in mental health scores to be displayed on the interface in real time, and avoiding lags in monitoring the mental state for both users and monitors. It can not only effectively prevent the leakage of mental health data but also meet the requirements of privacy protection, providing users with a higher level of trust. The design of this interface enables non-professional users to easily get started, reducing the difficulty for users to understand mental health scores, making the mental health monitoring function more popular and user-friendly, enhancing the response speed of mental health monitoring, providing a more timely and scientific basis for mental health intervention, and avoiding the exacerbation of mental health problems caused by delayed warnings.

[0155] Furthermore, storing the multimodal data generated from collection and analysis means storing the collected multimodal data and the generated mental health scores in a central database. The central database sorts the data in chronological order and marks corresponding tags, and simultaneously backs up the collected multimodal data and the generated mental health scores to the cloud, and regularly conducts integrity checks on the backup data.

[0156] This not only improves the storage efficiency of data but also provides systematic data support for mental health research, helping to discover potential patterns of mood fluctuations in large-scale data analysis. Through this management method, the mental health assessment system can more accurately capture changes in users' mood fluctuations, providing a data basis for the dynamic analysis and trend prediction of mental health scores, helping to discover potential mental health risks. The multiple storage of backup data ensures that the system can still quickly recover in case of emergencies, endowing the mental health monitoring system with a high level of data protection ability, ensuring the long-term stability of mental health score data, supporting the system to obtain highly reliable data in long-term tracking analysis, and improving the overall data quality of mental health monitoring. The data security system can protect mental health data from external attacks or hardware failures, ensuring the security and integrity of mental health score data, and thereby providing reliable support for users' mental health management. The centralized management method of large-scale data provides a solid foundation for mental health research, promoting the development of mental health assessment technology.

[0157] Example 2, refer to Figure 2, which is the second embodiment of the present invention. This embodiment is different from the previous one and provides a mental health early warning system for mental health education, including

[0158] A collection and fusion module, which is used to collect multimodal data and perform preprocessing, and extract and fuse the feature vectors of the preprocessed multimodal data;

[0159] A calculation and early warning module, which is used to construct a phase space trajectory point using delay embedding to calculate the trajectory complexity, and use weighted summation to calculate the mental health score and early warning;

[0160] A visualization and storage module, which is used to construct a visualization interface to display the mental health score in real time and store the multimodal data generated by collection and analysis.

[0161] If the above functions are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0162] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0163] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for instance, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0164] It should be understood that the various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, the multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

Claims

1. A mental health early warning method for mental health education, characterized in that: include, Collect and preprocess multimodal data, extract feature vectors of preprocessed multimodal data for fusion; Use delay embedding to construct phase space trajectory points to calculate trajectory complexity, and use weighted summation to calculate mental health scores and warnings; Build a visual interface to display mental health scores in real time, and store the multimodal data collected and analyzed.

2. The mental health early warning method for mental health education according to claim 1, characterized in that: The collecting and preprocessing of multimodal data refers to using a high-definition camera and a microphone to collect and preprocess the multimodal data of the user during the mental health education process; The multimodal data includes voice and facial expression data; The preprocessing includes adding a timestamp to the collected multimodal data, using a linear interpolation method to perform time synchronization correction on the multimodal data, using a bandpass filter to remove noise from the speech data, standardizing the denoised speech data, using a Gaussian filter to remove noise from the facial expression data, using a Laplace operator to sharpen the facial expression data, and using an affine transformation method to standardize and align the facial expression data.

3. The mental health early warning method of mental health education as claimed in claim 2, characterized in that: The extracting and fusing the feature vectors of the pre-processed multimodal data refers to collecting historical multimodal data and performing pre-processing and feature vector extraction to generate a training set; Use the pre-trained FER+ model to classify the emotions of historical speech data and assign corresponding emotion labels, and use the pre-trained OpenSMILE model to classify the emotions of historical facial expression data and assign corresponding emotion labels; Use the pre-trained OpenAI Whisper model to extract speech feature vectors from historical speech data, and use the pre-trained VGG-Face model to extract facial feature vectors from historical facial expression data; Random noise ∈ is introduced from the standard normal distribution. The historical speech and facial feature vectors are converted into mean and variance respectively using a fully connected network. The mean and variance are fused using weighted average fusion to obtain the mean and variance of the historical fusion. Based on the fused mean and variance, the variational distribution q(Z│F′)=U(μ,σ 2 ), where U is the symbol of normal distribution, Z is the historical fusion feature vector, μ is the mean after fusion, σ is the variance after fusion, and F′ is the speech and facial feature vector; Generate the historical fusion feature vector Z using the reparameterization formula; Use the standard normal distribution to define the prior distribution p(Z)=U(0,I), where I is the unit variance matrix; Calculate the KL divergence KL(q(Z│F′)‖p(Z)) between the variational distribution q(Z│F′) and the prior distribution p(Z), the formula is: Where K is the dimension of the historical fusion feature vector Z, and are the mean and variance of the variational distribution q(Z│F′) in the kth dimension respectively; Define the KL divergence KL(q(Z│F′)‖p(Z)) as the redundant information loss function L red , the formula is: L red =KL(q(Z│F′)‖p(Z)), Use the back propagation algorithm to optimize the model parameters of the KL divergence and iteratively minimize the redundant information loss function L red ; The number of principal components a is set using the cumulative variance contribution method, the covariance matrix of the speech and facial feature vectors is calculated respectively, the covariance matrix is ​​decomposed to obtain the eigenvalues ​​and eigenvectors of the speech and face, the eigenvalues ​​are arranged in descending order, the eigenvectors corresponding to the first a largest eigenvalues ​​are retained to form a principal component matrix, and the speech and facial feature vectors are projected using the principal component matrix to obtain the potential representation of the speech and facial feature vectors; The energy function E(Z,Y) is constructed using the energy benchmark model. The formula is: E(Z,Y)=-<Z,f(Y)>+b Z , Where Z is the historical fusion feature vector,<Z,f(Y)> is the inner product of the historical fusion feature vector and the linear mapping function f(Y), f(Y) is the linear mapping function, b Z is the energy reference bias term; The linear mapping function f(Y) is as follows: f(Y)=W y ·Y+b Y , Where W y is the parameter matrix of the mapping, Y is the emotion label, b Y is the bias term of the mapping; Randomly extract the historical fusion feature vector and the corresponding emotion label from the training set, which are recorded as positive sample pairs. Randomly select another emotion label from the historical fusion feature vector extracted from the same batch, which are recorded as negative sample pairs. Defining the energy difference loss function L based on contrastive learning info , the formula is: Where N is the total number of samples, Z i ′ is the historical fusion feature vector of the i-th sample, Y i ′ is the emotion label of the i-th sample, Y j ′ is the wrong emotion label paired with the i-th sample history fusion feature vector, exp(-E(Z i ′,Y i ′)) is the energy term of the positive sample, ∑ j≠i exp(-E(Z i ′,Y i ′)) is the negative sample energy term; Use the gradient descent method to optimize the parameters of the energy function and iteratively minimize the energy difference loss function L info ; Use multi-objective optimization to transform the redundant information loss function L red And the energy difference loss function L info Combined, it is defined as the information bottleneck loss function L IB , the formula is: L IB =L info +α·L red , Where α is the trade-off coefficient; Use the second-order optimization method to iteratively optimize the fused mean and variance; The multimodal data collected in real time is used with the trained information bottleneck loss function L IB , get the real-time fusion feature vector z; The real-time fusion feature vector z is sorted by time to form a time series z(t).

4. The mental health early warning method for mental health education according to claim 3, characterized in that: The use of delay embedding to construct phase space trajectory points to calculate trajectory complexity refers to using the mutual information method to calculate the optimal delay time τ of the time series; Use the FNN method to calculate the embedding dimension m; Based on the optimal delay time τ and embedding dimension m, the phase space trajectory point X(t) is constructed using delay embedding; The phase space trajectory points X(t) are sorted in time order to form a trajectory sequence {X(t i )}= {X(t1),X(t2),…,X(t A )}, A is the total number of phase space trajectory points, t i is the i-th time point of the time series; Connect each point in the phase space trajectory sequence in time order to form a closed polygon with vertices {X(t1),X(t2),…,X(t A )}, use the polygon area formula to calculate the trajectory area S; The trajectory length L of adjacent phase space trajectory points is calculated using the Euclidean distance accumulation method; The local curvature of adjacent phase space trajectory points is calculated using the local curvature formula; Calculate the local curvature C i The average value of the overall curvature C avg ; By using the trajectory length L, trajectory area S and overall curvature C avg The combination of defines the trajectory complexity D, and the formula is: Where D is the trajectory complexity.

5. The mental health early warning method for mental health education according to claim 4, characterized in that: The weighted sum is used to calculate the mental health score and the early warning index. The number of clusters is set to 3, and the K-means clustering algorithm is used to classify the emotional state of the trajectory complexity D, including stable, fluctuating and violently fluctuating; Use statistical analysis to calculate the emotional state time proportion of the classification results. i ; Use the information entropy formula to calculate the emotional state entropy value H; The mental health score R is calculated using weighted summation based on trajectory complexity D and emotional state entropy value H; The normal distribution method is used to set the health threshold ε, and the mental health score R is compared with the health threshold ε. If R>ε, it is judged as high risk, an early warning is issued and the monitor is notified via email. If R≤ε, it is judged as normal and multimodal data continues to be monitored.

6. The mental health early warning method for mental health education according to claim 5, characterized in that: The construction of a visual interface to display the mental health score in real time refers to using the front-end framework Vue.js to build a visual interface to display the mental health score in real time, displaying a risk alert icon above the visual interface, and the alert icon turns red and flashes when an early warning is issued; Users who have passed real-name verification are allowed to view the information.

7. The mental health early warning method for mental health education according to claim 6, characterized in that: The storage of the multimodal data collected and generated by analysis refers to storing the collected multimodal data and the mental health scores generated by the analysis in a central database, sorting the data in chronological order and marking the corresponding tags in the central database, synchronously backing up the collected multimodal data and the mental health scores generated by the analysis in the cloud, and regularly checking the integrity of the backup data.

8. A mental health early warning system for mental health education based on the mental health early warning method for mental health education according to any one of claims 1 to 7, characterized in that: include, The collection and fusion module is used to collect and preprocess multimodal data, extract the feature vectors of the preprocessed multimodal data for fusion; A calculation and early warning module is used to construct phase space trajectory points using delay embedding to calculate trajectory complexity, and to calculate mental health scores and early warnings using weighted summation; The visualization storage module is used to build a visualization interface to display mental health scores in real time and to store, collect and analyze multimodal data.

9. A computer device comprising: A memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, the steps of the mental health early warning method of mental health education described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the mental health early warning method for mental health education described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • MD patient emotion fluctuation monitoring and affective disorder state evaluation method and system

    CN115517681A

  • Multi-modal mental health prediction method and system

    CN118136256A

  • Psychological health condition general screening and evaluation method, system, equipment and medium

    CN118866364A

  • Multi-modal fusion-based VR emotion recognition and response system and method

    CN119066606A