Method for multi-modal and multi-dimensional analysis and early warning of psychological health of students by using AI (Artificial Intelligence)

Through multimodal data analysis of students' facial micro-expression, body language and speech information, combined with adaptive weighting algorithms and time convolutional networks, the problems of inaccurate emotion recognition and lack of personalized intervention in the existing technology are solved, and accurate monitoring and early warning of students' mental health are achieved.

CN120408371AInactive Publication Date: 2025-08-01HANGZHOU LINGDONG DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510500918.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing student mental health monitoring methods rely on a single data source, resulting in inaccurate emotional recognition and lack of dynamic adjustment of multimodal data, making it difficult to provide personalized interventions, and fail to achieve early warning and prevention.

Method used

Through multimodal data, students' facial micro-expression, body language and speech information are analyzed, and combined with adaptive weighting algorithms and time convolutional networks, a dynamic emotion trajectory tracking model is constructed to generate a personalized intervention plan.

Benefits of technology

Accurate monitoring and early warning of students' emotional changes is achieved, personalized psychological intervention is provided, and the deviation of a single mode is avoided, ensuring high accuracy and timely intervention in emotional recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408371A_ABST
    Figure CN120408371A_ABST
Patent Text Reader

Abstract

The invention discloses a method for multi-modal and multi-dimensional analysis of student mental health and early warning by using AI. The method comprises the following steps: acquiring three data streams of a student through a multi-source data sensor; respectively extracting data features, constructing a sub-modal emotion analysis model, and obtaining possibility probabilities of various emotions; dynamically adjusting the confidence coefficient weight of each mode through an adaptive weighting algorithm, and judging and obtaining the real emotion of the student; constructing a dynamic emotion trajectory tracking model based on a time convolutional network, calculating to obtain a student emotion stability index, and judging whether the student needs intervention and guidance based on the emotion stability index; and generating a personalized intervention scheme based on the dynamic emotion track of the student needing to be intervened and guided. The method has the advantages that facial micro-expressions, body languages and voice information of students are analyzed through multi-modal data, emotional fluctuation is accurately recognized, early warning is generated, and psychological health problems of the students are effectively intervened through dynamic emotional trajectory tracking and a personalized intervention scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to multimodal data processing, and particularly to a method for using AI multimodal and multi-dimensional analysis to analyze students' mental health and give early warnings. Background Art

[0002] With the rapid development of society and the continuous change of the educational environment, the mental health problems of students have become a topic of widespread concern. Especially in the context of increasingly fierce modern educational competition, many students are facing multiple pressures from family, school and society. Psychological problems such as anxiety, depression, mood swings, and social barriers are becoming increasingly common among students, having a serious impact on their physical and mental health. These psychological problems will not only lead to a decline in students' academic performance, but may also have a profound impact on their long-term development. Many students often lack effective emotion regulation and psychological coping strategies when facing academic pressure and interpersonal relationship problems, resulting in the further aggravation of mental health problems.

[0003] Most of the current mental health monitoring methods on the market usually rely on a single data source, such as facial expression or voice analysis, which often leads to inaccuracies in emotion recognition. For example, facial expressions may not fully reflect a student's inner state, especially when the emotional changes are relatively subtle; and relying solely on voice data is also easily affected by factors such as environmental noise and speaking speed, affecting the reliability of the results. In addition, existing emotion analysis methods usually do not dynamically adjust the modal weights by combining multimodal data, and may misjudge the true emotions of students in some cases. Many traditional methods lack long-term tracking and analysis of students' emotional changes, and it is also difficult to provide personalized intervention plans. Usually, they can only respond when the emotional problems are serious, and cannot achieve early warning and prevention. Therefore, the existing methods often cannot achieve comprehensive and accurate monitoring of students' emotional changes, and lack targeted and timely intervention measures. Summary of the Invention

[0004] In order to improve the existing method for generating early warning schemes for students' mental health, a method for using AI multimodal and multi-dimensional analysis to analyze students' mental health and give early warnings is provided. This method analyzes students' facial micro-expressions, body language and voice information through multimodal data, accurately identifies emotional fluctuations and generates early warnings, and uses dynamic emotion trajectory tracking and personalized intervention plans to effectively monitor and intervene in students' mental health problems.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A method for using AI multimodal and multi-dimensional analysis to analyze students' mental health and give early warnings, including:

[0007] Collect three data streams of students through multi-source data sensors, including: facial micro-expression video stream, limb behavior expression video stream, and interactive audio data stream;

[0008] Based on the three obtained data streams, extract data features respectively, and construct a multi-modal emotion analysis model to obtain the probability of various emotions;

[0009] Based on the probability of emotions of the three data streams, dynamically adjust the confidence weights of each modality through an adaptive weighting algorithm to judge and obtain the real emotions of students;

[0010] Based on the real emotions of students obtained from real-time analysis, construct a dynamic emotion trajectory tracking model based on a temporal convolutional network, calculate and obtain the student emotion stability index, and judge whether students need intervention and guidance based on the emotion stability index;

[0011] Based on the dynamic emotion trajectory of students who need intervention and guidance, generate a personalized intervention plan including cognitive behavioral therapy and mindfulness training.

[0012] Preferably, the three data streams of students collected through multi-source data sensors, including: facial micro-expression video stream, limb behavior expression video stream, and interactive audio data stream, specifically include:

[0013] Obtain the dynamic video stream of the student's facial area, enhance the local contrast through the CLAHE algorithm, and enhance the micro-expression;

[0014] Shoot the limb behavior of students through a multi-angle depth camera;

[0015] Obtain interactive audio data through a microphone array, suppress noise through beamforming, and enhance the directional sound source.

[0016] Preferably, the process of extracting data features respectively based on the three obtained data streams, constructing a multi-modal emotion analysis model, and obtaining the probability of various emotions specifically includes:

[0017] Based on the obtained facial micro-expression video stream, segment it into 0.5-second video clips;

[0018] Enhance the local contrast through the CLAHE algorithm to enhance the micro-expression, and locate the feature points of the face through the FAN network;

[0019] Based on the obtained limb behavior expression video stream, normalize the joint point coordinates to the position relative to the center of the torso;

[0020] Obtain key behavior features based on the action movement trajectory, including: joint angular velocity, limb movement amplitude, action frequency;

[0021] Based on the obtained behavioral characteristics, construct a behavioral coding dictionary, including various types of student actions.

[0022] Based on the obtained interactive audio data stream, perform multi-dimensional feature extraction, including yes / no features, frequency shift features, and semantic features.

[0023] Obtain the speech vitality index through the standard deviation of the speech rate and fundamental frequency, and perform silence interval analysis based on the silence duration.

[0024] Based on the feature data of the three obtained data streams, through a two-stream neural network, and combined with the existing emotion labels, train a cross-modal emotion analysis model.

[0025] Based on the trained cross-modal emotion analysis model, output the probability of various emotions of students in each data stream, and obtain the probability distribution.

[0026] Preferably, for the probability of emotions based on the three data streams, dynamically adjust the confidence weights of each modality through an adaptive weighting algorithm, and the specific steps for judging and obtaining the true emotions of students include:

[0027] Based on the probability distribution of emotions of the three obtained data streams, calculate the confidence of each modality.

[0028] Based on the confidence of each modality, assign weights to its corresponding modality data, and dynamically adjust the weight of each modality based on the change of modality confidence.

[0029] Obtain the final emotion prediction result by weighted fusion of the output probabilities of each modality.

[0030] Preferably, for calculating the confidence of each modality based on the probability distribution of emotions of the three obtained data streams, the specific steps include:

[0031] Calculate the classification loss based on the difference between the probability distribution of emotions of each modality in the historical data and the true label.

[0032] Evaluate the confidence of the modality based on the obtained classification loss value.

[0033] Add the probability distribution of emotions obtained from this confidence evaluation to the historical data to update the historical data.

[0034] Preferably, for constructing a dynamic emotion trajectory tracking model based on a temporal convolutional network based on the true emotions of students obtained from real-time analysis, calculating the student emotion stability index, and judging whether the student needs intervention and guidance based on the emotion stability index, the specific steps include:

[0035] Based on the real emotions of students obtained by analyzing each moment, construct a time series of students' real emotions, and train through a temporal convolutional network to construct a dynamic emotion trajectory tracking model;

[0036] Based on the trained dynamic emotion trajectory tracking model, quantify the change of students' emotions over time, and calculate the student stability index;

[0037] Based on the emotion stability index, set an intervention threshold, judge whether it is necessary to provide intervention guidance to students, and set an intervention warning.

[0038] Preferably, the process of quantifying the change of students' emotions over time and calculating the student stability index based on the trained dynamic emotion trajectory tracking model specifically includes:

[0039] Based on the trained dynamic emotion trajectory tracking model, obtain the emotion change trajectory function of students, and calculate the emotion fluctuation amplitude and the average emotion change rate;

[0040] Calculate and obtain the emotion stability index based on the emotion fluctuation amplitude and the average emotion change rate respectively;

[0041] Fuse the calculation results of the emotion fluctuation amplitude and the change rate to obtain the final emotion stability index.

[0042] Compared with the prior art, the advantages of the present invention are as follows:

[0043] By comprehensively integrating facial micro-expressions, body behaviors, and interactive audio data streams, the emotional expressions of students are comprehensively captured, avoiding biases that may be brought by a single modality. Secondly, the adaptive weighted algorithm is used to dynamically adjust the confidence of each modality, which can more accurately judge the real emotions of students and ensure high accuracy of emotion recognition. Furthermore, through the dynamic emotion trajectory tracking model based on the temporal convolutional network, the change of students' emotions can be monitored in real time, calculate the emotion stability index, and provide a scientific basis for early intervention. This method also combines cognitive behavioral therapy and mindfulness training to generate personalized intervention programs, which helps to provide effective psychological support at the initial stage of students' emotional problems and prevent the occurrence of serious mental health problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Schematic diagram of the method proposed by the present invention;

[0045] Figure 2 Schematic diagram of data collection proposed by the present invention;

[0046] Figure 3 Schematic diagram of obtaining the emotion probability distribution proposed by the present invention;

[0047] Figure 4 Schematic diagram of obtaining the real emotion proposed by the present invention;

[0048] Figure 5 Schematic diagram of calculating the confidence level of each modality proposed by the present invention;

[0049] Figure 6 Schematic diagram of intervention guidance judgment proposed by the present invention;

[0050] Figure 7 Schematic diagram of calculating the student stability index proposed by the present invention;

[0051] Figure 8 Architecture diagram of the electronic device in this solution;

[0052] Figure 9 Schematic diagram of the structure of the computer-readable storage medium in this solution. Detailed implementation manners

[0053] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.

[0054] Refer to Figure 1 As shown, the method for analyzing and warning students' mental health using AI multi-modal and multi-dimensional includes:

[0055] Step 1: Collect and obtain three data streams of students through multi-source data sensors, including: facial micro-expression video stream, limb behavior expression video stream, and interactive audio data stream;

[0056] Step 2: Based on the three obtained data streams, extract data features respectively, and construct a sub-modal emotion analysis model to obtain the probability of various emotions;

[0057] Step 3: Based on the probabilities of emotions of the three data streams, dynamically adjust the confidence weights of each modality through an adaptive weighting algorithm to judge and obtain the true emotions of students;

[0058] Step 4: Based on the true emotions of students obtained through real-time analysis, construct a dynamic emotion trajectory tracking model based on a temporal convolutional network, calculate and obtain the student emotion stability index, and judge whether students need intervention guidance based on the emotion stability index;

[0059] Step 5: Based on the dynamic emotion trajectories of students who need intervention guidance, generate a personalized intervention plan including cognitive behavioral therapy and mindfulness training.

[0060] Refer to Figure 2 As shown, collecting and obtaining three data streams of students through multi-source data sensors, including: facial micro-expression video stream, limb behavior expression video stream, and interactive audio data stream specifically includes:

[0061] Obtain the dynamic video stream of the student's facial area, enhance the local contrast through the CLAHE algorithm, and enhance the micro-expressions.

[0062] Capture the student's limb behavior through a multi-angle depth camera.

[0063] Obtain the interactive audio data through a microphone array, suppress the noise through beamforming, and enhance the directional sound source.

[0064] Specifically, use CLAHE to increase the contrast of the local area of the image while avoiding excessive enhancement of the global contrast. Let the local histogram of the image be h(x, y), where x and y are the pixel positions in the image. Calculate the local contrast enhancement, and the formula is:

[0065]

[0066] Among them, I is the pixel value of the original image, μ and σ are the mean and standard deviation of the local area respectively, and ∈ is a small constant used to avoid division-by-zero errors;

[0067] After enhancing the local details through CLAHE, the changes in micro-expressions can be seen more clearly, such as the subtle movements of the eyes or mouth;

[0068] Use multiple depth cameras to capture the student's limb behavior from different angles. The depth camera can obtain the depth information of each pixel, form a three-dimensional image, and can identify the three-dimensional structure and posture of the human body. Analyze the movement and posture of the limbs according to the relative position and angle changes between the joints. For example, use the angle change formula to calculate the rotation angle of the joint, and the formula is:

[0069]

[0070] Among them, θ is the rotation angle of the joint, v1 and v2 are two vectors connecting the joints, |v1| and |v2| are the magnitudes of the vectors, · is the dot product, and cos -1 Convert the scalar value of the dot product into an included angle;

[0071] When obtaining the interactive audio data, beamforming is used to weight the signals of the array microphones to enhance the sound from a specific direction and suppress the noise from other directions.

[0072] Refer to Figure 3 As shown, based on the three obtained data streams, extract the data features respectively, and build a multi-modal emotion analysis model to obtain the probability of various emotions. Specifically include:

[0073] Based on the obtained facial micro-expression video stream, segment it into 0.5-second video clips;

[0074] Enhance the local contrast through the CLAHE algorithm, enhance the micro-expressions, and locate the facial feature points through the FAN network;

[0075] Based on the obtained limb behavior expression video stream, normalize the joint point coordinates to the position relative to the torso center;

[0076] Obtain key behavior features based on the action movement trajectory, including: joint angular velocity, limb movement amplitude, action frequency;

[0077] Based on the obtained behavior features, construct a behavior encoding dictionary, including various types of student actions;

[0078] Based on the obtained interactive audio data stream, perform multi-dimensional feature extraction, including yes / no feature, frequency shift feature, semantic feature;

[0079] Obtain the speech vitality index through the standard deviation of the speech rate and the fundamental frequency, and perform silence interval analysis based on the silence duration;

[0080] Based on the feature data of the three obtained data streams, through a two-stream neural network, and combined with the existing emotion labels to train a multi-modal emotion analysis model;

[0081] Based on the trained multi-modal emotion analysis model, output the probability of various emotions of students in each data stream, and obtain the probability distribution.

[0082] Specifically, use the FAN network to locate the key points on the face. The FAN extracts deep features from the facial image through a convolutional neural network, and accurately locates 68 or 98 key points on the face, calculates the two-dimensional coordinates of each key point, such as the corners of the mouth, the corners of the eyes, etc. By calculating the facial region features, it is possible to better detect micro-expression changes;

[0083] Obtain the joint point coordinates of each frame, and estimate the joint angular velocity by calculating the angular change rate of the joint. The formula is:

[0084]

[0085] where ω j is the joint angular velocity, θ j (t) is the joint angle at time t, and Δt is the time interval;

[0086] Obtain the movement amplitude of the joint by calculating the change in joint position. The formula is:

[0087]

[0088] Wherein, A is the movement amplitude of the joint, (x(0), y(0)) is the initial position of the joint, and (x(t), y(t)) is the position at time t;

[0089] Based on the periodic motion of the joint, frequency-domain information is extracted through Fourier transform to obtain the frequency distribution of the motion. The formula is:

[0090]

[0091] Wherein, is the frequency distribution, x(t) is the joint position in the time domain, and ω is the frequency;

[0092] According to features such as joint angle, angular velocity, and movement amplitude, the student's actions are classified into different categories. Each action category is represented by a vector, and the elements in the vector represent different action features. A label is assigned to each action category, and the action features are encoded to construct a behavior encoding dictionary, which can be stored in the form of a matrix or a table. Each row represents a different action category, and each column represents a feature;

[0093] The video (facial micro-expressions, body behaviors) and audio data streams are processed simultaneously by a two-stream neural network. Each data stream is respectively subjected to feature extraction through its own neural network, and then the features of each stream are concatenated (fused) to a shared layer for emotion analysis;

[0094] Existing emotion labels include joy, anger, sadness, etc.

[0095] Refer to Figure 4 As shown, based on the probability of emotions of the three data streams, the confidence weights of each modality are dynamically adjusted through an adaptive weighting algorithm. Judging and obtaining the true emotions of the student specifically includes:

[0096] Based on the probability distribution of emotions of the three obtained data streams, the confidence of each modality is calculated;

[0097] Based on the confidence of each modality, weights are assigned to its corresponding modality data, and the weights of each modality are dynamically adjusted based on the change of the modality confidence;

[0098] The final emotion prediction result is obtained by weighted fusion of the output probabilities of each modality.

[0099] Specifically, as time goes by, the confidence of each modality may change. In order to dynamically adjust the weight of each modality, a moving average method based on a time window is designed to smooth the change of confidence. The formula is:

[0100]

[0101] Wherein, is the normalized confidence at time t, λ is the smoothing factor, which determines the weight ratio between the historical confidence and the current confidence, and α m (t) is the confidence before normalization;

[0102] Based on the weight ω of each modality m , that is, the normalized confidence, it is fused with the emotion probability distribution of each modality to obtain the final emotion prediction probability distribution. The formula is:

[0103]

[0104] where, P m is the emotion probability distribution of each modality, P final is the final emotion prediction probability distribution, ω m is the weight of each modality, and M is the total number of modalities.

[0105] Refer to Figure 5 As shown, based on the emotion likelihood probability distributions of the three data streams obtained, calculating the confidence of each modality specifically includes:

[0106] Calculating the classification loss based on the difference between the emotion likelihood probability distribution of each modality in the historical data and the true label;

[0107] Evaluating the confidence of the modality based on the obtained classification loss value;

[0108] Adding the emotion likelihood probability distribution obtained from this confidence evaluation to the historical data to update the historical data.

[0109] Specifically, the classification loss generally uses the cross-entropy loss function, which can measure the difference between the predicted probability distribution and the true label distribution. The formula is:

[0110]

[0111] where, L m is the difference between the predicted probability distribution and the true label distribution, y c is the encoding of the true label. If the true label is category c, then y c = 1, otherwise it is 0, is the predicted probability of modality m for emotion category c, and C is the total number of emotions;

[0112] Evaluating the confidence of each modality based on the classification loss. The smaller the loss, the more confident the model is in the emotion prediction of this modality. Therefore, the confidence of the modality can be evaluated by the reciprocal of the loss value;

[0113] After each emotion prediction, we add the confidence of this modality and the corresponding emotion probability distribution to the historical data. The update of the historical data includes adding the emotion probability distribution of the current modality; adding the confidence of this modality; adding the current true label.

[0114] Refer to Figure 6 As shown, based on the real emotions of students obtained from real-time analysis, a dynamic emotion trajectory tracking model based on a temporal convolutional network is constructed, the student emotion stability index is calculated, and based on the emotion stability index, it is judged whether students need intervention guidance, specifically including:

[0115] Based on the real emotions of students obtained at each moment, a time series of students' real emotions is constructed and trained through a temporal convolutional network to construct a dynamic emotion trajectory tracking model;

[0116] Based on the trained dynamic emotion trajectory tracking model, the situation of students' emotions changing over time is quantified, and the student stability index is calculated;

[0117] Based on the emotion stability index, an intervention threshold is set to judge whether students need intervention guidance and an intervention warning is set.

[0118] Specifically, according to the emotion probability distribution at each moment, an emotion time series E = {P1, P2, …, P T} of students is constructed, specifically:

[0119]

[0120] Among them, each row P t corresponds to the emotion prediction probability at the t-th moment;

[0121] The emotion stability index can be obtained by calculating the volatility of the emotion time series. The formula is:

[0122]

[0123] Among them, is the probability of the c-th type of emotion of the student at the t-th moment, is the average emotion probability at the t-th moment. The smaller the emotion stability index S, the more stable the student's emotion and the smaller the fluctuation; the larger S is, the greater the emotion fluctuation. T is the total length of the time period, and C is the total number of emotions;

[0124] To further quantify the stability of emotions, the emotion trajectory can be smoothed to obtain a smoothed emotion trajectory sequence;

[0125] Based on the emotion stability index S, an intervention threshold can be set. When the emotion stability index S of a student exceeds this threshold, it indicates that the student's emotion has a large fluctuation and may need intervention.

[0126] Refer to Figure 7 As shown, based on the trained dynamic emotion trajectory tracking model, the quantification of the change of students' emotions over time and the calculation of the student stability index specifically include:

[0127] Based on the trained dynamic emotion trajectory tracking model, obtain the emotion change trajectory function of students, and calculate the emotion fluctuation amplitude and the average emotion change rate;

[0128] Calculate and obtain the emotion stability index based on the emotion fluctuation amplitude and the average emotion change rate respectively;

[0129] Fuse the calculation results of the emotion fluctuation amplitude and the change rate to obtain the final emotion stability index.

[0130] Specifically, the emotion fluctuation amplitude is the maximum change amplitude of the emotion within a given time interval, indicating the fluctuation degree of the students' emotions. The average emotion change rate is an index measuring the change speed of the emotion over the entire time period, reflecting the average change of the students' emotions per unit time. By means of weighted average, the two stability indexes are fused to finally obtain a comprehensive emotion stability index.

[0131] Furthermore, the method according to the implementation manner of the present application can also be implemented by means of Figure 8 the architecture of the electronic device shown. As Figure 8 shown, the electronic device 500 may include a bus 501, one or more CPUs 502, a read-only memory (ROM) 503, a random access memory (RAM) 504, a communication port 505 connected to the network, an input / output component 506, a hard disk 507, etc. The storage device in the electronic device 500, such as the ROM 503 or the hard disk 507, may store the method for using AI multi-modal and multi-dimensional analysis of students' mental health and early warning provided by the present application. The electronic device 500 may further include a terminal interface 508. Of course, Figure 8 the architecture shown is only exemplary. When implementing different devices, one or more components shown in the electronic device may be omitted according to actual needs. Figure 8

[0132] Figure 9 is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of the present application. As Figure 9 ​As shown, there is a computer-readable storage medium 600 according to an embodiment of the present application. Computer-readable instructions are stored on the computer-readable storage medium 600. When the computer-readable instructions are run by a processor, the method of using AI multi-modal and multi-dimensional analysis to analyze students' mental health and give early warnings according to the embodiments of the present application described with reference to the above drawings can be executed. The storage medium 600 includes but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0133] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.

[0134] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized.

[0135] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for analyzing students' mental health and giving early warnings using AI multi-modal and multi-dimensional analysis, characterized in that, Including: Collect three data streams of students through multi-source data sensors, including: facial micro-expression video stream, limb behavior expression video stream, and interactive audio data stream; Based on the three obtained data streams, extract data features respectively, and construct a multi-modal emotion analysis model to obtain the probability of various emotions; Based on the probability of emotions of the three data streams, dynamically adjust the confidence weights of each modality through an adaptive weighting algorithm to judge and obtain the true emotions of students; Based on the true emotions of students obtained from real-time analysis, construct a dynamic emotion trajectory tracking model based on a temporal convolutional network, calculate and obtain the student emotion stability index, and judge whether students need intervention and guidance based on the emotion stability index; Based on the dynamic emotion trajectory of students who need intervention and guidance, generate a personalized intervention plan including cognitive behavioral therapy and mindfulness training.

2. The method for analyzing and warning students' mental health using AI multi-modal and multi-dimensional analysis according to claim 1, characterized in that, The three data streams of students collected through multi-source data sensors, including: facial micro-expression video stream, limb behavior expression video stream, and interactive audio data stream, specifically include: Obtain the dynamic video stream of the student's facial area, enhance the local contrast through the CLAHE algorithm, and enhance the micro-expression; Shoot the student's limb behavior through a multi-angle depth camera; Obtain interactive audio data through a microphone array, suppress noise through beamforming, and enhance the directional sound source.

3. The method for analyzing and warning students' mental health using AI multi-modal and multi-dimensional analysis according to claim 1, characterized in that, Based on the three obtained data streams, extract data features respectively, and construct a multi-modal emotion analysis model to obtain the probability of various emotions, specifically including: Based on the obtained facial micro-expression video stream, segment it into 0.5-second video clips; Enhance the local contrast through the CLAHE algorithm, enhance the micro-expression, and locate the feature points on the face through the FAN network; Based on the obtained limb behavior expression video stream, normalize the joint coordinates to the position relative to the center of the torso; Obtain key behavior features based on the action movement trajectory, including: joint angular velocity, limb movement amplitude, action frequency; Based on the obtained behavior features, construct a behavior coding dictionary, including various point-type student actions; Based on the obtained interactive audio data stream, perform multi-dimensional feature extraction, including yes / no feature, frequency shift feature, semantic feature; Obtain the speech vitality index through the standard deviation of speech rate and fundamental frequency, and perform silence interval analysis based on the silence duration; Based on the feature data of the three obtained data streams, through a two-stream neural network, and combined with existing emotion labels to train a multi-modal emotion analysis model; Based on the trained multi-modal emotion analysis model, output the probability of various emotions of students in each data stream, and obtain the probability distribution.

4. The method for using AI multi-modal and multi-dimensional analysis to analyze students' mental health and give early warnings according to claim 1, wherein Based on the probability of emotions of the three data streams, dynamically adjust the confidence weights of each modality through an adaptive weighting algorithm to judge and obtain the true emotions of students, specifically including: Based on the probability distribution of emotions of the three obtained data streams, calculate the confidence of each modality; Assign weights to the corresponding modality data based on the confidence of each modality, and dynamically adjust the weight of each modality based on the change of modality confidence; Obtain the final emotion prediction result by weighted fusion of the output probabilities of each modality.

5. The method for using AI multi-modal and multi-dimensional analysis to analyze students' mental health and give early warnings according to claim 4, wherein, Calculating the confidence of each modality based on the obtained probability distribution of emotional possibilities of the three data streams specifically includes: Calculating the classification loss based on the difference between the probability distribution of emotional possibilities of each modality in the historical data and the true label; Evaluating the confidence of the modality based on the obtained classification loss value; Adding the probability distribution of emotional possibilities obtained from this confidence evaluation to the historical data to update the historical data.

6. The method for using AI multi-modal and multi-dimensional analysis to analyze students' mental health and give early warnings according to claim 1, characterized in that, Constructing a dynamic emotional trajectory tracking model based on a temporal convolutional network based on the real-time analyzed true emotions of students, calculating the student emotional stability index, and determining whether students need intervention and guidance based on the emotional stability index specifically includes: Constructing a time series of the true emotions of students based on the true emotions of students analyzed at each moment, and training through a temporal convolutional network to construct a dynamic emotional trajectory tracking model; Quantifying the change of students' emotions over time based on the trained dynamic emotional trajectory tracking model, and calculating the student stability index; Based on the emotional stability index, setting an intervention threshold, determining whether it is necessary to intervene and guide students, and setting an intervention warning.

7. The method for using AI multi-modal and multi-dimensional analysis to analyze students' mental health and give early warnings according to claim 6, characterized in that, Quantifying the change of students' emotions over time based on the trained dynamic emotional trajectory tracking model, and calculating the student stability index specifically includes: Based on the trained dynamic emotional trajectory tracking model, obtaining the emotional change trajectory function of students, and calculating the emotional fluctuation amplitude and the average emotional change rate; Calculating and obtaining the emotional stability index based on the emotional fluctuation amplitude and the average emotional change rate respectively; Fusing the calculation results of the emotional fluctuation amplitude and the change rate to obtain the final emotional stability index.

8. An electronic device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for analyzing and warning students' mental health using AI multi-modal and multi-dimensional analysis as described in any one of claims 1-7.

9. A computer-readable storage medium storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by the processor, the method for analyzing and warning students' mental health using AI multi-modal and multi-dimensional analysis as described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on micro-expressions, body movements and voices

    CN113469153A

  • Non-contact psychological state assessment method and system based on multi-modal fusion technology

    CN117936032A

  • AI-based pet emotion recognition system

    CN119049086A

  • Multi-modal classroom emotion recognition method and system based on modal adaptive learning

    CN119418725A

  • Artificial intelligence-based system for analyzing student behavior

    DE202024107622U1

Cited By

  • Virtual reality-based depression adjuvant therapy system

    CN121528447A