Campus psychological assessment multi-modal emotion recognition and privacy protection method and system
By collecting multimodal data non-contactly and performing X-Fi fusion and TinyMultiNet recognition, the system addresses the issues of insufficient multimodal fusion and privacy protection in campus psychological assessment systems. This enables efficient emotion recognition and personalized intervention on terminal devices, improving the accuracy and privacy protection of mental health assessments.
Patent Information
- Application Number
- CN202510840663.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-11-21
AI Technical Summary
Existing campus mental health assessment systems suffer from problems such as insufficient multimodal integration, weak privacy protection, high deployment difficulty, and rigid intervention mechanisms, making it difficult to achieve multi-dimensional psychological state description, privacy protection, and low-resource-consumption terminal deployment.
Physiological signals, voice signals, and facial expression data are collected in a non-contact manner. Multimodal fusion is performed using the X-Fi model, and emotion recognition is achieved using the TinyMultiNet lightweight model. On-device encryption and data blurring are also performed to generate personalized visualization reports and emotion guidance.
It improves the accuracy of emotion recognition, ensures data privacy, reduces resource consumption, enables real-time deployment and personalized psychological intervention on terminal devices, and provides the ability to identify and intervene early.
Smart Images

Figure CN120998385A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of campus psychological assessment, and in particular to a campus psychological assessment multi-modal emotion recognition and privacy protection method and system. BACKGROUND
[0002] With the increasing attention to mental health problems in primary and secondary schools and colleges, traditional psychological assessment methods (such as questionnaires or interviews) face the problems of relying on subjective filling, difficulty in real-time monitoring, and low participation rate. At the same time, artificial intelligence sensing technology (such as physiological signal data recognition, speech emotion recognition, and face analysis) has been gradually applied to psychological state recognition, but the existing system still has the following shortcomings:
[0003] 1) The limitations of single-modal information in the comprehensive description of mental state: Many existing mental health assessment systems rely on natural language processing (NLP) techniques to analyze text or use artificial intelligence tools to make judgments based on single information input such as facial expressions or voice recognition. For example, chat records are analyzed to screen for depression, or social media and wearable device data are analyzed to assess mental state. However, mental state is a dynamic attribute of humans that is easily influenced by small external factors, so it should be considered from multiple dimensions to achieve a more comprehensive description of mental state. 2) Neglect of privacy protection, especially the ethical and legal risks of facial image data in educational settings: Although emotion recognition technology based on facial expression analysis has shown great potential in mental assessment systems, the sensitivity of facial image data as biological information should be given special attention. For vulnerable groups such as adolescents, misuse of facial data can have serious consequences. Existing systems usually store data centrally, which faces the risk of hacking and internal personnel leaks. 3) The complexity of the system model and the problem of resource consumption limit its deployment in terminal scenarios (such as classrooms and counseling rooms): Mental assessment systems that incorporate artificial intelligence models usually require a large amount of computing resources (GPU / NPU) and memory support, and need to be deployed across platforms. The differences in instruction sets, compilers, and programming interfaces of different chips make model transplantation and optimization costly. Existing systems mostly use cloud deployment, relying on network transmission and centralized computing, which cannot adapt to terminal scenarios without network or low bandwidth. In the cloud scenario, the memory of GPU / NPU and other acceleration resources is also extremely limited. 4) Lack of automated feedback and intervention mechanisms, making it difficult to achieve a "measurement-evaluation-intervention" closed loop: Most current mental assessment systems mainly present results in static visual reports, such as standardized visual reports, which rely heavily on textual descriptions and lack dynamic interpretation of mental state and interaction with the evaluated person. The system cannot match the evaluation results to targeted and personalized intervention therapies, and the suggestions made by users are too general and single, such as "exercise more" and "maintain a positive attitude," making it difficult to truly play a psychological counseling role.
[0004] For example, the invention application with document number CN119153039B discloses a psychological intervention method, system, device, and medium based on multi-modal emotion recognition, which can improve the accuracy of emotion recognition and effectively improve the accuracy of psychological intervention. However, the method also has the following problems: the multi-modal information fusion mechanism is not sufficient and still mainly relies on splicing processing; the privacy protection mechanism is missing, and the exposure risk of face and voice data is high; the model deployment complexity is high, making it difficult to adapt to the campus terminal environment; and the intervention mechanism lacks individualization and closed-loop feedback capabilities.
[0005] Therefore, the current campus mental health assessment system based on artificial intelligence has broken through the limitations of traditional methods to a certain extent, but still faces problems such as insufficient multi-modal fusion, weak privacy protection, high deployment difficulty, and rigid intervention mechanism, and a campus mental assessment multi-modal emotion recognition and privacy protection method is needed to solve the above problems. SUMMARY
[0006] In view of the above problems, the purpose of the present application is to provide a campus mental assessment multi-modal emotion recognition and privacy protection method and system, which fuses multi-modal data and improves the accuracy of emotion recognition.
[0007] The embodiment of the present application provides a campus mental assessment multi-modal emotion recognition and privacy protection method and system.
[0008] The first aspect is a campus mental assessment multi-modal emotion recognition and privacy protection method, comprising:
[0009] S1, collecting physiological signal data, speech signal data and facial expression data in a non-contact manner, and preprocessing the collected data;
[0010] S2, aligning and processing the preprocessed signal data to extract a signal data feature vector;
[0011] S3, inputting the signal data feature vector into a modal invariant base model for multi-modal fusion;
[0012] S4, inputting the fused data into a lightweight multi-modal emotion recognition model for emotion recognition;
[0013] S5, generating a user emotion state and emotion intensity visualization report according to the emotion recognition result;
[0014] S6, calculating a DASS-21 index according to the visualization report, generating a standard assessment scale, and providing emotion relief.
[0015] Optionally, S1 includes collecting physiological signal data by a millimeter wave radar, collecting speech signal data by a microphone array, and collecting facial expression data by a video, wherein the physiological signal data includes respiratory signal data and heartbeat signal data;
[0016] The physiological signal data is subjected to low-pass filtering, the speech signal data is subjected to band-pass filtering, and the facial expression data is subjected to privacy protection.
[0017] Optionally, the method for preprocessing the collected data in S1 comprises:
[0018] The respiratory signal data and body motion interference are separated by empirical mode decomposition (EMD), and the respiratory signal data features are extracted;
[0019] Through micro-Doppler analysis, the heartbeat signal data features are extracted by using the slight frequency offset of chest vibration;
[0020] Through the deep learning model RNNoise, the speech signal data features are extracted by denoising the speech signal data;
[0021] The expression signal data features are extracted by using the deep learning model 3D-CNN.
[0022] Optionally, the S2 extracts the signal data feature vector, and the step includes:
[0023] S21, aligning the timestamps of the preprocessed modal data;
[0024] S22, performing time domain and frequency domain analysis on the preprocessed physiological signal data to obtain heart rate variability, respiratory rate and chest vibration amplitude features;
[0025] S23, processing the preprocessed speech signal data by using the librosa model to extract pitch, speech rate, energy and MFCC features;
[0026] S24, processing the expression signal data extracted by 3D-CNN by using the neural network model MicroExpNet to obtain micro-expression features.
[0027] Optionally, the modal invariant base model in S3 is an X-Fi model, wherein the X-Fi model includes a data input layer, a cross-modal interaction layer and an output layer.
[0028] The X-Fi model extracts feature information of each modal data through the modal specific encoder of the data input layer, calculates the similarity of the features between the multi-modal through the cross-modal interaction layer and generates weights, performs multi-modal fusion according to the weights, and outputs by the output layer.
[0029] Optionally, the emotional state and emotional intensity visualization report in S5 includes a user emotional state portrait, a radar chart showing emotional core indicators and a line chart showing user emotional stability and fluctuation trend.
[0030] Optionally, the standard evaluation scale includes three subscales of depression, anxiety and stress, each subscale has seven questions, and the score of each question adopts four levels: never, sometimes, often and always, and the score is mapped to four severity levels: normal, mild, moderate and severe. If it enters the mild range, the system will issue a warning and trigger the psychological counseling mechanism.
[0031] Second aspect: a campus psychological evaluation multi-modal emotion recognition and privacy protection system, the system comprises:
[0032] The data acquisition preprocessing module is used for collecting physiological signal data, voice signal data and facial expression signal data collected in a non-contact manner, and pre-processing the collected signal data.
[0033] The feature extraction module is used for aligning and processing the pre-processed signal data to extract a signal data feature vector.
[0034] The multi-modal fusion model module is used for inputting the signal data feature vector into a modal invariant basic model for multi-modal fusion.
[0035] The lightweight emotion recognition model module is used for inputting the fused data into a lightweight multi-modal emotion recognition model for emotion recognition.
[0036] The emotion state portrait construction module is used for generating a user emotion state and emotion intensity visual report according to the emotion recognition result.
[0037] The evaluation and feedback module is used for calculating a DASS-21 index according to the visual report, generating a standard evaluation scale, and providing emotion relief.
[0038] The privacy protection sub-module performs blurring processing and identity feature desensitization on the video data, and performs end-side encryption on the collected data.
[0039] The third aspect: an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to realize the steps of the method provided in the first aspect.
[0040] The fourth aspect: a non-transitory computer readable storage medium, having a computer program stored thereon, wherein the computer program is executed by a processor to realize the steps of the method provided in the first aspect.
[0041] The beneficial effects of the present application are as follows:
[0042] 1、The present application fuses physiological signals, voice signals and facial expression data, and the system can capture emotion features from multiple angles to improve the emotion recognition accuracy. This multi-modal fusion method can more comprehensively depict the psychological state and assist in more accurately locating the root cause of psychological problems. Meanwhile, the system extracts features such as heart rate variability (HRV) and respiratory rate, extracts features such as pitch, speech rate, energy and MFCC, and fuses facial micro-expression data. Through the feature-level extraction and fusion mechanism, the modal data is independently processed and then interacted, which improves the expression ability and robustness of the system.
[0043] 2、The present application adopts de-identification processing of personal information, blurs the original image during face data collection, desensitizes the identity features on the basis of retaining the action unit, encrypts the collected data on the terminal side, and ensures the confidentiality of the data during transmission. At the same time, the physiological signals are collected by the millimeter wave radar in a non-contact manner, which reduces the resistance of students due to the fear of privacy leakage, and can monitor the psychological state of students without their awareness.
[0044] 3、The TinyMultiNet lightweight deep learning model of the present application is used as an emotion recognition model, and through model compression technology and hardware adaptation optimization, the system can be deployed in terminal devices (such as classrooms and psychological counseling rooms) in a lightweight manner, without relying on cloud computing resources. The system can run in real time on low-power teaching terminal devices, adapt to terminal scenarios without network and low bandwidth, and reduce deployment costs and resource consumption.
[0045] 4、The present application can generate visual reports such as user emotional state portrait, emotional core index radar chart, and emotional stability change trend line chart in real time, so that users can have a clearer understanding of their own psychological state. At the same time, combined with the generated emotional indicators and changes, a standard evaluation scale is generated, and an emotional guidance option is provided. For users detected to have psychological problems, the system can guide users to adjust their emotions in real time through the camera, realizing early identification and early intervention.
[0046] 5、The system of the present application adopts a non-contact data collection method, which does not require students to actively cooperate, thereby reducing the resistance of students and lowering the participation threshold. At the same time, the system can generate personalized suggestions according to multiple evaluation results, dynamically monitor the emotional changes of students, and provide more comprehensive mental health information for schools and parents, helping them better understand the psychological state of adolescents. BRIEF DESCRIPTION OF DRAWINGS
[0047] Figure 1 It is a flowchart of the emotion recognition and privacy protection method of the present application;
[0048] Figure 2 It is a schematic diagram of the principle structure of the emotion recognition and privacy protection method of the present application;
[0049] Figure 3 It is a schematic diagram of the structure of the emotion recognition and privacy protection system of the present application;
[0050] Figure 4 It is a schematic diagram of the lightweight neural network TinyMultiNet structure of the present application;
[0051] Figure 5 It is a schematic diagram of the structure of the electronic device of the present application. DETAILED DESCRIPTION
[0052] Embodiments of the present application are described below in detail, examples of which are shown in the drawings, wherein the same or similar symbols represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary, only for the purpose of explaining the present application, and cannot be understood as a limitation of the present application.
[0053] Current campus psychological assessment methods have many limitations: first, the psychological assessment system relies on NLP or single information input, such as facial expression or speech recognition, but the psychological state is complex and variable, and needs to be considered from multiple dimensions. Second, facial image data is sensitive, especially in the education scene, privacy protection and ethical legal risks need to be paid attention to. Third, the psychological assessment system model is complex, and the resource consumption is large, which limits the deployment in terminal scenes such as classrooms. Fourth, there is a lack of automatic feedback and intervention mechanism, and it is difficult to realize the complete "measurement-evaluation-intervention" process.
[0054] Embodiment one:
[0055] In view of the above problems, the present application provides a campus psychological assessment multi-modal emotion recognition and privacy protection method, Figure 1 The flowchart of the method of the present application, the method comprises:
[0056] S1, collecting physiological signal data, speech signal data and facial expression signal data in a non-contact manner, and pre-processing the collected signal data.
[0057] Non-contact acquisition includes using millimeter wave radar acquisition, microphone array acquisition and video acquisition, wherein:
[0058] The millimeter wave radar realizes non-contact, non-interference and low-load detection and collection of human physiological signal data such as respiration and heartbeat. Respiration and heartbeat will cause periodic displacement of the chest cavity. By radar emission signal data, receiving human echo and corresponding series of signal data processing methods, respiration and heartbeat signal data are extracted from the original signal data. Millimeter wave radar can capture these micron-level movements through Doppler effect or phase change.
[0059] At the same time, video acquisition is used to collect object facial expression pictures. Based on 3D-CNN, it is recognized that changes in facial expression pictures (such as smiling and frowning) will cause slight changes in skin and muscle.
[0060] For the collection of speech signal data, a microphone array composed of multiple microphones is used to capture far-field speech, and at the same time, beamforming technology is used to suppress environmental noise, so as to realize the collection of speech signal data.
[0061] The collected physiological signal data is subjected to low-pass filtering processing, the speech signal data is subjected to band-pass filtering processing, the facial expression data is subjected to de-identification processing of personal information, the original image is subjected to fuzzy security processing during facial data collection, the identity features are desensitized on the basis of retaining the action unit, and the collected data is subjected to end-side encryption, so as to ensure the confidentiality of the data in the transmission process and realize the protection of personal privacy.
[0062] The physiological signal data mainly includes respiratory signal data and heartbeat signal data, wherein the respiratory signal data, the heartbeat signal data, the expression signal data and the speech signal data jointly constitute multi-modal signal data.
[0063] When the respiratory signal data is extracted, the respiratory signal data is separated from body motion interference by empirical mode decomposition (EMD);
[0064] When the heartbeat signal data is extracted, the heartbeat signal data is extracted by micro-Doppler analysis using the micro frequency offset of chest vibration;
[0065] When the speech signal data is extracted, the speech signal data is captured by a microphone array, and the speech signal data is subjected to noise reduction processing by a deep learning model RNNoise to extract the speech signal data.
[0066] When the facial expression data is extracted, a deep learning model 3D-CNN is used to extract dynamic expression features, which can recognize more complex expressions, and the facial expression data is collected.
[0067] The respiratory signal data is stored in the form of a floating-point array by separating the frequency difference between the respiratory signal data and the body motion interference. Micro-Doppler analysis reveals the high-frequency micro-Doppler component corresponding to the heartbeat, and these data are stored as a multi-dimensional numerical array. The audio processed by RNNoise retains its original encoding format, usually PCM encoding. In addition, the time dimension is also considered when processing the facial expression data, and the data after 3D-CNN processing presents as a multi-dimensional floating-point tensor.
[0068] S2, aligning the pre-processed signal data to extract signal data feature vectors.
[0069] The pre-processed multi-modal data are aligned by time stamp.
[0070] The pre-processed physiological signal data are subjected to time domain and frequency domain analysis to obtain heart rate variability (HRV), respiratory rate and chest vibration amplitude characteristics.
[0071] The collected speech signal data is processed by librosa to extract pitch, speech rate, energy and MFCC features; wherein the pitch, speech rate and energy can reflect the degree of emotion, and the MFCC represents the timbre and can reflect the emotion category.
[0072] The expression signal data extracted by the 3D-CNN is processed by using a neural network model MicroExpNet to obtain micro-expression features.
[0073] S3, inputting the signal data feature vector into a modal invariant foundation model for multi-modal fusion.
[0074] The modal invariant foundation model adopts an X-Fi model, the X-Fi model can realize cross-modal independent feature alignment through a Transformer architecture, adopts an adaptive weight mechanism to avoid the problem of modal loss, and simultaneously performs multi-modal fusion according to the weight to output a fused feature vector.
[0075] When the X-Fi deep learning model is used for fusion, the multi-modal data fine-grained correlation modeling is highlighted, and the accuracy of emotion recognition and psychological evaluation and the system universality are improved.
[0076] The X-Fi (Modality-Invariant Foundation Model) model comprises a data input layer, a cross-modal interaction layer and a data output layer.
[0077] The X-Fi is a feature interaction and fusion model for multi-modal data, the data input layer separately extracts feature information of each modal data based on a modality-specific encoder (Modality-Specific Encoders), the cross-modal interaction layer (Cross-modal Interaction Layer) calculates the similarity of features between modalities and generates a weight, and multi-modal information is fused according to the weight.
[0078] In the time dimension, the X-Fi applies a DTW algorithm to align the time sequence difference of asynchronous time sequence data; in the semantic dimension, the feature of each modal data is projected into a unified semantic space to share a latent space mapping. The signal data of the data input layer is low-level feature encoded, and the output layer outputs high-level semantic encoding of psychological feature mapping. Based on this, the X-Fi realizes cross-modal fine-grained correlation modeling, and the multi-modal psychological data is connected from low-level signal data to high-level semantics, providing an extensible framework for accurate psychological evaluation.
[0079] The X-Fi cross-modal interaction layer independently calculates the Query-Key-Value interaction of the multi-modal feature vector, merges the results to obtain cross-modal features based on fine-grained correlation, generates a gating weight according to the importance of each modality, the feature quality also affects the weight, and the features are fused by weighting. The data processed by the cross-modal interaction layer has modal complementarity and emotion discriminability, and is data of fused semantic features with high-level semantics, which is output by the output layer.
[0080] S4, input the fused data into a lightweight multi-modal emotion recognition model to perform emotion recognition.
[0081] TinyMultiNet is a lightweight neural network that performs lightweight processing on input data. In combination with the lightweight TinyMultiNet, real-time operation in edge devices is supported. TinyMultiNet uses quantization-aware training (QAT) in model compression technology to simulate quantization errors during the training process and to ensure data accuracy as much as possible.
[0082] TinyMultiNet is used to perform emotion recognition on the fused semantic feature data output by the output layer of X-Fi, as shown in Figure 4 TinyMultiNet is composed of fused semantic features, grouped attention fusion, shared feature layers, and emotion classification heads and intensity regression heads.
[0083] TinyMultiNet measures the difference between the predicted results and the true labels through a cross-entropy loss function, and adjusts the model parameters based on this difference, so that the features of different modalities can be better aligned during the fusion process to improve the classification accuracy. The training core is the label smoothing cross-entropy technique, which softens the labels by introducing a smoothing parameter, solves overfitting, and enhances the generalization ability. The present application uses Smooth L1 as the loss function in the regression task for loss optimization, balances stability and robustness, approximates the true value, and uses Softmax as the activation function of the last layer to convert the linear output of the model into class probability through exponential operation and normalization.
[0084] TinyMultiNet model is trained based on a dataset to improve the accuracy of emotion classification. The emotion categories are divided into seven categories: happiness, sadness, anger, fear, surprise, disgust, and neutral, and the emotion intensity is normalized to [0-1].
[0085] S5, according to the emotion recognition result, a user emotion state and emotion intensity visualization report is generated.
[0086] Based on the model output emotion categories, the user's emotional and psychological state is summarized, and the system automatically generates an individual emotion state portrait. The visualization report is in a visual form, which is convenient for users to understand. The emotion user emotion portrait is constructed, and the core indicators are displayed in a radar chart, and the stability and fluctuation trend is displayed in a line chart.
[0087] S6, according to the visualization report, the DASS-21 index is calculated, a standard evaluation scale is generated, and emotional counseling is provided.
[0088] In combination with the generated emotion indicators and change visualization report, the system automatically calculates the DASS-21 index, generates a standard evaluation scale, and provides emotional counseling.
[0089] The DASS-21 index is a standardized scale for evaluating three mental states of individuals, depression, anxiety and stress, comprising 21 questions, divided into three subscales, each with 7 questions, and each question is scored in four levels: never, sometimes, often and always, and the scores are mapped to four severity levels: normal, mild, moderate and severe, and if a dimension enters the mild range or above, the system will issue a warning and trigger a psychological counseling mechanism.
[0090] For users detected to have psychological problems, the system will provide emotional counseling options for users to choose from, while real-time guidance can be provided through the camera to help users adjust their emotions. When an anomaly is detected, a composite risk index is constructed based on multi-modal data, and a warning is issued when a certain threshold is reached; for a user, the results of multiple evaluations are combined to generate personalized recommendations, enabling early identification and intervention.
[0091] Visual reporting can help users clearly understand the results of psychological assessments and better understand their own unconscious emotional changes. For teenagers themselves, they may not be aware of underlying psychological problems, and the method can guide students to pay attention to their own emotions and psychological state; for schools, it can help schools preliminarily screen out students who may have emotional distress, excessive stress or psychological abnormalities, and prevent problems from worsening due to lack of awareness; the evaluation results can provide data reference for psychological teachers to develop individualized counseling plans based on the psychological characteristics of different students and provide timely and effective learning strategy guidance for students under heavy academic pressure; the evaluation results can also serve as a starting point for school-parent communication, helping parents understand their children's psychological state in school, such as whether they are adapting to group life and how their peer relationships are, and avoiding conflicts between parents and children due to neglect or misunderstanding.
[0092] Embodiment two:
[0093] The present application also provides a campus psychological assessment multi-modal emotion recognition and privacy protection system, as shown in Figure 3 The system comprises:
[0094] The data acquisition and preprocessing module is used to collect physiological signal data, voice signal data and facial expression signal data collected in a non-contact manner, and to preprocess the collected signal data.
[0095] The millimeter wave radar reflection communication sub-module is used to non-contact capture the user's breathing rate, heart rate variation and body micro-movement; the voice acquisition sub-module is used to acquire voice data using an array microphone; and the video acquisition sub-module is used to acquire face video data, supporting high frame rate facial motion analysis.
[0096] The privacy protection module performs blurring processing and identity feature desensitization on the video data, and performs end-side encryption on the collected data.
[0097] For the psychological assessment system, it is important to ensure the confidentiality and integrity of sensitive mental health data. Through the privacy protection module, during the data collection stage, the personal information is processed using de-identification during user registration, and the original image is blurred during face data collection; during the data transmission stage, the data is encrypted for transmission; during data storage, a hierarchical data storage strategy is adopted, with data retention periods of one week, one year, and long-term; the data subject has the right to request deletion of data, and when the data storage period expires, the data destruction mechanism is automatically triggered, meeting the ethical and data compliance requirements in the education scenario.
[0098] The feature extraction module is configured to align and extract signal data feature vectors from the preprocessed signal data.
[0099] The radar signal data is extracted for respiratory rate, heart rate variability (HRV), and chest vibration amplitude, the speech signal data is extracted for pitch, speech rate, energy, and MFCC, and the video image is extracted for micro-expression action units (AUs) and facial key points. All feature vectors are normalized and input into the multi-modal fusion model.
[0100] The multi-modal fusion model module is configured to deploy a multi-modal fusion model for inputting signal data feature vectors into a modal-invariant base model for multi-modal fusion.
[0101] The multi-modal fusion model can use the X-Fi model, which implements cross-modal feature alignment through a Transformer architecture, and an adaptive weight mechanism to avoid modal missing problems and output fused general semantic embedding representations.
[0102] The lightweight emotion recognition model module is configured to input the fused data into a lightweight multi-modal emotion recognition model for emotion recognition.
[0103] The TinyMultiNet model is used to perform emotion recognition on the fused semantic features, outputting discrete emotion categories and continuous intensity values (0-1).
[0104] The emotion state portrait construction module is configured to generate a user emotion state and emotion intensity visualization report based on the emotion recognition results, and to convert the output into a user emotion state portrait, an emotion index radar chart, an emotion stability level, and an emotion fluctuation trend by using a rule-based reasoning and classification mapping combined structure.
[0105] The assessment and feedback module is configured to calculate the DASS-21 index based on the visualization report and generate a standard assessment scale, and the system will automatically select emotion relief resources (meditation videos, mindfulness training), online consultation portals, and risk warnings to provide emotion relief.
[0106] The application also provides an electronic device, Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the application is shown in the figure. Figure 5 As shown in the figure, the electronic device can include a processor, a communications interface, a memory and a communications bus, wherein the processor, the communications interface and the memory complete communications with each other through the communications bus. The processor can invoke logical instructions in the memory, for example to execute the following method:
[0107] S1, physiological signal data, speech signal data and facial expression data are collected in a non-contact manner, and the collected data is preprocessed;
[0108] S2, the preprocessed signal data is aligned to extract a signal data feature vector;
[0109] S3, the signal data feature vector is input into a modal invariant base model for multi-modal fusion;
[0110] S4, the fused data is input into a lightweight multi-modal emotion recognition model for emotion recognition;
[0111] S5, a user emotion state and emotion intensity visualization report is generated according to the emotion recognition result;
[0112] S6, a DASS-21 index is calculated according to the visualization report, a standard evaluation scale is generated, and emotional counseling is provided.
[0113] In addition, the logical instructions in the memory described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the application or the part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and various program code storage media.
[0114] The embodiment of the application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method provided by each of the above embodiments, for example including:
[0115] S1, physiological signal data, speech signal data and facial expression data are collected in a non-contact manner, and the collected data is preprocessed;
[0116] S2, the preprocessed signal data is aligned to extract a signal data feature vector;
[0117] S3, the signal data feature vector is input into a modal invariant base model for multi-modal fusion;
[0118] S4, the fused data is input into a lightweight multi-modal emotion recognition model for emotion recognition;
[0119] S5, according to the emotion recognition result, a user emotion state and emotion intensity visualization report is generated;
[0120] S6, according to the visualization report, a DASS-21 index is calculated, a standard evaluation scale is generated, and emotional counseling is provided.
[0121] The system embodiments described above are only illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0122] From the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary general hardware platforms, and of course, they can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in each embodiment or some parts of the embodiment.
[0123] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A campus psychological assessment multi-modal emotion recognition and privacy protection method, characterized in that, The method comprises the following steps: S1, collecting physiological signal data, speech signal data and facial expression data in a non-contact manner, and preprocessing the collected data; S2, aligning and processing the preprocessed signal data to extract signal data feature vectors; S3, inputting the signal data feature vectors into a modal invariant basic model for multi-modal fusion; S4, inputting the fused data into a lightweight multi-modal emotion recognition model for emotion recognition; S5, generating a user emotion state and emotion intensity visualization report according to the emotion recognition result; S6, calculating the DASS-21 index according to the visualization report, generating a standard evaluation scale, and providing emotional counseling.
2. The method of claim 1, wherein, The S1 comprises: Collecting physiological signal data by millimeter wave radar, collecting speech signal data by microphone array, and collecting facial expression data by video, wherein the physiological signal data includes respiratory signal data and heartbeat signal data; performing low-pass filtering on the physiological signal data, band-pass filtering on the speech signal data, and privacy processing on the facial expression data.
3. The method of claim 2, wherein, The method for preprocessing the collected data in S1 comprises: Separating the respiratory signal data from body motion interference by empirical mode decomposition (EMD) to extract the respiratory signal data features; Extracting heartbeat signal data features by micro-Doppler analysis using the micro frequency offset of chest vibration; Performing noise reduction processing on the speech signal data by a deep learning model RNNoise to extract speech signal data features; Extracting expression signal data features by a deep learning model 3D-CNN.
4. The method of claim 3, wherein, The S2 extracts signal data feature vectors, and the steps comprise: S21, aligning the timestamps of the preprocessed modal data; S22, performing time domain and frequency domain analysis on the preprocessed physiological signal data to obtain heart rate variability, respiratory rate and chest vibration amplitude features; S23, processing the preprocessed speech signal data by a librosa model to extract pitch, speech rate, energy and MFCC features; S24, processing the expression signal data extracted by 3D-CNN by a neural network model MicroExpNet to obtain micro-expression features.
5. The method of claim 1, wherein, The modal invariant basic model in S3 is an X-Fi model, wherein the X-Fi model comprises a data input layer, a cross-modal interaction layer and an output layer; The X-Fi model extracts feature information of each modal data by a modal-specific encoder of the data input layer, calculates the similarity of features between multi-modal and generates weights through the cross-modal interaction layer, performs multi-modal fusion according to the weights, and outputs by the output layer.
6. The method of claim 1, wherein, The emotion state and emotion intensity visualization report in S5 comprises: a user emotion state portrait, a radar chart showing emotion core indicators, and a line chart showing user emotion stability and fluctuation trend.
7. The method of claim 6, wherein, The standard evaluation scale comprises: Three subscales of depression, anxiety and stress, each subscale has 7 questions, and each question is scored by four levels: never, sometimes, often and always, and the score is mapped to four severity levels: normal, mild, moderate and severe. If it enters the mild range, the system will issue a warning and trigger the psychological counseling mechanism.
8. A campus psychological assessment multi-modal emotion recognition and privacy protection system applied to the method of any one of claims 1 to 7, characterized in that, The system comprises: The data acquisition preprocessing module is configured to collect physiological signal data, voice signal data and facial expression signal data collected in a non-contact manner, and to preprocess the collected signal data. The feature extraction module is configured to perform alignment processing on the preprocessed signal data to extract a signal data feature vector. The multi-modal fusion model module is configured to input the signal data feature vector into a modal-invariant base model for multi-modal fusion. The lightweight emotion recognition model module is configured to input the fused data into a lightweight multi-modal emotion recognition model for emotion recognition. The emotion state portrait construction module is configured to generate a user emotion state and emotion intensity visual report according to the emotion recognition result. The evaluation and feedback module is configured to calculate a DASS-21 index according to the visual report, generate a standard evaluation scale, and provide emotional counseling. The privacy protection submodule is configured to perform blurring processing and identity feature desensitization on the video data, and to perform end-side encryption on the collected data.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps of the method of any one of claims 1-7. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1-7.
Citation Information
Patent Citations
Psychological intervention methods, systems, equipment and media based on multimodal emotion recognition
CN119153039B
Cited By
Psychological state assessment parameter determination method and device, electronic equipment and storage medium
CN121587725A
Psychological state evaluation parameter determination method and device, electronic equipment and storage medium
CN121587725B
Psychological state assessment method and system based on multi-source heterogeneous perception
CN121647677A