Methods, devices, equipment, media, and programs for identifying student mental states

By acquiring students' multi-dimensional behavioral feature sequences and analyzing them using a multi-branch attention fusion neural network, the problem of difficulty in identifying students' psychological states in real time and comprehensively in existing technologies has been solved, realizing dynamic identification and early warning of psychological states.

CN122074984APending Publication Date: 2026-05-26CHINA MOBILE GROUP DESIGN INST +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE GROUP DESIGN INST
Filing Date
2026-01-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies are insufficient for real-time, comprehensive, and dynamic identification of students' psychological states, and lack effective early detection and warning capabilities, especially for non-acute, progressive psychological problems.

Method used

By acquiring multi-dimensional behavioral feature sequences of target students, including facial expression feature sequences, behavioral trajectory feature sequences, and voice emotion feature sequences, a multi-branch attention fusion neural network is used for feature encoding and fusion analysis to output psychological state representation information and identify the students' psychological state.

Benefits of technology

It enables dynamic identification of students' psychological states, improves the ability to capture subtle fluctuations and gradual changes, enhances the ability to detect and warn of non-acute, gradual psychological problems in the early stages, and reduces response lag.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122074984A_ABST
    Figure CN122074984A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, medium, and program product for identifying student psychological states, aiming to address the problems in existing technologies such as the difficulty in real-time, comprehensive, and dynamic identification of students' psychological states, and the lack of effective early detection and warning capabilities. The method includes: acquiring a multi-dimensional behavioral feature sequence of the target student, which includes at least an expression feature sequence, a behavioral trajectory feature sequence, and a voice emotion feature sequence; inputting the multi-dimensional behavioral feature sequence into a trained psychological state identification model to output psychological state representation information for characterizing the target student's psychological state; and identifying the target student's psychological state based on the psychological state representation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, medium and program product for recognizing students' psychological states. Background Technology

[0002] Currently, the campus student mental health management system mainly relies on traditional methods such as periodic psychological scale assessments, teacher subjective observation, and home visits to identify students' mental states. However, these traditional methods have three inherent flaws: First, the identification results of students' mental states are highly dependent on teachers' personal experience and students' self-reports at specific moments, and are easily influenced by the evaluator's personal experience and subjective factors such as students' concealment behavior; second, traditional methods usually identify students' mental states on a weekly, monthly, or even semester basis, which, due to the low sampling frequency, fails to capture subtle fluctuations and gradual changes in students' mental states; third, traditional methods often only intervene passively after students' mental problems appear or worsen, resulting in a severely delayed response.

[0003] In conclusion, traditional methods are insufficient for real-time, comprehensive, and dynamic understanding of students' psychological state, especially for some non-acute, progressive psychological problems, where there is a lack of effective early detection and warning capabilities. Summary of the Invention

[0004] This application provides a method for identifying students' psychological states, which addresses the problems in the prior art of difficulty in real-time, comprehensive, and dynamic identification of students' psychological states, as well as the lack of effective early detection and warning capabilities.

[0005] This application also provides a student psychological state recognition device, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] The embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide a method for identifying student psychological states, including: Obtain the multi-dimensional behavioral feature sequence of the target student. The multi-dimensional behavioral feature sequence includes at least the facial expression feature sequence, the behavioral trajectory feature sequence, and the voice emotion feature sequence. The multi-dimensional behavioral feature sequence is input into the trained psychological state recognition model to output psychological state representation information used to characterize the psychological state of the target student. Identify the psychological state of target students based on psychological state representation information.

[0007] Optionally, obtain a multi-dimensional behavioral feature sequence of the target student, including: By deploying sensing devices in public areas of the campus, multimodal behavioral feature data containing target students is collected. The multimodal behavioral feature data includes at least video data and audio data. Feature extraction is performed based on multimodal behavioral feature data to obtain multidimensional behavioral feature sequences of target students.

[0008] Optionally, the mental state recognition model is a multi-branch attention fusion neural network, including at least: The input layer includes a first input terminal, a second input terminal, and a third input terminal, which are used to receive facial expression feature sequences, behavioral trajectory feature sequences, and speech emotion feature sequences, respectively. The feature encoding layer includes a first encoding sub-network, a second encoding sub-network, and a third encoding sub-network, which are respectively connected to the first input terminal, the second input terminal, and the third input terminal. It is used to perform sequence encoding on the facial expression feature sequence, the behavioral trajectory feature sequence, and the speech emotion feature sequence to obtain the first sequence representation, the second sequence representation, and the third sequence representation. The attention fusion layer is used to adaptively weight and fuse the first sequence representation, the second sequence representation, and the third sequence representation based on self-attention and / or cross-attention mechanisms to obtain a fused representation; The output layer is used to output psychological state representation information based on the fusion representation to characterize the psychological state of the target student.

[0009] Optionally, a multi-dimensional behavioral feature sequence is input into the trained mental state recognition model to output mental state representation information for characterizing the mental state of the target student, including: The facial expression feature sequence is input into the first encoding sub-network to obtain the first sequence representation; The behavioral trajectory feature sequence is input into the second encoding sub-network to obtain the second sequence representation; The speech emotion feature sequence is input into the third coding sub-network to obtain the third sequence representation; The first sequence representation, the second sequence representation, and the third sequence representation are input into the attention fusion layer to obtain the fused representation; Based on the fusion representation, the output layer outputs mental state representation information.

[0010] Optionally, the psychological state of the target student can be identified based on psychological state representation information, including: The target students' emotional stability index, social participation index, loneliness trend index, and behavioral abnormality index were calculated based on their psychological state representation information. The psychological state of target students is identified based on the emotional stability index, social participation index, loneliness trend index, and behavioral abnormality index.

[0011] Optional, an emotional stability index is used to characterize the dispersion of the target student's facial and vocal emotional fluctuations within a preset time window; The Social Engagement Index is used to characterize the frequency and duration of target students' spatial proximity or interaction with others in public areas of the campus. The Loneliness Trend Index is used to characterize the rate of decline or persistently low value of the social engagement index over several consecutive days. The behavioral abnormality index is used to characterize the degree to which a target student's behavioral trajectory or activity rhythm deviates from their personal historical baseline or the statistical distribution of their peer group.

[0012] Secondly, embodiments of this application provide a student psychological state recognition device, including an acquisition module, a processing module, and a recognition module, wherein: The acquisition module is used to acquire the multi-dimensional behavioral feature sequence of the target student. The multi-dimensional behavioral feature sequence includes at least the facial expression feature sequence, the behavioral trajectory feature sequence, and the voice emotion feature sequence. The processing module is used to input multi-dimensional behavioral feature sequences into the trained psychological state recognition model, so as to output psychological state representation information for representing the psychological state of the target student. The identification module is used to identify the psychological state of the target student based on psychological state representation information.

[0013] Thirdly, embodiments of this application provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the student psychological state recognition method as described above.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the student psychological state recognition method described above.

[0015] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the student mental state recognition method as described above.

[0016] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The method provided in this application acquires multi-dimensional behavioral feature sequences, such as facial expression feature sequences, behavioral trajectory feature sequences, and vocal emotion feature sequences, of the target student. These sequences are then fused and analyzed using a trained psychological state recognition model to output psychological state representation information. Finally, the psychological state of the target student is identified based on this representation information. This comprehensive modeling of multi-dimensional behavioral feature sequences to form a quantitative representation of the student's psychological state improves the ability to capture subtle fluctuations and gradual changes in psychological state, enabling dynamic identification of the student's psychological state. Furthermore, identifying the target student's psychological state through psychological state representation information allows for the output of recognition results before psychological problems become apparent, thereby enhancing the early detection and warning capabilities for non-acute, gradual psychological problems and reducing response lag. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic diagram illustrating the implementation process of a student psychological state recognition method provided in this application embodiment; Figure 2 A schematic diagram illustrating the implementation process of identifying the psychological state of a target student based on psychological state representation information, provided in an embodiment of this application; Figure 3 This application provides a schematic diagram of the specific structure of a student psychological state recognition device. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] It should be understood that the training and prediction processes of the AI ​​models involved in the various embodiments of this specification all adhere to multiple legal and compliant principles, including legal data sources, compliant data content, compliant data governance, compliant training objectives and schemes, compliant training processes, compliant training environments and tools, and compliant ethical verification of training results, and comply with the requirements of Article 5 of the Patent Law. Among them: Data source legitimacy: All datasets used for AI model training were obtained through legal means, covering three categories: publicly authorized data, data authorized by partners, and self-collected compliant data. Publicly authorized data comes from compliant data sources following open-source licenses such as Apache 2.0, with complete copyright attribution and authorization scope clearly marked, and no unauthorized open-source code or data reuse. Data authorized by partners has been subject to formal data usage agreements, clearly defining the scope, duration, and confidentiality obligations, and possessing a complete authorization chain. For self-collected data involving personal information, strict informed consent procedures have been followed, and anonymization processes (including but not limited to field masking, feature anonymization, and differential privacy technology applications) have been implemented to remove personally identifiable information, fully complying with the requirements of relevant laws and regulations such as the "Interim Measures for the Administration of Generative Artificial Intelligence Services" and the "Personal Information Protection Law."

[0020] Data content compliance: The AI ​​model's dataset undergoes multiple screenings and cleaning processes to remove all content that may violate social morality or harm public interests. It contains no obscene, pornographic, violent, discriminatory, or information that endangers national or public safety, nor does it involve the illegal acquisition or use of genetic resources. For data in sensitive fields (such as healthcare and finance), an additional privacy-preserving computation module (including federated learning and secure multi-party computation technologies) ensures that the data is "usable but not visible," avoiding compliance risks during the original data transmission process and ensuring that the data application scenarios and uses comply with public order and good morals and industry regulatory requirements.

[0021] Data governance norms: A complete data traceability system is established during the AI ​​model training process to automatically record the source, collection time, annotation process, cleaning rules, and permission allocation of training data, generating traceable compliance reports to ensure that the data is verifiable throughout its entire lifecycle. The dataset annotation process for AI models is completed by a professional human R&D team, clearly defining the proportion of human creative contributions and avoiding reliance on AI-generated data that has not undergone substantial human modification, thus meeting the examination requirements for "human main contributions" in AI patent applications.

[0022] Training objectives and plans are compliant: The AI ​​model training aims to provide auxiliary support for mental health education, psychological counseling resource allocation, and campus care services through artificial intelligence technology. The model's output is solely for risk warnings or statistical representations and does not constitute a medical diagnosis. It will not be used to make punitive, discriminatory, or other automated decisions that significantly impact students' legitimate rights. The training plan and final output do not violate any mandatory provisions of laws or administrative regulations, do not harm public interests or the legitimate rights of others, and pose no potential risk of being used for illegal activities, privacy violations, or disruption of public safety. It strictly adheres to the ethical principle of "intelligent for good."

[0023] Training process compliance: A closed-loop training framework is adopted to ensure compliance and controllability of the training process. The specific process is as follows: First, training samples are obtained through compliant data sources. After the aforementioned data cleaning and desensitization, they are input into the neural network model to generate preliminary training results. Second, an expert system is introduced to verify the preliminary results. Based on preset rules and human expert experience, the feasibility of the results is evaluated, and outputs that may pose ethical risks or compliance hazards are corrected (such as removing decision-making logic that violates public order and good morals, and adjusting model parameters that do not comply with safety regulations). Finally, the loss function weights are dynamically optimized based on expert system feedback to strengthen the model's learning of compliant results, avoid overfitting errors or non-compliant labels, and form a closed-loop control of "data input - model training - expert verification - parameter optimization - result feedback" to ensure that the entire training process complies with A5 ethical review requirements.

[0024] Training environment and tool compliance: AI model training is implemented using nationally licensed chips and a compliant training platform. All open-source frameworks and components used in the training process have obtained their corresponding licenses, and copyright statements and patent citation information are fully retained, with no instances of infringement or reuse. The training environment is built using virtual devices (containers / virtual machines) with fixed random seeds and initial parameter configurations to ensure the reproducibility of the training process. Furthermore, through access control and operation log recording, risks such as data leakage and parameter tampering during training are prevented, ensuring the security and compliance of the training process.

[0025] Training results ethical verification compliance: After the model is trained, it undergoes additional third-party ethical compliance assessment and algorithm filing review to verify that the model output does not violate social morality or harm public interests. For potentially sensitive scenarios (such as public services and intelligent decision-making), a special result verification mechanism is established to ensure that the model always complies with Article 5 of the Patent Law and relevant laws and regulations in practical applications.

[0026] In summary, the data and training process used in the AI ​​model of this specification strictly comply with the relevant provisions of Article 5 of the Patent Law and the Patent Examination Guidelines (2023 Edition), and there are no violations of laws, social ethics, public interests, or illegal use of genetic resources. It fully meets the compliance requirements for patent authorization.

[0027] To address the challenges of real-time, comprehensive, and dynamic identification of students' psychological states in existing technologies, as well as the lack of effective early detection and warning capabilities, this application provides a method for identifying students' psychological states.

[0028] The execution subject of this method can be various types of computing devices, or it can be an application or app installed on the computing device. The computing device can be a user terminal such as a mobile phone, tablet computer, or smart wearable device, or it can be a server.

[0029] For ease of description, this application uses a server as the execution subject of the method in its embodiments to illustrate the method. Those skilled in the art will understand that this embodiment uses a server as an example to describe the method, which is merely an illustrative example and does not limit the scope of protection of the corresponding claims.

[0030] Specifically, the implementation flow of the method provided in this application embodiment is as follows: Figure 1 As shown, it includes the following steps: Step 102: Obtain the multi-dimensional behavioral feature sequence of the target student. The multi-dimensional behavioral feature sequence includes at least the facial expression feature sequence, the behavioral trajectory feature sequence, and the voice emotion feature sequence.

[0031] In some implementations, considering that students' daily activities on campus are scattered across multiple typical public areas such as classrooms, cafeterias, corridors, playgrounds, and libraries, their psychological states may be manifested through differentiated behavioral patterns in different scenarios. Therefore, to build a continuous monitoring data foundation covering students' daily activities across all scenarios, sensing devices can be deployed in multiple typical public areas when acquiring the multi-dimensional behavioral feature sequences of target students. Then, using the sensing devices deployed in the campus public areas, multi-modal behavioral feature data containing the target students is collected, wherein the multi-modal behavioral feature data includes at least video and audio data; finally, feature extraction is performed based on the multi-modal behavioral feature data to obtain the multi-dimensional behavioral feature sequences of the target students.

[0032] Optionally, the deployment locations can fully consider the differences in behavioral representations across different scenarios: deploying equipment in classrooms to capture learning focus, classroom interaction, and immediate emotional responses; observing social patterns and dining habits in the cafeteria; monitoring sustained focus and abnormal behaviors (such as prolonged sleeping) in the library / study room; analyzing sports participation and group interaction in the playground and public activity areas; and identifying behaviors such as loitering and lingering that may reflect emotional fluctuations in transitional areas such as corridors and dormitory entrances.

[0033] The sensing devices mainly include AI camera modules and voice acquisition units. The AI ​​camera module can be a network camera supporting wide dynamic range (WDR) and starlight-level night vision, equipped with a CMOS sensor of at least 1 / 2.8 inches and 2K (2560×1440) resolution to ensure clear capture of facial details and body movements under different lighting conditions. The high-sensitivity voice acquisition unit can use an array microphone with a sampling rate of at least 44.1kHz and a bit depth of at least 16bit to ensure the sound quality required for voice emotion and analysis.

[0034] Secondly, the sensing device can support both infrared and visible light modes, enabling continuous data acquisition around the clock. It also has edge computing capabilities, allowing for preliminary processing such as face detection and skeletal key point extraction at the edge, thus reducing the burden on backend transmission and computing.

[0035] To ensure data validity and reliability, in some embodiments, the following principles can be followed when deploying sensing devices: the camera installation angle and height should avoid obstruction to ensure complete capture of students' faces, upper bodies, and movement trajectories; the voice collector should be placed close to the sound source in a location with low environmental interference. All devices are connected to the backend server via low-latency, high-stability networks such as Gigabit Ethernet or 5G to support real-time data transmission and processing. Through the above deployment strategy and device selection, the system establishes a sustainable, multi-dimensional, and high-quality foundation for student behavior data collection.

[0036] Specifically, when acquiring the multi-dimensional behavioral feature sequence of the target student, the AI ​​camera module can continuously collect raw video streams containing the target student's face, posture, and movement path based on a preset frame rate. Through integrated face detection and multi-object tracking algorithms, it completes preliminary face localization, identity association (identified by an anonymous ID), and skeletal key point extraction on the device, generating a preliminary visual structured data stream containing timestamps. The synchronously deployed audio pickup array uses beamforming and noise reduction processing to directionally collect raw audio streams of the area where the target student is located, and filters out effective speech segments through voice activity detection (VAD).

[0037] Subsequently, the backend analysis server receives the pre-processed visual structured data stream and audio data stream, and performs fine-grained feature extraction: for the video data stream, a temporal convolutional network is used to extract facial expression features (such as pleasure and tension scores) and behavioral trajectory features (such as coordinates, speed, and distance to nearest neighbors) sampled at fixed time intervals (e.g., per second); for the audio data stream, speech emotion features (such as tone, speech rate, and energy profile) aligned with the same temporal primitives are extracted. Finally, the cross-modal temporal features belonging to the target student are aligned and concatenated along the time axis to construct a multi-dimensional behavioral feature sequence for the student, providing structured input for subsequent fusion analysis.

[0038] In some implementations, when acquiring the facial expression feature sequence of the target student, temporal facial expression analysis can be performed on the acquired video stream using a visual analysis unit deployed on a server or edge node. Specifically: First, a deep learning-based face detection algorithm (such as MTCNN or RetinaFace) is used to process each video frame, accurately selecting the face regions within it. Then, the detected faces undergo geometric correction and illumination normalization to align them to standard poses and eliminate the influence of illumination differences, providing standardized input for subsequent analysis.

[0039] Next, convolutional neural network models (such as ResNet and VGG-Face) pre-trained on large-scale facial expression datasets (such as FER2013 and AffectNet) are used to extract high-dimensional deep facial expression features from the normalized face images. These features are then fed into a multi-classifier (such as a Softmax classification layer) to identify the classification results corresponding to the seven basic emotions: happiness, sadness, anger, surprise, fear, disgust, and calmness, and output the intensity quantification score of each emotion (usually represented by a score from 0 to 100).

[0040] Finally, the above recognition results are sampled and organized at fixed time intervals (e.g., every second), and the emotional scores of the target students in a continuous time period are arranged in chronological order to construct an expression feature sequence that reflects the pattern of their emotional changes.

[0041] In some implementations, when acquiring the behavioral trajectory feature sequence of a target student, firstly, a multi-target tracking algorithm can be used to process continuous video frames. Through human detection, trajectory prediction, and data association, continuous tracking of the target student can be achieved. Optionally, to address cross-camera scenes and occlusion issues, Re-ID technology can be integrated for identity re-identification, thereby ensuring the consistency and continuity of the target student in the campus multi-camera network. Secondly, for the tracked target student, the coordinates of 18 to 25 key points on their body are extracted in real time using a human pose estimation model. Then, based on the tracked key point coordinates, the spatial positions of the same target student at consecutive time points can be sequentially used to generate their continuous behavioral trajectory. Next, the trajectory is analyzed in depth to extract the following multi-dimensional features: Dwelling point identification: By setting the spatial clustering radius and minimum duration threshold, identify students' long-term stay events in specific areas (such as corners or in front of windows).

[0042] Movement pattern quantification: Calculates the instantaneous velocity, average velocity, acceleration, and frequency of change in direction of movement of the trajectory to identify patterns such as abnormal loitering, slow movement, or sudden running.

[0043] Social interaction metrics: Real-time calculation of the Euclidean distance between the student and all other students in the field of vision, combined with the group density in the scene, to calculate the duration of the interaction within different social distance thresholds, with particular attention to the student's marginalized status or atypical long-distance maintenance behavior in the crowd.

[0044] Determining Alone Status: Based on social distance calculations, the system calculates the cumulative time a student spends within a specific time period without entering a preset social radius with anyone.

[0045] Finally, the various indicators obtained from the above analysis, including spatial coordinate sequence, velocity sequence, dwell time, social distance sequence, and time spent alone, are aligned, encoded, and organized according to a unified time axis to obtain the behavioral trajectory feature sequence of the target student. This sequence comprehensively represents the target student's movement, stay, and social interaction patterns in physical space in the form of structured time-series data.

[0046] In some implementations, when acquiring the speech emotion feature sequence of the target student, speech activity detection can be performed on the original audio stream first. By analyzing short-time energy, zero-crossing rate, or model-based VAD algorithms, segments containing valid speech can be accurately distinguished from silent or background noise segments, thereby filtering out clean speech data that can be analyzed.

[0047] Secondly, to protect the privacy and security of the target students and the accuracy of the psychological state recognition results, this application embodiment can, under strict authorization, use a voiceprint model based on i-vector or x-vector technologies to confirm the speaker's identity for speech segments, so as to achieve accurate association between speech and specific target students.

[0048] Subsequently, acoustic features were extracted from the selected speech segments, including the following features: Prosodic features: such as the statistical values ​​of the fundamental frequency (F0) (mean, variance, dynamic range), speech rate (number of syllables per unit time), and the frequency and duration distribution of pauses; Sound quality characteristics: such as formant frequency and bandwidth, spectral tilt; Energy characteristics: such as short-time energy and its entropy value; Spectral characteristics: such as Mel frequency cepstral coefficients (MFCC) and their first and second order differences.

[0049] These features together constitute the acoustic characteristics of a speech signal.

[0050] Next, the extracted acoustic feature sequences are input into a pre-trained deep learning model (such as LSTM or Transformer) for emotion classification. The model outputs the emotion category (such as calm, anger, sadness, happiness, anxiety) corresponding to the speech segment and its confidence or intensity score, and simultaneously analyzes the variation patterns of speech rate and intonation within the segment.

[0051] Finally, the aforementioned emotion categories, intensity scores, and derived indicators such as speech rate and tone are sampled and aligned according to fixed time windows (e.g., every 5 seconds or by speech segment) to form a speech emotion feature sequence arranged in an orderly manner over time. This sequence, as an important component of the multi-dimensional behavioral feature sequence, is input together with the visual feature sequence into the subsequent fusion analysis module for comprehensive psychological state assessment.

[0052] It should be noted that the method provided in this application strictly complies with laws and regulations such as the "Personal Information Protection Law of the People's Republic of China" and the "Law on the Protection of Minors," and a privacy protection mechanism is set up in all aspects of data collection, processing, and use.

[0053] Specifically, at the data processing level, data can be anonymized and minimized at the data acquisition or edge computing ends: real-time anonymization (such as pixelation, blurring, or cartoonization) of facial images in video streams is performed, and no recognizable original facial images are stored in non-alert states, only feature vectors such as facial expressions are extracted and saved; audio data undergoes voiceprint anonymization, removing acoustic features that can be used for direct identification, and retaining only non-identification features related to emotional state, such as speech rate, tone, pitch, and volume. Simultaneously, feature vector data and analysis results are encrypted using high-strength encryption algorithms (such as AES-256) throughout the entire transmission and storage process, and transmitted via secure protocols such as HTTPS / SSL to prevent data leakage and tampering.

[0054] At the access control level, a role-based access control (RBAC) model is adopted to implement hierarchical authorization: ordinary teachers / class teachers can only view the de-identified psychological state trend reports of their students and can only obtain limited student information related to the warning when a warning is triggered; psychological counselors have higher privileges and can view more detailed risk profiles when a warning is triggered, but they are not allowed to access the original audio and video or any information that can directly identify the student when there is no warning; system administrators are only responsible for system operation and maintenance and have no right to view specific student psychological data by default, unless they have obtained explicit authorization and meet the compliance review conditions, in which case they can access the data within a controlled scope.

[0055] Furthermore, the method provided in this application establishes the core privacy principle of "monitoring only, not replaying": for the purpose of real-time analysis and early warning, except for anonymized and aggregated data used for model training, the original video and audio are destroyed or irreversibly anonymized after feature extraction is completed, and are not stored or replayed for a long time; any request for retrospective viewing must be subject to strict approval and multi-party authorization, and may only be allowed in extreme emergency situations such as life safety, and the entire process is audited.

[0056] At the same time, before deploying sensing devices, students and their guardians should be fully informed of the system's purpose, working principle, data collection scope, privacy protection measures, data processing methods, and storage period, and monitoring can only be activated after obtaining written consent from students and their guardians; for underage students, it is essential to ensure that their guardians obtain valid consent in order to protect the students' legitimate rights and ensure compliant operation.

[0057] Step 104: Input the multi-dimensional behavioral feature sequence into the trained psychological state recognition model to output psychological state representation information for characterizing the psychological state of the target student.

[0058] Among them, the psychological state representation information is a high-dimensional real-time vector, including... t The first moment i Emotional characteristics (such as anxiety scores), the first j Behavioral characteristics (such as duration of time alone), and the first k Various speech features (such as speech rate variation rate).

[0059] In one specific implementation, the mental state recognition model can be a multi-branch attention fusion neural network specifically designed for multimodal temporal data fusion analysis. This network aims to fully mine and fuse deep temporal information from different behavioral modalities to construct a comprehensive and dynamic representation of mental states.

[0060] Specifically, the input layer of this mental state recognition model can contain three independent input terminals: a first input terminal, a second input terminal, and a third input terminal, which are responsible for receiving the preprocessed facial expression feature sequence, behavioral trajectory feature sequence, and speech emotion feature sequence, respectively. This separate input design can ensure the structural integrity of the original data of each modality, laying the foundation for subsequent modality-specific coding.

[0061] The feature encoding layer of this mental state recognition model consists of three parallel encoding sub-networks: a first encoding sub-network, a second encoding sub-network, and a third encoding sub-network. These sub-networks are connected one-to-one with the three input terminals. Each encoding sub-network is a temporal feature extractor, such as a Long Short-Term Memory network, a temporal convolutional network, or a Transformer encoder structure. Its task is to perform high-order abstraction and context-dependent modeling on the original input feature sequence. Through encoding, the original temporal features are transformed into sequence representations containing rich semantic information; that is, the facial expression feature sequence is encoded as the first sequence representation, the behavioral trajectory feature sequence is encoded as the second sequence representation, and the speech emotion feature sequence is encoded as the third sequence representation.

[0062] The attention fusion layer of this mental state recognition model receives sequence representations from three coding sub-networks and performs deep interaction and fusion based on self-attention and / or cross-attention mechanisms. The self-attention mechanism allows the model to focus on important information at different time steps within each modality sequence; while the cross-attention mechanism allows a sequence representation of one modality to query relevant information in the sequence representation of another modality, thereby establishing cross-modal semantic associations. Through this attention mechanism, the mental state recognition model can adaptively assign differentiated weights to features at different modalities and time points, achieving selective and weighted fusion of information, ultimately generating a unified and information-rich fused representation. This fused representation integrates the multidimensional, temporal variation patterns of students' emotions, behaviors, and speech.

[0063] Finally, the output layer of the mental state recognition model (e.g., one or more fully connected layers) is used to further map and integrate the above fusion representation, outputting the final mental state representation information. This information can be a multi-dimensional vector (e.g., a mental health state vector), where each dimension corresponds to a quantified mental state indicator (e.g., an emotional stability index, a social engagement index, etc.), thus comprehensively and structurally representing the current mental state of the target student.

[0064] Based on the network structure of the mental state recognition model described above, in some embodiments, when inputting multi-dimensional behavioral feature sequences into the trained mental state recognition model to output mental state representation information for representing the mental state of the target student, the expression feature sequence can be input into the first encoding sub-network to obtain the first sequence representation; the behavioral trajectory feature sequence can be input into the second encoding sub-network to obtain the second sequence representation; the voice emotion feature sequence can be input into the third encoding sub-network to obtain the third sequence representation; then, the first sequence representation, the second sequence representation, and the third sequence representation can be input into the attention fusion layer to obtain the fused representation; finally, the mental state representation information is output through the output layer based on the fused representation.

[0065] Step 106: Identify the psychological state of the target student based on psychological state representation information.

[0066] In some embodiments, when identifying the psychological state of a target student, such as Figure 2 As shown, it includes the following steps: Step 202: First, calculate the target student's emotional stability index, social participation index, loneliness trend index, and behavioral abnormality index based on the psychological state representation information.

[0067] Among them, the emotional stability index is used to characterize the dispersion of the target student's facial and vocal emotional fluctuations within a preset time window.

[0068]

[0069] in, Variance representing emotional characteristics; Indicates the frequency of changes in emotional characteristics; Indicates the duration of extreme emotions; This indicates an index of emotional stability.

[0070] The Social Engagement Index is used to characterize the frequency and duration of target students' spatial proximity or interaction with others in public areas of the campus.

[0071]

[0072] in, Indicates social engagement index; Indicates social distancing; Indicates the duration of time spent alone; This indicates frequent group interaction.

[0073] The Loneliness Trend Index is used to characterize the rate of decline or persistently low value of the social engagement index over several consecutive days.

[0074]

[0075] in, Indicator of loneliness trends; Indicates the duration of time spent alone; This indicates a trend of declining social engagement; This indicates abnormal loitering behavior.

[0076] The behavioral abnormality index is used to characterize the degree to which a target student's behavioral trajectory or activity rhythm deviates from their personal historical baseline or the statistical distribution of their peer group.

[0077] The Behavioral Abnormality Index (BAI) can be calculated by combining abnormal wandering and disordered work and rest patterns; Optionally, these metrics are calculated using machine learning models combined with empirical rules from psychological experts, where the machine learning models include support vector machines (SVM) or random forests (RF).

[0078] Step 204: Identify the psychological state of the target students based on the emotional stability index, social participation index, loneliness trend index, and behavioral abnormality index.

[0079] In this embodiment, the psychological state of the target student can be identified based on the emotional stability index, social participation index, loneliness trend index, and abnormal behavior index, along with preset early warning rules. Risk assessment can then be performed and early warning information generated.

[0080] Optionally, in one implementation, the warning rule is a composite rule based on a combination of multi-level thresholds and durations set by dynamic indicators.

[0081] Specifically, the early warning rules include: If the loneliness trend index And the duration exceeds If this occurs within an hour, it triggers a high-risk warning for loneliness; If the emotional stability index exist The decline within a certain period of time exceeded If the expression shows a consistently high score for sadness or anxiety, a warning of a sharp emotional fluctuation is triggered. If a student lingers abnormally in a specific area for an extended period without interacting with others, a warning is triggered based on facial emotion recognition results, indicating a possible need for help or low mood.

[0082] Optionally, when the system detects that dynamic indicators meet the early warning rules, it generates a risk profile containing information such as student desensitization labels, risk type, trigger indicators, duration, current psychological state representation information, abnormal indicator change trends, and suggested points of attention. This profile is then pushed to relevant personnel such as psychological counselors, class teachers, or grade managers through an encrypted channel.

[0083] Optionally, the data acquisition and / or feature extraction steps may also include privacy protection steps: real-time face anonymization processing of the acquired image data, and no original identifiable face images are stored after feature extraction; voiceprint anonymization processing of the acquired voice data, retaining only non-identifiable acoustic features used for emotion recognition; and encryption processing of all feature data during transmission and storage.

[0084] Optionally, in another implementation, an access control strategy is also included: the system adopts a role-based access control model, ordinary teachers can only view the de-identified psychological state trend report, psychological counselors can only view the detailed risk profile when an alert is issued, and system administrators do not have the right to view students' specific psychological data.

[0085] The method provided in this application acquires multi-dimensional behavioral feature sequences, such as facial expression feature sequences, behavioral trajectory feature sequences, and vocal emotion feature sequences, of the target student. These sequences are then fused and analyzed using a trained psychological state recognition model to output psychological state representation information. Finally, the psychological state of the target student is identified based on this representation information. This comprehensive modeling of multi-dimensional behavioral feature sequences to form a quantitative representation of the student's psychological state improves the ability to capture subtle fluctuations and gradual changes in psychological state, enabling dynamic identification of the student's psychological state. Furthermore, identifying the target student's psychological state through psychological state representation information allows for the output of recognition results before psychological problems become apparent, thereby enhancing the early detection and warning capabilities for non-acute, gradual psychological problems and reducing response lag.

[0086] To address the shortcomings of existing technologies in providing real-time, comprehensive, and dynamic identification of students' psychological states, as well as the lack of effective early detection and warning capabilities, this application provides a student psychological state identification device. A schematic diagram of the device's specific structure is shown below. Figure 3 As shown, it includes an acquisition module 31, a processing module 32, and a recognition module 33. The functions of each module are as follows: The acquisition module 31 is used to acquire the multi-dimensional behavioral feature sequence of the target student. The multi-dimensional behavioral feature sequence includes at least the facial expression feature sequence, the behavioral trajectory feature sequence, and the voice emotion feature sequence. Processing module 32 is used to input multi-dimensional behavioral feature sequences into the trained psychological state recognition model, so as to output psychological state representation information for representing the psychological state of the target student. The identification module 33 is used to identify the psychological state of the target student based on psychological state representation information.

[0087] Optionally, module 31 is used for: By deploying sensing devices in public areas of the campus, multimodal behavioral feature data containing target students is collected. The multimodal behavioral feature data includes at least video data and audio data. Feature extraction is performed based on multimodal behavioral feature data to obtain multidimensional behavioral feature sequences of target students.

[0088] Optionally, the mental state recognition model is a multi-branch attention fusion neural network, including at least: The input layer includes a first input terminal, a second input terminal, and a third input terminal, which are used to receive facial expression feature sequences, behavioral trajectory feature sequences, and speech emotion feature sequences, respectively. The feature encoding layer includes a first encoding sub-network, a second encoding sub-network, and a third encoding sub-network, which are respectively connected to the first input terminal, the second input terminal, and the third input terminal. It is used to perform sequence encoding on the facial expression feature sequence, the behavioral trajectory feature sequence, and the speech emotion feature sequence to obtain the first sequence representation, the second sequence representation, and the third sequence representation. The attention fusion layer is used to adaptively weight and fuse the first sequence representation, the second sequence representation, and the third sequence representation based on self-attention and / or cross-attention mechanisms to obtain a fused representation; The output layer is used to output psychological state representation information based on the fusion representation to characterize the psychological state of the target student.

[0089] Optionally, processing module 32 is used for: The multi-dimensional behavioral feature sequence is input into the trained psychological state recognition model to output psychological state representation information for characterizing the psychological state of the target student, including: The facial expression feature sequence is input into the first encoding sub-network to obtain the first sequence representation; The behavioral trajectory feature sequence is input into the second encoding sub-network to obtain the second sequence representation; The speech emotion feature sequence is input into the third coding sub-network to obtain the third sequence representation; The first sequence representation, the second sequence representation, and the third sequence representation are input into the attention fusion layer to obtain the fused representation; Based on the fusion representation, the output layer outputs mental state representation information.

[0090] Optionally, the recognition module 33 is used for: Identifying the psychological state of target students based on psychological state representation information includes: The target students' emotional stability index, social participation index, loneliness trend index, and behavioral abnormality index were calculated based on their psychological state representation information. The psychological state of target students is identified based on the emotional stability index, social participation index, loneliness trend index, and behavioral abnormality index.

[0091] Optional, an emotional stability index is used to characterize the dispersion of the target student's facial and vocal emotional fluctuations within a preset time window; The Social Engagement Index is used to characterize the frequency and duration of target students' spatial proximity or interaction with others in public areas of the campus. The Loneliness Trend Index is used to characterize the rate of decline or persistently low value of the social engagement index over several consecutive days. The behavioral abnormality index is used to characterize the degree to which a target student's behavioral trajectory or activity rhythm deviates from their personal historical baseline or the statistical distribution of their peer group.

[0092] Using the device provided in this application embodiment, multi-dimensional behavioral feature sequences such as facial expression feature sequences, behavioral trajectory feature sequences, and voice emotion feature sequences of target students are acquired. These sequences are then fused and analyzed using a trained psychological state recognition model to output psychological state representation information. Finally, the psychological state of the target student is identified based on the psychological state representation information. In this way, by comprehensively modeling the multi-dimensional behavioral feature sequences and forming a quantitative representation of the student's psychological state, the ability to capture subtle fluctuations and gradual changes in psychological state can be improved, enabling dynamic identification of the student's psychological state. Furthermore, by identifying the psychological state of the target student through psychological state representation information, identification results can be output before psychological problems become apparent, thereby improving the early detection and warning capabilities for non-acute, gradual psychological problems and reducing response lag.

[0093] Figure 4 To illustrate the hardware structure of an electronic device according to various embodiments of this application, the electronic device 400 includes, but is not limited to, components such as: a radio frequency unit 401, a network module 402, an audio output unit 403, an input unit 404, a sensor 405, a display unit 406, a user input unit 407, an interface unit 408, a memory 409, a processor 410, and a power supply 411. Those skilled in the art will understand that... Figure 4 The electronic device structures shown are not intended to limit the electronic device. An electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. In the embodiments of this application, the electronic device includes, but is not limited to, mobile phones, tablets, laptops, PDAs, in-vehicle terminals, wearable devices, and pedometers.

[0094] The processor 410 is used to acquire a multi-dimensional behavioral feature sequence of the target student, which includes at least an expression feature sequence, a behavioral trajectory feature sequence, and a voice emotion feature sequence; input the multi-dimensional behavioral feature sequence into a trained psychological state recognition model to output psychological state representation information for representing the psychological state of the target student; and identify the psychological state of the target student based on the psychological state representation information.

[0095] The memory 409 is used to store a computer program that can run on the processor 410, which, when executed by the processor 410, implements the aforementioned functions implemented by the processor 410.

[0096] It should be understood that, in this embodiment, the radio frequency unit 401 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink data from the base station and processes it with the processor 410; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 401 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc. Furthermore, the radio frequency unit 401 can also communicate with networks and other devices through a wireless communication system.

[0097] The electronic device provides users with wireless broadband internet access through network module 402, such as helping users send and receive emails, browse web pages, and access streaming media.

[0098] The audio output unit 403 can convert audio data received by the radio frequency unit 401 or the network module 402 or stored in the memory 409 into audio signals and output them as sound. Furthermore, the audio output unit 403 can also provide audio output related to specific functions performed by the electronic device 400 (e.g., call signal reception sound, message reception sound, etc.). The audio output unit 403 includes a speaker, a buzzer, and a receiver, etc.

[0099] Input unit 404 is used to receive audio or video signals. Input unit 404 may include a graphics processing unit (GPU) 4041 and a microphone 4042. The GPU 4041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on display unit 406. The image frames processed by GPU 4041 can be stored in memory 409 (or other storage medium) or transmitted via radio frequency unit 401 or network module 402. Microphone 4042 can receive sound and process such sound into audio data. The processed audio data can be converted into a format that can be transmitted to a mobile communication base station via radio frequency unit 401 in telephone call mode.

[0100] The electronic device 400 also includes at least one sensor 405, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 4061 according to the ambient light level, and the proximity sensor can turn off the display panel 4061 and / or backlight when the electronic device 400 is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used to identify the posture of the electronic device (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. The sensor 405 may also include a fingerprint sensor, pressure sensor, iris sensor, molecular sensor, gyroscope, barometer, hygrometer, thermometer, infrared sensor, etc., which will not be described in detail here.

[0101] The display unit 406 is used to display information input by the user or information provided to the user. The display unit 406 may include a display panel 4061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0102] User input unit 407 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of electronic devices. Specifically, user input unit 407 includes a touch panel 4071 and other input devices 4072. Touch panel 4071, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 4071). Touch panel 4071 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 410, which receives and executes commands from the processor 410. In addition, touch panel 4071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. Besides touch panel 4071, user input unit 407 may also include other input devices 4072. Specifically, other input devices 4072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here.

[0103] Furthermore, the touch panel 4071 can cover the display panel 4061. When the touch panel 4071 detects a touch operation on or near it, it transmits the information to the processor 410 to determine the type of touch event. Subsequently, the processor 410 provides corresponding visual output on the display panel 4061 based on the type of touch event. Although in Figure 4 In this embodiment, the touch panel 4071 and the display panel 4061 are two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 4071 and the display panel 4061 can be integrated to realize the input and output functions of the electronic device. The specific implementation is not limited here.

[0104] Interface unit 408 serves as an interface for connecting external devices to electronic device 400. For example, external devices may include a wired or wireless headphone port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 408 can be used to receive input from external devices (e.g., data, power, etc.) and transmit the received input to one or more components within electronic device 400, or it can be used to transmit data between electronic device 400 and external devices.

[0105] The memory 409 can be used to store software programs and various data. The memory 409 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 409 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0106] The processor 410 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 409, and by calling data stored in the memory 409, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. The processor 410 may include one or more processing units; preferably, the processor 410 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 410.

[0107] The electronic device 400 may also include a power supply 411 (such as a battery) that supplies power to various components. Preferably, the power supply 411 can be logically connected to the processor 410 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system.

[0108] In addition, the electronic device 400 includes some functional modules not shown, which will not be described in detail here.

[0109] Preferably, this application embodiment also provides an electronic device, including a processor 410, a memory 409, and a computer program stored in the memory 409 and executable on the processor 410. When the computer program is executed by the processor 410, it implements the various processes of the above-described student psychological state recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0110] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described student psychological state recognition method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0111] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0112] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0113] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0114] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0115] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0116] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0117] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0118] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0119] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for identifying students' psychological states, characterized in that, include: Obtain a multi-dimensional behavioral feature sequence of the target student, wherein the multi-dimensional behavioral feature sequence includes at least an expression feature sequence, a behavioral trajectory feature sequence, and a voice emotion feature sequence; The multi-dimensional behavioral feature sequence is input into the trained psychological state recognition model to output psychological state representation information for characterizing the psychological state of the target student. The psychological state of the target student is identified based on the psychological state representation information.

2. The method as described in claim 1, characterized in that, Obtain the multi-dimensional behavioral feature sequence of the target student, including: By using sensing devices deployed in public areas of the campus, multimodal behavioral feature data containing the target students is collected, wherein the multimodal behavioral feature data includes at least video data and audio data; Feature extraction is performed based on the multimodal behavioral feature data to obtain the multidimensional behavioral feature sequence of the target student.

3. The method as described in claim 1, characterized in that, The mental state recognition model is a multi-branch attention fusion neural network, which includes at least: The input layer includes a first input terminal, a second input terminal, and a third input terminal, which are respectively used to receive the facial expression feature sequence, the behavior trajectory feature sequence, and the voice emotion feature sequence; The feature encoding layer includes a first encoding sub-network, a second encoding sub-network, and a third encoding sub-network, which are respectively connected to the first input terminal, the second input terminal, and the third input terminal, and are used to perform sequence encoding on the facial expression feature sequence, the behavioral trajectory feature sequence, and the speech emotion feature sequence to obtain a first sequence representation, a second sequence representation, and a third sequence representation. An attention fusion layer is used to adaptively weight and fuse the first sequence representation, the second sequence representation, and the third sequence representation based on self-attention and / or cross-attention mechanisms to obtain a fused representation; The output layer is used to output psychological state representation information based on the fusion representation to characterize the psychological state of the target student.

4. The method as described in claim 3, characterized in that, The multi-dimensional behavioral feature sequence is input into the trained psychological state recognition model to output psychological state representation information for characterizing the psychological state of the target student, including: The facial expression feature sequence is input into the first encoding sub-network to obtain the first sequence representation; The behavioral trajectory feature sequence is input into the second encoding sub-network to obtain the second sequence representation; The speech emotion feature sequence is input into the third coding sub-network to obtain the third sequence representation; The first sequence representation, the second sequence representation, and the third sequence representation are input into the attention fusion layer to obtain the fused representation; Based on the fusion representation, the psychological state representation information is output through the output layer.

5. The method as described in claim 1, characterized in that, Identifying the psychological state of the target student based on the aforementioned psychological state representation information includes: The target student's emotional stability index, social participation index, loneliness trend index, and behavioral abnormality index are calculated based on the psychological state representation information. The psychological state of the target student is identified based on the emotional stability index, the social engagement index, the loneliness trend index, and the behavioral abnormality index.

6. The method as described in claim 5, characterized in that, The emotional stability index is used to characterize the degree of dispersion of the target student's facial expressions and vocal emotional fluctuations within a preset time window. The social engagement index is used to characterize the frequency and duration of the target student's spatial proximity or interaction with others in public areas of the campus; The loneliness trend index is used to characterize the rate of decline or persistently low value of the social engagement index over several consecutive days. The behavioral anomaly index is used to characterize the degree to which the target student's behavioral trajectory or activity rhythm deviates from his / her personal historical baseline or the statistical distribution of his / her peer group.

7. A student psychological state recognition device, characterized in that, It includes an acquisition module, a processing module, and a recognition module, among which: The acquisition module is used to acquire a multi-dimensional behavioral feature sequence of the target student, wherein the multi-dimensional behavioral feature sequence includes at least an expression feature sequence, a behavioral trajectory feature sequence, and a voice emotion feature sequence. The processing module is used to input the multi-dimensional behavioral feature sequence into the trained psychological state recognition model, so as to output psychological state representation information for representing the psychological state of the target student. The identification module is used to identify the psychological state of the target student based on the psychological state representation information.

8. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the student mental state recognition method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the student mental state recognition method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the student mental state recognition method according to any one of claims 1 to 6.