Information processing methods and devices for psychological crisis screening

By employing a hierarchical processing logic that combines low-resolution initial screening of the population with high-resolution precise analysis of individuals, and integrating multimodal feature fusion and the Transformer structure, the problems of high cost, numerous misjudgments, and poor perceptibility in campus psychological crisis monitoring have been solved, achieving accurate, continuous, and perceptible campus psychological crisis screening.

CN120998534BActive Publication Date: 2026-01-30浙江连信科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511508218.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-30
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing technologies are insufficient to achieve full-time, comprehensive, and accurate monitoring of psychological crises among students on campus. Furthermore, existing systems are costly, have significant data redundancy, and single-modal analysis can easily lead to misjudgments and over-intervention, lacking user-friendliness for seamless monitoring.

Method used

We employ a hierarchical processing logic that combines low-resolution initial screening of the group with on-demand high-definition precise individual analysis. By integrating multimodal feature fusion and a multi-temporal Transformer structure with attention mechanisms, we construct a psychological crisis screening method, which includes initial screening of group behavioral characteristics, precise analysis of individual characteristics, and classroom interaction verification to generate abnormal behavior reports.

Benefits of technology

Significantly reduces computing power consumption and costs, improves the accuracy of psychological crisis identification, reduces excessive intervention, enables continuous dynamic monitoring, adapts to campus scenarios for seamless monitoring, and ensures timely identification of early hidden risk signals.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention discloses an information processing method and apparatus for psychological crisis screening, comprising the following steps: acquiring group behavioral characteristic information of students through information collection devices installed on campus; determining the preliminary psychological risk level of an individual based on the acquired group behavioral characteristic information; when the preliminary psychological risk level of an individual exceeds a set threshold, performing precise analysis on the individual characteristic information, extracting multimodal specific features and performing time-series modeling to determine the precise psychological risk level of the individual; when the precise psychological risk level exceeds a set threshold, focusing on the abnormal behavior areas of the individual through a heatmap visualization model, generating an abnormal behavior report containing the abnormal behavior areas, psychological risk level, and determination criteria, and sending the behavior report to the administrator for reminder. This invention achieves low-cost, low-intrusive acquisition of behavioral information for psychological analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information data processing technology, and specifically relates to an information processing method and device for psychological crisis screening. Background Technology

[0002] In recent years, mental health problems among students on campus have become increasingly prevalent. In daily campus settings, psychological crises among students often manifest as latent behaviors, such as prolonged periods of silence with their heads down in class, curling up alone in a corner during breaks, refusing to participate in group activities, and frequently avoiding eye contact when communicating with peers. If these latent signals are not identified in a timely manner through continuous and accurate monitoring, psychological crises can easily develop from early latent states to later overt and extreme consequences.

[0003] Current approaches to handling psychological crises among students on campus mainly fall into two categories: traditional manual monitoring and technology-based early warning systems based on image recognition. Traditional methods rely on teachers' daily observations and periodic questionnaires. The former is limited by teachers' energy and subjective experience, making it difficult to achieve round-the-clock, comprehensive monitoring; the latter suffers from low frequency and delayed feedback, failing to capture real-time changes in psychological states. Existing technology-based early warning methods primarily employ continuous high-definition image acquisition and analysis, requiring high-performance hardware to process massive amounts of video data. This not only leads to high hardware procurement and operating costs but also generates a large amount of redundant data, increasing the processing burden.

[0004] Furthermore, such technologies often focus on a single modality, failing to achieve multimodal correlation analysis of body behavior, voice signals, and facial expressions. This makes it difficult to accurately assess psychological crises through the coordinated features of actions, language, and expressions. More importantly, existing systems frequently trigger risk assessments due to abnormally high frequencies of a single feature, which can easily lead to excessive intervention by teachers. This not only disrupts normal teaching order but also makes students clearly aware of being monitored, causing resistance and discomfort. It lacks the user-friendliness of seamless monitoring adapted to the campus setting, which is not conducive to the advancement of long-term screening work. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides an information processing method and apparatus for psychological crisis screening. It mainly relies on existing camera equipment to acquire information and performs two parts: initial screening and precise screening in a low-cost manner. It integrates the multimodal features of the object to improve the relevance of the information and increase the accuracy.

[0006] The technical solution adopted in this invention is as follows:

[0007] In a first aspect, the present invention provides an information processing method for psychological crisis screening, which collects and analyzes data from individuals on campus, including the following steps:

[0008] S100. Obtain group behavioral characteristic information of students through information collection devices installed on campus;

[0009] S200. Based on the obtained information on group behavior characteristics, conduct an initial screening analysis to determine the preliminary psychological risk level of an individual.

[0010] S300. When the initial psychological risk level of a single object exceeds the set threshold, the individual characteristic information of the single object is obtained for precise analysis. Multimodal specific feature extraction and temporal modeling are performed on the individual characteristic information to construct the behavioral profile of the object. Then, a multi-temporal Transformer structure is used to fuse multimodal specific features, and an attention mechanism is introduced to focus on abnormal behavioral fragments to determine the precise psychological risk level of the object.

[0011] S400: When the precise psychological risk level exceeds the set threshold, the abnormal behavior area of ​​the subject is monitored through the heat map visualization model, and an abnormal behavior report containing the abnormal behavior area, psychological risk level and judgment basis is generated. The behavior report is then sent to the manager for reminder.

[0012] In conjunction with the first aspect, the present invention provides a first embodiment of the first aspect, wherein the group behavior characteristic information in step S100 includes the interactive action characteristics, regional clustering characteristics, and activity level characteristics of the student group.

[0013] In conjunction with the first aspect, the present invention provides a second implementation of the first aspect, wherein the individual feature information obtained for a single object in step S300 includes the object's facial image features, eye movement and gaze behavior features, body behavior features and voice features;

[0014] Among them, facial image features are subtle facial expression information in high-resolution facial images, eye movement and gaze behavior features are eye movement and gaze state information of the object, body behavior features are the amplitude, frequency and posture information of the object's limb movements, and speech features are non-semantic acoustic information and semantic information of the object when speaking independently.

[0015] In conjunction with the first aspect, the present invention provides a third embodiment of the first aspect, wherein the specific process of multimodal specific feature extraction and temporal modeling in step S300 includes:

[0016] Facial micro-expression units were extracted by using the OpenFace tool in conjunction with a facial micro-expression unit recognition model to extract the intensity and frequency of 17 facial micro-expression units of the object.

[0017] Then, eye movement and fixation behavior features were extracted, including fixation duration, number of fixation avoidances, blink frequency, eyelid closure time percentage, and eye saccade speed.

[0018] Next, body behavior features are extracted. The 2D or 3D body key points of the object are extracted using the OpenPose algorithm or BlazePose algorithm. Based on the body key points, the amplitude of limb movements, the frequency of movements, and the asymmetry of movements are calculated.

[0019] Speech feature extraction involves denoising the speech signal of the subject and extracting non-semantic acoustic features such as pitch, speech rate, and number of pauses. At the same time, semantic analysis is performed on the speech content to extract semantic features related to psychological crisis.

[0020] The extracted facial micro-expression units, eye movement and gaze behavior features, body behavior features and voice features are sliced ​​according to a 5-second time window with a 50% overlap ratio to construct the feature time sequence of the object and form a behavior profile.

[0021] In conjunction with the first implementation of the first aspect, the present invention provides a fourth implementation of the first aspect, wherein the acquisition of group behavior characteristic information in step S100 uses the video stream output by the campus information collection device as the core analysis carrier, and the specific processing method includes:

[0022] First, low-resolution multi-frame video streams are used to analyze the interactive action characteristics of the group. The resolution of the sampled low-resolution video streams is no higher than 480P. A lightweight limb keypoint detection algorithm is used to extract the limb displacement and action frequency of each student in the group based on the obtained interactive action feature information to determine whether the interactive actions are abnormal. At the same time, in the initial state, single frames are selected from the low-resolution video stream at fixed time intervals of 5 minutes / frame and upscaled to high-resolution to extract the facial features of individual students. The preliminary psychological risk level is analyzed based on the interactive action features and facial features.

[0023] If the initial screening reveals an increase in the preliminary psychological risk level of a single student, the interval for acquiring high-definition images is shortened to 2 minutes / frame; if the preliminary psychological risk level further rises to high risk, the interval for acquiring high-definition images is shortened to 30 seconds / frame; when the preliminary psychological risk level of a single student exceeds the set threshold, the adjustment of the high-definition image acquisition interval is stopped, triggering the precise analysis process in step S300.

[0024] In conjunction with the first embodiment of the first aspect, the present invention provides a fifth embodiment of the first aspect, wherein in step 100, obtaining the group behavior characteristic information is specifically as follows:

[0025] First, low-cost preliminary data collection is carried out. The initial collection resolution of the collection equipment deployed on campus is set to a low resolution mode of 360P-480P. A fixed-size sliding window of 512×384 pixels or 640×480 pixels is used. The low-resolution video stream is divided into areas at a sliding frequency of 2-3 times per second. Each sliding window corresponds to a fixed physical area on campus.

[0026] Then, a preliminary screening analysis of large-scale behavior of multiple people in a region is performed. For multiple human targets in each sliding window, large-scale limb movement features are extracted using a lightweight behavior recognition algorithm. The large-scale movement features include limb displacement amplitude, movement frequency, and duration of abnormal posture. If at least one of the large-scale movement features is detected in the same region in two consecutive sliding windows, the region is marked as a region to be accurately detected.

[0027] Next, precise individual data collection is carried out. The resolution of the collection device corresponding to the area to be precisely detected is temporarily increased to a high-definition mode of 720P-1080P, and the video stream of the area is focused. The YOLO algorithm is used to detect and distinguish the individuals in the high-definition video stream. Combined with the extracted large-scale motion features, the specific individuals in the area who have made abnormal large-scale motions are matched and located.

[0028] For each specific individual located, their facial image, body posture, and voice signal are extracted from the high-definition video stream. Face alignment and voice noise reduction preprocessing are performed before proceeding to subsequent specific feature extraction and temporal modeling steps.

[0029] In conjunction with the first aspect, the present invention provides a sixth embodiment of the first aspect, wherein the criteria for determining the abnormal behavior report in step S400 include:

[0030] The report includes the specific types of abnormal features in the multimodal specific features, the duration of the abnormal features, and the synergistic occurrence of abnormal features in each modality. The abnormal behavior report also includes key audio and video screenshots from the initial screening and precise analysis process of the subject. After receiving the report, the administrator will manually review it based on the screenshots and heatmap. After the review is confirmed, the psychological intervention process for the student will be initiated.

[0031] In conjunction with the first aspect, the present invention provides a seventh embodiment of the first aspect, which further includes a classroom interaction embedded dynamic verification step before the precise analysis is initiated in step S300, specifically as follows:

[0032] Based on the potential abnormal characteristics corresponding to the object's preliminary psychological risk level, and combined with the subject teaching plan for the remaining courses of the day, personalized interactive tasks that are adapted to the teaching objectives are generated.

[0033] Personalized interactive tasks are synchronized with the instructors and embedded in the classroom teaching process; multimodal data of the subject during the interaction process is continuously collected through campus information collection devices, including body coordination, facial expression changes, and voice expression during the interaction.

[0034] Feature extraction is performed on the collected interactive multimodal data, and the changes in the object's preliminary psychological risk characteristics before and after the interaction are compared to adjust its preliminary psychological risk level.

[0035] If the adjusted preliminary psychological risk level still exceeds the set threshold, the precise analysis in step S300 is triggered; if the adjusted preliminary psychological risk level is lower than the set threshold, the precise analysis is not started for the time being, and steps S100-S200 and this dynamic verification step are repeated in the next teaching cycle to continuously monitor the psychological risk status of the subject.

[0036] Secondly, the present invention provides an apparatus for acquiring and analyzing information on objects in a campus and generating abnormal behavior reports for feedback in the information processing method for psychological crisis screening described in any of the above claims. The apparatus includes a server and a data acquisition module. The data acquisition module acquires group behavior characteristic information and individual characteristic information and sends them to the server for processing and analysis, and the server generates abnormal behavior reports.

[0037] The beneficial effects of this invention are as follows:

[0038] (1) This invention can significantly reduce computing power consumption and cost investment, and solve the problems of high cost and data redundancy in existing technologies. Through the hierarchical processing logic of low-resolution group screening and on-demand high-definition individual precise analysis, the initial screening stage only uses low-resolution video stream to analyze group behavior characteristics, without the need to process high-definition data at all times. The resolution of the target area is increased only when identifying potential risk objects, which greatly reduces unnecessary computing power occupation and data redundancy, effectively reduces hardware procurement and long-term operating costs, and is more suitable for the cost budget requirements of campuses.

[0039] (2) This invention realizes the correlation analysis of multimodal features, improves the accuracy of psychological crisis identification, and makes up for the limitations of single-modal analysis in existing technologies. By simultaneously extracting multi-dimensional features such as students' physical behavior, facial expressions, and voice signals, and combining them with a multi-temporal Transformer structure and attention mechanism to construct a fusion model, it can judge psychological crisis through the collaborative correlation of multimodal features, avoid misjudgment caused by abnormal single features, and improve the accuracy of identification;

[0040] (3) This invention can effectively reduce excessive intervention, alleviate students’ discomfort from being monitored, and solve the problems of high frequency of misjudgment and lack of user-friendliness in existing systems. By adding dynamic verification steps to classroom interaction, personalized interactive tasks adapted to normal classroom segments are generated based on the teaching plan. Student interaction data is collected and risk levels are adjusted in natural teaching scenarios. This can filter out misjudgment signals with single abnormal features, reduce unnecessary precise analysis and teacher intervention, and avoid exposing students’ potential risk attributes. This allows students to complete monitoring in a state of unconsciousness, reduce resistance, and ensure the long-term sustainable progress of screening work.

[0041] (4) This invention can realize continuous and dynamic monitoring of psychological crises, make up for the shortcomings of traditional manual monitoring in terms of insufficient coverage and delayed feedback. Through the closed-loop process of initial screening-dynamic verification-precise analysis, combined with the daily teaching cycle, it can continuously track changes in students' psychological state. It does not rely on teachers' subjective experience, and can capture the dynamic development process of psychological crises through time-series feature modeling, ensuring that early hidden risk signals are identified in a timely manner, buying time for subsequent intervention, and reducing the probability of extreme behaviors. Detailed Implementation

[0042] The present invention will be further explained below with reference to specific embodiments.

[0043] Example 1:

[0044] This embodiment discloses an information processing method for psychological crisis screening, which collects and analyzes characteristic information of students on campus to determine psychological risks.

[0045] This method involves connecting all information collection devices within the campus that can acquire information, obtaining information directly or by connecting to information storage hard drives, and then analyzing the information. Based on existing campus equipment, the method can be implemented on existing processing servers through software integration, or a dedicated system can be built to achieve the corresponding technical effects.

[0046] Specifically, this method mainly consists of two parts: initial screening and precision screening, and includes the following steps:

[0047] First, information on the group behavior characteristics of students is obtained through information collection devices set up on campus; then, preliminary analysis based on the obtained group behavior characteristics is used to determine the initial psychological risk level of individual individuals.

[0048] When the initial psychological risk level of a single object exceeds the set threshold, the individual characteristic information of the single object is obtained for precise analysis. Multimodal specific feature extraction and temporal modeling are performed on the individual characteristic information to construct the behavioral profile of the object. Then, a multi-temporal Transformer structure is used to fuse multimodal specific features, and an attention mechanism is introduced to focus on abnormal behavioral segments to determine the precise psychological risk level of the object.

[0049] When the precise psychological risk level exceeds the set threshold, the abnormal behavior area of ​​the subject is monitored through a heat map visualization model, and an abnormal behavior report containing the abnormal behavior area, psychological risk level and judgment basis is generated. The behavior report is then sent to the manager for reminder.

[0050] Furthermore, this embodiment is applied to a school, hereinafter referred to as the target school. The school is a medium-sized campus in the county, with 3 grades, 8 classes in each grade, 45-50 students in each class, a total of 24 classes, and about 1200 students. The campus public activity areas include the teaching building corridor, a standard plastic track, a canteen that can accommodate 500 people, and a small library. The daily teaching hours are from 7:30 to 17:30.

[0051] (a) Existing monitoring equipment resources

[0052] The target school already has a traditional security monitoring system deployed, so there is no need to add new dedicated cameras. Existing equipment can be reused directly, with the following specific configuration:

[0053] Teaching Classes: Each classroom has a 2-megapixel network surveillance camera installed on the back wall near the window corner. It supports switching between 1080P high-definition resolution and 480P low-resolution resolution, with a fixed frame rate of 25fps. It has a built-in microphone to support audio capture and comes with a 1TB local hard drive.

[0054] Public areas: One 2-megapixel camera of the same model is installed at each end of the corridor on each floor of the teaching building; one 3-megapixel PTZ camera is installed at each of the four corners of the playground; and two 2-megapixel cameras are installed in the canteen and library, respectively covering the dining area / reading area and the entrance / exit.

[0055] Data transmission: All cameras are connected to the campus gigabit Ethernet, supporting real-time transmission of audio and video streams to designated servers via TCP / IP protocol, and historical data can also be retrieved through local hard drives.

[0056] (II) Server Setup and System Architecture

[0057] Based on the target school's existing monitoring network, one new edge computing server will be added as the data processing core to build an integrated system for data acquisition, analysis, and output. The specific configuration and architecture are as follows:

[0058] Server hardware configuration: The PowerEdge R750 edge server is selected, with hardware parameters including: Intel Xeon Gold 6330 CPU, NVIDIA A10 GPU, 64GB DDR4 memory, 2TB SSD and 16TB HDD. This capacity stores key audio and video clips and abnormal behavior reports, with a retention period of 30 days. The server is deployed in the school's information center computer room and is equipped with UPS power supply to ensure that core analysis tasks are not interrupted in the event of a power outage.

[0059] System software architecture: The server is built on a Linux operating system and deploys three layers of functional modules, which work together through internal interfaces.

[0060] Data acquisition layer: Runs camera access and control program, receives 480P low-resolution audio and video streams from all cameras in real time, and supports remotely sending resolution switching and area focusing commands to designated cameras;

[0061] Data processing layer: integrates a group initial screening module, a precise analysis module and a dynamic verification module. The group initial screening module has a built-in simplified version of the OpenPose algorithm and a basic facial expression recognition model. The precise analysis module integrates the OpenFace tool, a self-developed facial micro-expression unit recognition model and a multi-temporal Transformer model.

[0062] Results Output Layer: Deploy report generation and push program, support drawing heatmaps of abnormal behavior, generate PDF reports, and push warning information to the mobile app of psychology teachers / class teachers via WebSocket protocol, while keeping a backup of the report on the local server.

[0063] II. Implementation Process of Psychological Crisis Screening Methods Based on the Above System

[0064] The following section details the implementation process of this method in the context of the target school's daily teaching environment, using the screening of student B in a certain class as an example:

[0065] Step S100: Obtaining Group Behavioral Characteristics

[0066] Data collection is triggered during daily teaching hours. All surveillance cameras are in 480P low-resolution mode by default, continuously collecting audio and video streams at 25fps and transmitting them to the server in real time via the campus LAN. The server's data acquisition layer receives data in the order of priority for classes and secondary priority for public areas, extracting one key frame every five frames (to reduce data processing load), while filtering out invalid audio data during noisy breaks, retaining only the audio signal of a single student speaking independently.

[0067] Group behavior feature extraction: The server's initial group screening module analyzes keyframes and valid audio:

[0068] Interactive action characteristics: The simplified OpenPose algorithm is used to extract key points of students' bodies in the class, such as head, torso, and arms, to analyze students' interactions during breaks. For example, whether student B participates in the classmate's jump rope game or actively reaches out to pass stationery during group discussions. If student B is detected to be sitting alone in his seat for 10 consecutive minutes without physical or verbal interaction with the students around him, he is marked as having abnormal interactive actions.

[0069] Facial features: By analyzing the facial contours of student B using a basic facial recognition model, if the proportion of eye closure time exceeds 60% and the corners of the eyes continue to drop in three consecutive keyframes, it is marked as abnormal facial features.

[0070] Speech content characteristics: When the teacher asks student B a question in class, the server extracts the audio of student B's answer, identifies semantic keywords such as "I don't want to answer," "I don't know," and "meaning," and marks them as having a negative tendency in speech content.

[0071] Step S200: Preliminary determination of psychological risk level

[0072] Based on the three types of abnormal features extracted above, the server's initial screening module determines the preliminary psychological risk level of student B by referring to the basic assessment indicators in the student psychological crisis screening guidelines.

[0073] Assessment logic: Abnormal interactive actions, abnormal facial expressions, and negative speech content meet the criteria for medium risk assessment.

[0074] Low risk: single abnormal feature; medium risk: two or more abnormal features lasting for at least 5 minutes; high risk: three abnormal features accompanied by self-harm tendencies.

[0075] Results storage: The server associates student B's preliminary risk level, abnormal characteristic type, and occurrence time with his anonymous ID, temporarily stores it in the SSD cache, and triggers the high-definition image interval adjustment, shortening the high-definition image acquisition interval of a certain class's camera from the default 5 minutes / frame to 2 minutes / frame.

[0076] Step S300: Precise Analysis

[0077] Threshold trigger: The server continuously monitors student B's initial risk level. If the initial risk level remains medium risk in three rescreenings within one hour, it will automatically send a command to the surveillance camera of a certain class to switch the resolution to 1080P high-definition mode and focus on the seat area of ​​student B in the second column of the third row.

[0078] Individual Feature Information Acquisition: A high-definition camera focuses and captures a 1080P audio and video stream of student B. The server's precision analysis module extracts multimodal specific features from this stream data.

[0079] Facial micro-expression unit extraction: Using the OpenFace tool combined with a self-developed model, 17 facial micro-expression units of student B were extracted, such as AU15, the intensity of the corner of the mouth pulling down is 0.7 / 1.0, and the frequency is 6 times / minute; AU4, the intensity of the frown is 0.5 / 1.0, and the frequency is 4 times / minute.

[0080] Eye movement and gaze behavior feature extraction: Using eye movement tracking algorithms, the gaze duration of student B is extracted, such as the gaze duration of student B only accounting for 15% of the class time, which is lower than the class average of 40%. The number of gaze avoidances is also extracted, such as the number of times student B avoids eye contact when the teacher looks at student B, which is 8 times per class. The blinking frequency is also extracted, such as 20 times per minute, which is higher than the class average of 12 times per minute.

[0081] Body behavior feature extraction: The OpenPose algorithm is used to extract 2D body key points of student B and calculate the range of limb movements, such as the range of the hand gripping the table seam is 3cm and the frequency is 8 times / minute; movement asymmetry, such as the left shoulder being 2cm lower than the right shoulder when sitting, and a slight curled posture.

[0082] Speech feature extraction: Noise reduction was performed on the audio of student B answering questions in the 15:00 English class, and pitch, speech rate, and number of pauses were extracted. At the same time, semantic analysis revealed negative expressions such as "I feel that I can't do anything well".

[0083] Temporal modeling and risk assessment: The server slices the above features into 5-second time windows to construct a feature sequence containing 240 temporal segments, forming a behavioral profile of student B. Then, the sequence is encoded by four branch encoders of a multi-temporal Transformer structure. The fusion layer focuses on the collaborative abnormal segments of picking at the table seam, drooping corners of the mouth, and low voice through the Cross-Attention mechanism, and finally outputs the accurate psychological risk level of student B as medium risk.

[0084] Step S400: Abnormal Behavior Report Generation and Alerts

[0085] Report Generation: The server's output layer automatically generates an abnormal behavior report, including multi-dimensional visualizations. Specifically, it uses a heatmap to display hotspots in student activity areas within a specified time period; it uses keyframes to dynamically highlight abnormal actions of student B, such as picking at the seams of the table, drooping the corners of their mouth, and hunching their shoulders; and it attaches screenshots of these keyframes for administrators to review.

[0086] The following is a sample report:

[0087] Anonymous student ID, Grade 7, Class 2-02;

[0088] Precise psychological risk level: Medium risk;

[0089] The criteria for judgment include, for example, the frequency of AU15's mouth corner pulling down 6 times / minute, which exceeds the normal behavior threshold (2.4 times / minute); the number of speech pauses 5 times / sentence, which exceeds the normal behavior threshold (1.5 times / sentence).

[0090] Three screenshots of key abnormal segments are shown, corresponding to abnormal behaviors at 14:35, 15:00, and 15:30 respectively.

[0091] Report push: The server pushes the report in real time to the mobile phones of school psychologist A and class teacher B via WeChat mini program, and sends a pop-up reminder that the student in Class 2 of Grade 7 has medium-risk psychological signals and should be reviewed in time.

[0092] Manual review and intervention: After receiving the report, Teacher A reviewed it using heatmaps and screenshots. After class that day, Teacher A invited student B to the counseling room for a talk and confirmed that B had developed self-doubt due to recent exam failures. Subsequently, a two-week follow-up counseling plan was developed to prevent the risk from escalating further.

[0093] Furthermore, based on the extraction of group behavior features, contextualized extraction logic and quantitative judgment criteria are added for interactive action features, facial expression features, and speech content features to improve the accuracy of feature extraction. Among them, interactive action features are extracted quantitatively for different scenarios, and then differentiated analysis logic for break-time scenarios and classroom group scenarios is added.

[0094] During breaks: The frequency of physical contact was extracted using a simplified version of the OpenPose algorithm, such as the number of times student B touched his partner's arm or stood side by side. The normal threshold is no less than 3 times per 10 minutes, but student B only did 0 times. Active participation was also measured, such as whether student B actively moved towards the group. Student B remained within a 1-meter radius of his seat and did not move actively.

[0095] Classroom group scenario: Extract body coordination actions, such as whether to pass stationery or turn to face group members. Student B kept his head down the whole time and did not turn to face group members. The number of body coordination actions was 0. The abnormal interaction action was defined as two or more indicators being lower than 50% of the class average in a single scenario.

[0096] Facial features: Supplement specific judgment thresholds

[0097] Additional quantitative standards were added for the proportion of eye closure time and the downward pull of the corners of the mouth in the original embodiment:

[0098] Eye closure time percentage: Based on 480P keyframes, the normal threshold is no more than 30% calculated by the distance between key points on the eyelids. Student B's PERCLOS for three consecutive keyframes were 62%, 65%, and 60%, respectively, all exceeding the threshold.

[0099] Mouth corner direction: The angle between the line connecting the key points on both sides of the mouth corner and the tip of the nose is used to determine the direction. The normal angle is not less than 15°. Student B's angle is -8° and has not changed for 5 frames, which is judged as abnormal facial expression.

[0100] Speech content characteristics: clear semantic analysis logic

[0101] Supplementing the negative semantic keyword database and semantic matching rules:

[0102] Keyword database: Pre-set negative words related to student psychological crisis, including more than 50 core words such as "don't want to study", "nobody pays attention to me", "it's boring", "too tired", "don't want to live", etc., which are divided into mild, moderate and severe negative according to the degree of risk.

[0103] Analysis rules: The audio-to-text content of student B's answer was segmented using the jieba word segmentation tool in Python. Three mildly negative words were matched: "don't want to answer", "can't answer well", and "uninteresting". The semantic negative proportion reached 60%, which was judged to indicate that the content of the speech had a negative tendency.

[0104] Furthermore, regarding the refinement and optimization of individual feature information and multimodal specific feature extraction, based on the individual feature information acquisition in the above implementation, additional technical parameters and extraction logic are added for facial image features, eye movement and gaze behavior features, body behavior features, and speech features:

[0105] Facial image features: The high resolution is clearly 1080P (1920×1080). After the camera focuses, it can clearly capture 56 key points on the face, including subtle facial expressions such as slight frowning of the eyebrows, slight trembling of the eyelids, and slight upward turning of one corner of the mouth. These features cannot be recognized at the low resolution of 480P and require the support of 1080P high-definition images.

[0106] Eye movement and fixation behavior characteristics: Supplementing the specific extraction methods and quantitative indicators of eye movements and fixation states:

[0107] Fixation duration: The fixation duration of student B on the teacher was calculated by using the iris center point for localization. The teacher's standing area was preset to the image coordinates (800, 400) - (1000, 600). The cumulative fixation duration on the teacher was only 6 minutes and 45 seconds in one class, with an average of 18 minutes per class.

[0108] Number of gaze avoidances: The number of times the center point of student B's iris deviates from the teacher's direction when the teacher's gaze turns to student B, up to 8 times in one class period.

[0109] Blinking frequency: 20 blinks per minute, with a class average of 12 blinks.

[0110] Percentage of eyelid closure time: The average value within 10 minutes is 58%, and the normal threshold is no more than 30%.

[0111] Eye saccade speed: Calculated by the distance and time the iris moves in the horizontal / vertical direction. The horizontal saccade speed is 15° / s, which is lower than the class average of 25° / s.

[0112] Physical behavioral characteristics: Clarify the calculation basis for the amplitude, frequency and posture of limb movements.

[0113] Range of motion: Based on 25 body key points extracted using the OpenPose algorithm, including 1 for the head, 3 for the torso, 6 for the arms, and 15 for the legs, the displacement distance of the hand key points was then calculated. The displacement range of student B's action of picking at the seam of the table was 3cm.

[0114] Movement frequency: The hand-picking motion on the table seam is repeated 8 times per minute, and the leg-shaking motion is repeated slightly 12 times per minute.

[0115] Posture: Judging by the angle of the line connecting key points of the torso, such as the acromion, lumbar spine, and hip, when student B is sitting, the angle between the line connecting the acromion and lumbar spine and the horizontal direction is 65°, while the normal sitting posture is about 85°, indicating that the student is slightly curled up.

[0116] Speech features: Refine the logic for extracting non-semantic acoustic information and semantic information during independent speech.

[0117] Non-semantic acoustic information: Extract pitch, speech rate, and number of pauses from student B's audio response.

[0118] Semantic information: The audio-to-text content was semantically analyzed using the BERT pre-trained model. Two moderately negative semantics were matched: "can't do anything well" and "can't learn well". The semantic risk score was 0.72 out of 1.0. A score of not less than 0.6 was considered a semantic anomaly.

[0119] For the refinement and optimization of multimodal specific feature extraction and temporal modeling, based on the multimodal feature extraction in the original implementation, tool parameters, extraction result examples, and temporal modeling details are added for the four types of feature extraction steps specified by the user:

[0120] Facial micro-expression unit extraction:

[0121] Using the OpenFace 2.0 tool combined with a self-developed student micro-expression optimization model, and taking into account the characteristics of children's small facial contours and low expression amplitude, the AU intensity judgment threshold was adjusted to extract 17 facial micro-expression units, including AU1-inner brow lift, AU2-outer brow lift, AU4-brow furrow, AU5-levator palpebrae superioris muscle, AU6-cheek lift, AU7-eyelid compression, AU9-nasolabial fold muscle, AU10-upper lip lift, AU12-corner of mouth lift, AU14-corner of mouth stretch, AU15-corner of mouth pull-down, AU17-chin lift, AU20-lip stretch, AU23-lip compression, AU25-lip separation, AU26-jaw droop, and AU43-eyelid closure.

[0122] The key AU data for student B are as follows: AU4 frowning intensity 0.5 / 1.0, frequency 4 times / minute; AU15 corner of mouth pulling down intensity 0.7 / 1.0, frequency 6 times / minute; AU43 eyelid closure intensity 0.8 / 1.0, frequency consistent with blinking frequency.

[0123] Eye movement and fixation behavior feature extraction:

[0124] Using iris localization and eye-tracking algorithms, in addition to the indicators in the original embodiment, fixation dispersion is added. For example, if student B's fixation is dispersed across the desktop area, the dispersion value is 18, while the class average is 8, forming a 5-dimensional eye-tracking feature vector, including fixation duration, number of fixation avoidances, blink frequency, PERCLOS, and fixation dispersion.

[0125] Extraction of physical and behavioral features:

[0126] The BlazePose algorithm was used to extract 33 3D body key points, with a focus on calculating asymmetry indicators. These included the vertical height difference between the left and right acromion key points and the horizontal distance difference between the left and right knee key points, ultimately forming 3D body behavior characteristics in terms of movement amplitude, frequency, and asymmetry.

[0127] Speech feature extraction:

[0128] Non-semantic acoustic features were extracted using the Librosa library, with the fundamental frequency standard deviation added in addition to the original indicators; semantic features were obtained by combining keyword matching and semantic sentiment scores, with sentiment tendency labels added in addition to the original analysis, forming a 6-dimensional speech feature vector, namely pitch, speech rate, number of pauses, fundamental frequency standard deviation, percentage of negative semantics, and sentiment tendency label.

[0129] Temporal modeling:

[0130] The data was sliced ​​into 5-second time windows, meaning one feature window was generated every 2.5 seconds. High-definition data of student B was continuously collected for 10 minutes, resulting in a total of 240 feature windows. Each window integrated 17-dimensional AU features, 5-dimensional eye-tracking features, 3-dimensional body behavior features, and 6-dimensional speech features to form a 31-dimensional feature vector. Finally, a 240×31 feature time sequence was constructed as the core data for the behavioral profile. The behavioral profile also includes the percentage of abnormal features, such as the percentage of abnormal feature windows in student B reaching 72%.

[0131] Furthermore, in order to implement a low-cost deployment solution in existing schools using existing hardware such as surveillance cameras, the method is optimized and limited.

[0132] It should be noted that this optimization is aimed at addressing the issues of insufficient adaptability of the aforementioned solutions, difficulty in balancing computational costs and recognition accuracy, and low individual positioning efficiency in the two core scenarios of classrooms and public areas for schools of different sizes. The optimization achieves the following effects through refined design:

[0133] First, the solution can be flexibly adapted to different campus scenarios, including classrooms of different sizes and public areas with different layouts; second, it improves the anomaly identification rate in the initial screening and the accuracy of individual location in the fine screening while controlling computing power costs; third, it avoids data redundancy caused by high-definition collection at all times, ensures the seamless nature of the screening process, and reduces student resistance.

[0134] One implementation method is to use a primary and secondary screening scheme that combines low-resolution multi-frames with dynamic high-definition interval adjustment. This scheme is mainly suitable for various school classroom scenarios.

[0135] In the low-resolution multi-frame acquisition and group interaction motion analysis phase, the initial resolution of the acquisition equipment in the school classrooms was set to no higher than 480P. The specific resolution could be adjusted according to the classroom size: 360P for small classrooms less than 40 square meters, and 480P for medium to large classrooms between 40 and 60 square meters. The frame rate was maintained at 20 to 25 fps to ensure both controllable data volume and continuity of motion capture.

[0136] Lightweight limb keypoint detection algorithms, such as the MobileNet-optimized OpenPose or BlazePose-Lite, are used to extract 10 to 15 core limb keypoints for each student in the group, including the head, torso, arms, and legs. The cumulative limb displacement of each student is calculated in 30-second intervals to reflect activity level, and the frequency of effective movements is also recorded; a displacement exceeding 5 centimeters is considered a valid movement. Each student's statistical results are compared with the average level of the same class in the same scenario; those below 50% of the average are marked as abnormal interaction movements.

[0137] During low-resolution multi-frame acquisition, high-resolution frames for analyzing facial expressions are also acquired at intervals. In the initial stage of high-resolution frame acquisition and facial expression feature extraction, high-resolution frames are acquired at fixed time intervals, typically 5 minutes. Schools can adjust this to 3 to 5 minutes according to their own teaching pace. The frame with the most obvious changes in motion is selected from the low-resolution video stream to avoid static frames affecting feature extraction. The resolution of this frame is temporarily increased: 720P for small and medium-sized classrooms and 1080P for large lecture halls.

[0138] Based on high-definition frames, facial features of individual students are extracted, with a focus on analyzing the eyes and corners of the mouth. For the eyes, the degree of eyelid closure and blinking frequency are analyzed, while for the corners of the mouth, the direction and angle, and whether they are continuously pulled downwards, are analyzed. The extracted features are compared with preset normal facial expression ranges. The normal range for eyelid closure is 30% to 50%, and the normal range for the direction and angle of the corners of the mouth is 10° to 30°. Features exceeding these ranges are marked as abnormal facial expressions.

[0139] The high-definition frame interval is dynamically adjusted based on the initial psychological risk level. Students marked as abnormal undergo a preliminary psychological risk level re-screening every 15 to 20 minutes, and schools can customize the re-screening cycle. The initial psychological risk level is divided into low risk, medium risk, and high risk. Low risk is characterized by a single abnormal feature, medium risk by two abnormal features, and high risk by three or more abnormal features. If the re-screening level rises from low risk to medium risk, the high-definition frame acquisition interval is shortened to 2 minutes; if it rises from medium risk to high risk, it is shortened to 30 seconds to 1 minute; if the level drops back during this period, the interval returns to the value corresponding to the previous level, achieving dynamic adaptation where higher risk corresponds to more intensive monitoring.

[0140] When a student's initial psychological risk level meets specific conditions, a precise analysis process is triggered. These conditions include two consecutive high-risk screenings, or medium risk lasting for one hour with three or more abnormal features detected in high-definition frames. Schools can adjust the thresholds according to their own risk management needs. Once the conditions are met, the adjustment of the high-definition frame interval stops, and a continuous high-definition acquisition and area focusing command is sent to the acquisition device. The command will focus on the student's seat area, rather than the entire classroom in high definition, and simultaneously initiate subsequent multimodal specific feature extraction and precise risk analysis.

[0141] Taking the aforementioned class as an example, the classroom area is 54 square meters and can accommodate 45 students. A single 2-megapixel camera is installed in the corner of the back wall, 3.5 meters above the ground, with a focal length of 4 millimeters, which can cover the entire classroom's field of vision.

[0142] When implementing the first approach in this scenario, the initial resolution of the acquisition device was set to 480P, with a frame rate of 25fps. One keyframe was extracted every 10 frames for analysis. Using a MobileNet-optimized version of OpenPose, 13 limb keypoints were extracted. Statistics showed that student A's cumulative limb displacement within 30 seconds was only 28 cm, while the class average was 145 cm. Student A's effective movement frequency was 1 time per 30 seconds, compared to the class average of 7 times per 30 seconds. Therefore, student A was marked as having abnormal interactive movements.

[0143] Initially, keyframes were selected at 5-minute / frame intervals, temporarily upscaled to 720P, and the average eyelid closure of student A was 68%, which is outside the normal range. The corner of the mouth was -8°, showing a downward pull, which was marked as abnormal facial features and initially determined to be low risk.

[0144] After a second screening 15 minutes later, student A expressed reluctance to learn when answering questions, and the new negative characteristics of the speech content raised the initial psychological risk level to medium risk, and the interval for obtaining high-definition frames was shortened to 2 minutes; after another 15 minutes of screening, student A's interactive actions, expressions and speech content were all abnormal, the level was raised to high risk, and the interval was further shortened to 30 seconds.

[0145] After two subsequent screenings, 15 minutes apart, Student A was classified as high-risk. At this point, the camera switched to 1080P for continuous recording and focused on the seat area in the 3rd row, 2nd column where Student A was located. This area occupied 70% of the field of view, and the precise analysis process was initiated.

[0146] Another implementation method is to use a sliding window area division combined with precise positioning after initial screening of the area for preliminary screening and fine screening. This method is mainly suitable for public areas of various schools.

[0147] In the low-cost initial data collection phase, the initial resolution of the acquisition equipment in the campus's public areas was set to 360P to 480P. 360P was selected for relatively narrow spaces such as corridors and libraries, while 480P was selected for open spaces such as playgrounds and canteens. A fixed-size sliding window was used to divide the low-resolution video stream into zones. The window size was selected based on the width of the public area: 512×384 pixels for narrow areas less than 3 meters wide, and 640×480 pixels for wide areas between 3 and 5 meters wide. The sliding frequency was set to 2 to 3 times per second to ensure comprehensive coverage of the public areas. Each sliding window corresponded to a fixed physical zone within the campus. Calibration was performed using the camera's focal length and installation height, with 1 pixel corresponding to 0.005 to 0.007 meters, ensuring accurate correspondence between the zone divisions and the actual physical space.

[0148] During the initial screening and analysis of large-scale behavior of multiple individuals in the area, a lightweight behavior recognition algorithm was adopted, specifically YOLO-Lite combined with an action classifier, to extract the large-scale limb movement features of multiple individuals within each sliding window. The large-scale limb movement features have clear quantitative standards: first, the limb displacement amplitude exceeds 1 / 3 of the sliding window height; second, the movement frequency exceeds 2 times per minute; and third, the abnormal posture lasts for more than 10 seconds, including curling up, prolonged stillness, etc.

[0149] The analysis results of each sliding window are recorded in real time. If two consecutive sliding windows detect at least one large-scale motion feature in the same area, the interval between consecutive windows is adjusted according to the sliding frequency. The interval is 0.5 seconds for two slides per second and 0.3 seconds for three slides per second. At this time, the area is marked as the area to be accurately detected.

[0150] Entering the precise individual data collection stage, the resolution of the collection devices corresponding to the areas to be precisely detected is temporarily increased. Areas such as corridors and libraries are increased to 720P, while areas such as playgrounds and canteens are increased to 1080P. At the same time, the camera focus is adjusted to make the areas to be precisely detected the core collection area. The core area occupies more than 80% of the field of view, while other areas maintain the initial low resolution to avoid wasting computing power caused by full-area high-definition collection.

[0151] The YOLOv5s algorithm is used to detect people in high-definition video streams. This algorithm can achieve real-time detection at 25 frames per second on an edge server. It distinguishes target individuals through multiple features, including clothing features, such as the color and style of school uniforms worn by students (the algorithm has been pre-trained to feature school uniform color features); body features, such as height and build, combined with the age characteristics of students on campus; and location features, such as the time the target individual stays in the area and the range of activity, which distinguishes them from other people who pass by briefly. Ultimately, it accurately locates the specific individual who made the abnormally large movements.

[0152] For each specific individual located, their facial image, body pose, and speech signal are extracted from the high-definition video stream and preprocessed. Face alignment uses the MTCNN algorithm to detect the individual's facial region, rotates the face to a horizontal pose, and scales it to a standard size of 224×224 pixels to eliminate the impact of pose and size differences on subsequent feature extraction. Speech denoising uses spectral subtraction to process the acquired speech signal, filtering out environmental noise such as footsteps, crowd noise, and equipment operation in public areas, improving the signal-to-noise ratio of the speech signal, typically from 15dB to over 30dB, meeting the requirements of subsequent speech feature extraction. After preprocessing, the process proceeds to subsequent specific feature extraction and temporal modeling steps.

[0153] Taking the school teaching building corridor in the above embodiment as an example, the corridor is 40 meters long and 2.5 meters wide, with 10 classroom doors on one side. A single 2-megapixel camera is installed on the ceiling in the middle of the corridor, 4 meters above the ground, with a focal length of 6 millimeters, which can cover the entire corridor's field of vision.

[0154] The initial resolution of the acquisition device was set to 360P, using a sliding window of 512×384 pixels. This window corresponds to an actual physical area of ​​the corridor of approximately 3 meters × 2.3 meters. Through camera parameter calibration, 1 pixel corresponds to an actual distance of 0.006 meters. The sliding frequency was set to 2 times per second. The entire corridor was divided into 8 fixed physical zones, each corresponding to the area of ​​2 classroom doors.

[0155] YOLO-Lite combined with an action classifier was used to extract large-scale action features. During the monitoring process, the 5th second corresponded to sliding window number 3, and the 5.5th second also corresponded to sliding window number 3. Both consecutive windows detected student B huddled in a corner of a classroom doorway, with the abnormal posture lasting for 12 seconds, meeting the marking criteria. Therefore, the area corresponding to sliding window number 3 was marked as the area to be precisely detected. Subsequently, the acquisition resolution of this area was temporarily increased to 720P, and the camera focus was adjusted so that this area occupied 85% of the field of view. YOLOv5s algorithm was used for detection. Based on student B's red school uniform jacket, height of approximately 130 cm, and position in a corner of a classroom doorway, student B was accurately located. After extracting student B's face image, body posture, and speech signal, the MTCNN algorithm was used for face alignment, and spectral subtraction was used for speech denoising. The processed speech signal-to-noise ratio was improved from 14dB to 32dB. The subsequent steps were specific feature extraction and temporal modeling.

[0156] Furthermore, to address the issue of insufficient identification of implicit psychological risks such as social avoidance during the initial screening process, a new action association determination function has been added for multi-person interaction scenarios. By identifying objects that are within the effective interaction range of the group but do not interact or interact infrequently, this implicit characteristic is included in the risk level determination dimension of the initial screening, further improving the accuracy of capturing psychological crisis signals such as social avoidance and loneliness tendencies, and avoiding missed judgments due to focusing only on a single action or facial expression feature.

[0157] Specifically, in classroom group interaction scenarios, a new three-step action association determination is added, combining the existing low-resolution multi-frame analysis:

[0158] Defining the effective interaction range: Based on the seating arrangement in the low-resolution video stream, with the center seat of the group as the origin, define an effective interaction range with a radius of 1.2-1.5 meters to ensure that it covers the space for all group members to pass items and turn their heads to talk normally.

[0159] Group interaction status confirmation: A lightweight body keypoint detection algorithm is used to determine whether the group is in a continuous interaction state. If there are no less than 3 instances of body passing actions between members within 5 minutes, such as passing stationery, pointing to a book, or turning their heads to talk, the group is considered an active interactive group.

[0160] Individual Interaction Frequency Statistics: For target students within the effective interaction range, the frequency of their proactive interactions within 5 minutes is counted. Proactive interaction actions include reaching out to pass items, turning one's head to a member and speaking, and responding to others' actions (such as nodding or hand gestures when speaking). A single complete action is counted as 1 interaction. If the frequency is less than 2 times (below 30% of the average interaction frequency in the same group), or if there are no proactive interaction actions throughout the entire process, even if the student has no other abnormal facial expressions or speech, it is marked as abnormal interaction behavior.

[0161] Meanwhile, in response to the scenario where students gather and communicate in classrooms during breaks without fixed seats, the original interaction range determination method cannot adapt to dynamic gathering areas. It accurately identifies students who are in the communication area around their seats but do not interact at all. This incorporates the implicit social avoidance characteristics in this scenario into the initial screening, further reducing the error in determining normal quietness and abnormal lack of interaction.

[0162] First, low-resolution video streams are used to identify the core groups interacting in the classroom, avoiding misjudging single individuals stopping or briefly passing by as conversational scenarios. The judgment criteria are divided into two steps:

[0163] Number threshold: If at least two students are present in a certain area of ​​the classroom at the same time, and the time spent there is not less than 30 seconds, since a single interaction during a break usually lasts more than 30 seconds, excluding short-term passing by, it is preliminarily determined that there is a potential interaction group in that area.

[0164] Interaction behavior confirmation: For the initially identified group, a lightweight limb keypoint detection algorithm is used. If more than one secondary head turning towards each other is detected within 30 seconds (i.e., the angle between the line connecting the head keypoint of one member and the head keypoint of another member is no greater than 30°, and the duration is no less than 2 seconds), or if there is limb interaction, then the group is officially identified as the core interactive group. The seating area of ​​this group becomes the benchmark for subsequent movement and interaction.

[0165] Based on the seating positions of the core interactive group and the characteristics of classroom seating spacing during breaks, a dynamic movement and interaction range is defined, with specific rules as follows:

[0166] Horizontal boundary: Starting from the left edge of the seat of the leftmost member in the core interactive group and ending at the right edge of the seat of the rightmost member, extend outward by 0.5 meters, covering the distance that students in adjacent seats can participate in the interaction. For example, if the core group is in seats 3-5 in row 4, the horizontal range is from the left edge of seat 2 in row 4 to the right edge of seat 6 in row 4.

[0167] Vertical boundary: Starting from the front edge of the seat of the frontmost member in the core interactive group and ending at the back edge of the seat of the last member, extend outward by 0.3 meters, covering the distance that students in adjacent seats in the front and back rows can turn sideways to participate in the communication. For example, if the core group is in the 4th row, the vertical range is from the back edge of the 3rd row to the front edge of the 5th row.

[0168] Range validity: If a student's seat is completely within the boundaries defined above in the horizontal and vertical directions, that is, the center coordinates of the seat are within the range coordinates, then the student is determined to be within the movement interaction range, regardless of whether the student gets up or not. As long as the seat is within the range, the student is included in the interaction behavior monitoring.

[0169] For students within the mobile interaction range, based on their own seats and without leaving their seats, their interaction behavior is monitored for 30 seconds. No interaction behavior is considered to be achieved if all of the following conditions are met:

[0170] No active interaction actions: No head turning towards the core group, no hand raising towards the core group, and no body leaning forward.

[0171] No passive interaction response: When a member of the core interaction group initiates an interaction with the student, there is no response action such as nodding, speaking, or receiving an object within 3 seconds.

[0172] Excluding normal exceptions: If a student is writing with their head down and the core group does not interact with them, it is considered normal quiet and not included in the non-interaction exception; if they only have their head down without writing, or if the core group initiates interaction but there is no response, it is still considered non-interactive behavior. If a student is resting while lying down, first record their location information, and then if they repeatedly lie down to rest during breaks without interacting with surrounding students, it is considered non-interactive behavior.

[0173] Furthermore, in response to the problems of static screening lacking dynamic verification and easily causing student resistance, a new closed-loop mechanism is optimized, which combines pre-screening anchoring, interactive generation, classroom embedding, and dynamic verification, taking into account the characteristics of students' classroom activities as the core scenario and their sensitivity to special issues.

[0174] The core technical logic is based on the potential crisis characteristics (such as social avoidance, low mood, and voice tremor) pre-screened by the original video / audio scheme. The technology module automatically generates interactive content that can be embedded into the normal teaching process. Utilizing the natural scenarios of teacher-student interaction in the classroom, it further collects the student's behavior, facial expressions, and voice data to achieve non-invasive status confirmation. This avoids requiring special attention for students and supplements the shortcomings of static screening through dynamic interaction, solving the problems of the difficulty in quantifying and defining psychological crises and the limited accuracy of single screenings.

[0175] Specifically, a dynamic verification supplementary module is added to step S100 above.

[0176] Personalized interactive task generation

[0177] Based on the potential abnormal characteristics corresponding to the preliminary psychological risk level of the students to be verified, and combined with the subject teaching plans of the remaining courses for the day extracted from the school's teaching management system, personalized interactive tasks are generated, which must meet three core constraints:

[0178] Feature-Task Matching Constraints: If the potential abnormal feature is social avoidance, the task should be designed as a small-group collaboration of 2-3 people to avoid increasing the social pressure on students in large groups of 4 or more; if it is low mood, the task should be set as a low-pressure expression, such as reading a single sentence or making a short statement of opinion, to avoid complex and long expressions; if it is anxiety tendency, the difficulty of the task should be lower than the student's usual classroom performance level. For example, if the student can usually complete medium-difficulty exercises, the task should only be arranged to answer basic questions or repeat formulas.

[0179] Adaptation Requirements for Teaching Objectives: The task format must be fully integrated into the normal teaching process of the corresponding subject and cannot be designed as an independent task separate from the course content. For example, Chinese courses can incorporate group paragraph analysis and role-playing reading of passages; math courses can incorporate peer review of exercises and sharing of simple problem-solving strategies; and English courses can incorporate two-person dialogue practice and word dictation assistance. The task should appear consistent with other regular teaching activities involving other students, without special markings or additional procedures.

[0180] Task duration and frequency limits: The duration of a single interactive task should be controlled within 3-5 minutes, matching the regular duration of classroom group activities and recess interactions; each student to be verified should only have one interactive task embedded per day to avoid high-frequency tasks causing students to lose interest or become resistant; the task should include at least one monitoring node for active expression or physical cooperation to ensure that effective data for verifying psychological state can be collected.

[0181] Task synchronization and classroom embedding

[0182] Task synchronization restrictions: Anonymized task information is synchronized to the teachers of the remaining courses of the day through a lightweight mini-program on the teacher's side. The synchronized content only includes the specific form of the task, the key verification points that need to be focused on, and suggestions for guiding remarks. It does not mention the students to be verified, psychological risks, or other related statements. The synchronized information is automatically deleted after the task is completed to avoid privacy leaks.

[0183] Limited timing for classroom integration: Instructors must integrate tasks into the regular teaching process of the corresponding subject, such as group discussions in Chinese, problem-solving exchanges in mathematics, and collaborative learning in English. The integration time should be naturally connected with the class progress, without pausing the teaching process or setting up special segments. If students show resistance when the task is started, the teacher may only gently guide them once, using a gentle tone, such as "It's okay, just try it with your partner." Do not force or repeatedly urge them to avoid causing emotional fluctuations in students.

[0184] Data collection device adjustment restrictions: One minute before the task starts, the teacher clicks "Start Data Collection" via the mini-program, triggering the classroom information collection device to switch to directional focusing mode. Only the collection resolution of the area where the student to be verified is located is increased to 720P, while other areas maintain a low resolution of 480P. At the same time, the microphone's directional sound pickup function is activated to filter out environmental noise from other groups and prioritize the collection of the voice signals of the student to be verified and their partner, ensuring accurate data collection without increasing the overall computing power consumption.

[0185] Interactive multimodal data acquisition

[0186] The data collection device needs to continuously collect three types of multimodal data from the students to be verified throughout the entire interactive task (3-5 minutes), and the collection process must meet the following limitations:

[0187] Data collection of body coordination movements: Using a lightweight body key point detection algorithm, record whether the student to be verified has three types of movements: actively turning to the partner, actively coordinating body movements, and relaxing body posture. The frequency and duration of each type of movement need to be recorded.

[0188] Facial expression change data acquisition: One high-definition key frame is extracted from the acquisition stream every minute. Facial key point detection is used to analyze three types of facial expression features: eyes, corners of the mouth, and facial muscles. The results are compared with the abnormal facial expression data during the initial screening.

[0189] Voice expression status data collection: The voice signals of the students to be verified are collected through the directional sound recording function, and three types of non-semantic features are extracted: voice tone, speech rate and number of pauses. At the same time, the duration of active speaking and whether positive semantic expression occurs are recorded.

[0190] Interactive data verification analysis and risk level adjustment

[0191] The collected interactive multimodal data were subjected to multidimensional validation analysis, and the preliminary psychological risk level of the students to be validated was adjusted based on the analysis results, with the following specific limitations:

[0192] The verification analysis dimensions are limited to three core dimensions: improvement of abnormal features, active participation, and frequency of positive signals. Improvement of abnormal features requires comparison with the initial screening data. For social avoidance features, the frequency of proactive interaction actions must increase by at least 50%; for depressed mood features, the duration of positive expressions must be at least 1 minute; and for anxiety tendencies, the number of speech pauses must decrease by at least 40%. Active participation requires that the duration of proactive speaking accounts for at least 30% of the total task time, and there must be at least one instance of proactive physical coordination. The frequency of positive signals requires at least two instances of positive signals such as a raised corner of the mouth, an increased tone of voice, or proactively responding to a partner.

[0193] Risk level adjustment rules are as follows: If all three dimensions meet the criteria, the initial psychological risk level is reduced by 1 level; if any two dimensions meet the criteria, the initial psychological risk level remains unchanged; if only one dimension meets the criteria or none of the three dimensions meet the criteria, the initial psychological risk level is increased by 1 level; if the adjusted level is no risk or low risk and there are no abnormalities verified for two consecutive days, the S300 precise analysis will not be initiated for the time being, and the next teaching cycle will proceed with the regular initial screening and monitoring; if the adjusted level is still medium risk or above, the S300 precise analysis process will be triggered.

[0194] Comprehensive Example

[0195] Student A in Class 2 of Grade 7 has a preliminary psychological risk level of high risk, with potential abnormal characteristics including social avoidance and low mood. The remaining class that day is the second period of Chinese reading class in the afternoon, and the teaching plan is to have group discussion and role-playing reading of the excerpt "From the Hundred Herb Garden to the Three Flavors Study".

[0196] Personalized interactive task generation: Combining features and teaching plan, a task is generated for two students to read aloud in pairs, each with a different role. Student A is assigned the role of a young Lu Xun in "From the Hundred Herb Garden to the Three Flavors Study". The partner is the student who usually communicates with student A. The task is embedded in the group reading presentation by the Chinese teacher and is carried out simultaneously with other groups.

[0197] Task synchronization and classroom embedding: The Chinese teacher receives anonymized information through a mini-program and naturally guides the class 20 minutes into the reading lesson: Students are divided into groups by their deskmates and choose a passage from the reading to read aloud in roles. Deskmates A and B can try reading the passage by Little Lu Xun, which has very short lines. Then, the teacher clicks "start recording," and the classroom camera focuses on the area where the two students are seated, while the microphone picks up sound in a directional manner.

[0198] Interactive multimodal data collection: During the 3-minute task, student A was observed to turn to their deskmate once and point to the textbook dialogue once. The body posture angle recovered from 65° in the initial screening to 75°. Facial keyframes showed that the corners of the mouth turned up in the second minute, and the degree of eyelid closure decreased to 45%. The voice signal showed that the duration of active speaking was 1 minute and 20 seconds, the pitch was increased by 8Hz compared to the initial screening, the number of pauses decreased from 5 times / sentence to 2 times / sentence, and there were no negative semantic expressions.

[0199] Interactive data verification analysis and risk level adjustment: All three dimensions meet the standards. Student A's initial psychological risk level has been reduced from high risk to medium risk. S300 precise analysis will not be triggered for the time being. The next day's math class plan will embed a peer-to-peer exercise check task for continuous verification.

[0200] This invention is not limited to the optional embodiments described above, and anyone can derive other various forms of products based on the inspiration of this invention. The specific embodiments described above should not be construed as limiting the scope of protection of this invention; the scope of protection of this invention should be determined by the claims, and the specification can be used to interpret the claims.

Claims

1. An information processing method for mental crisis screening, data collection and analysis for campus objects, characterized in that, The method comprises the following steps: S100, obtaining group behavior characteristic information of students by an information collection device arranged in a campus, wherein the group behavior characteristic information comprises interaction action characteristics, regional gathering characteristics and activity degree characteristics of a student group; The group behavior characteristic information is analyzed based on a video stream output by the campus information collection device, and the specific processing mode comprises: First, the interaction action characteristics of the group are analyzed by using low-resolution multi-frame video streams, the resolution of the sampled low-resolution video stream is not higher than 480P, the interaction action characteristics are extracted by using a lightweight limb key point detection algorithm, the limb displacement and action frequency of each student in the group are extracted according to the obtained interaction action characteristics, and it is judged whether the interaction action is abnormal; meanwhile, in the initial state, a single frame image is selected from the low-resolution video stream at a fixed time interval of 5 minutes / frame and is upgraded to a high-definition resolution, which is used to extract the expression characteristics of a single student, and the preliminary psychological risk level is analyzed according to the interaction action characteristics and the expression characteristics; If the preliminary psychological risk level of a single student is found to be increased in the preliminary screening, the acquisition interval of the high-definition image is shortened to 2 minutes / frame; if the preliminary psychological risk level is further increased to a high risk, the acquisition interval of the high-definition image is shortened to 30 seconds / frame; when the preliminary psychological risk level of a single student exceeds a set threshold, the adjustment of the acquisition interval of the high-definition image is stopped, and the precise analysis process of step S300 is triggered; S200, determining the preliminary psychological risk level of a single object by preliminary screening analysis according to the obtained group behavior characteristic information; S300, when the preliminary psychological risk level of a single object exceeds a set threshold, individual characteristic information of the single object is acquired for precise analysis, multi-modal specific feature extraction and time sequence modeling are performed on the individual characteristic information, a behavior portrait of the object is constructed, then multi-modal specific features are fused by using a multi-time sequence Transformer structure, an attention mechanism is introduced to focus on an abnormal behavior segment, and the precise psychological risk level of the object is judged; Before starting the precise analysis, a classroom interaction embedded dynamic verification step is further included, which specifically comprises: According to the potential abnormal characteristics corresponding to the preliminary psychological risk level of the object, and combining the subject teaching plan of the remaining courses of the day, an individualized interaction task suitable for the teaching target is generated; The individualized interaction task is synchronized to a teacher, and the synchronization content only includes the specific form of the task, the verification points that need to be focused on, and the suggestion of the guide language, and does not mention the student to be verified and the psychological risk. In the classroom teaching process, the interaction task is embedded in the regular teaching link of the corresponding subject course, and the embedding time is naturally connected with the classroom progress, without separately pausing the teaching process or setting a special link; Multi-modal data of the object in the interaction process is continuously collected by the campus information collection device, including limb coordination action, facial expression change and voice expression state during the interaction; The collected interaction multi-modal data is subjected to feature extraction, the preliminary psychological risk characteristic change of the object before and after the interaction is compared, and the preliminary psychological risk level of the object is adjusted. If the adjusted preliminary psychological risk level still exceeds the set threshold, the precise analysis of step S300 is triggered; if the adjusted preliminary psychological risk level is lower than the set threshold, the precise analysis is not started, and the steps S100-S200 and the dynamic verification steps are repeated in the next teaching cycle to continuously monitor the psychological risk state of the object; S400, when the precise psychological risk level exceeds the set threshold, the abnormal behavior area of the object is focused on through the heat map visualization model, an abnormal behavior report containing the abnormal behavior area, the psychological risk level and the determination basis is generated, and the behavior report is sent to the manager for reminding.

2. The information processing method for psychological crisis screening according to claim 1, characterized in that: The individual feature information obtained for a single object in step S300 includes facial image features, eye movement and gaze behavior features, body behavior features and voice features of the object; Among them, the facial image features are subtle expression information in high-resolution facial images, the eye movement and gaze behavior features are eye movement and gaze state information of the object, the body behavior features are limb movement amplitude, frequency and posture information of the object, and the voice features are non-semantic acoustic information and semantic information when the object speaks independently.

3. The information processing method for psychological crisis screening according to claim 1, characterized in that: The specific process of multi-modal specific feature extraction and time sequence modeling in step S300 includes: Facial micro-expression unit extraction, through OpenFace tool combined with facial micro-expression unit recognition model, 17 facial micro-expression unit intensity and frequency of the object are extracted; Then the eye movement and gaze behavior feature extraction is performed, the gaze duration, gaze avoidance times, blink frequency, eyelid closure time proportion and eye scanning speed of the object are extracted; Then the body behavior feature extraction is performed, the 2D or 3D body key points of the object are extracted through OpenPose algorithm or BlazePose algorithm, and the limb movement amplitude, movement frequency and movement asymmetry are calculated based on the body key points; Voice feature extraction, noise reduction processing is performed on the voice signal of the object, non-semantic acoustic features such as pitch, speech rate and pause times are extracted, and semantic analysis is performed on the speaking content to extract semantic features related to psychological crisis; The above extracted facial micro-expression units, eye movement and gaze behavior features, body behavior features and voice features are sliced according to a 5-second time window, the slice overlap ratio is 50%, and the feature time sequence of the object is constructed to form a behavior portrait.

4. The information processing method for psychological crisis screening according to claim 1, characterized in that: In the step S100, the group behavior feature information is obtained as follows: First, low-cost preliminary screening data acquisition is performed, the initial acquisition resolution of the acquisition equipment deployed in the campus is set to 360P-480P low resolution mode, a fixed size sliding window of 512×384 pixels or 640×480 pixels is used, and the low resolution video stream is divided into areas according to a sliding frequency of 2-3 times per second, each sliding window corresponds to a fixed physical area in the campus; Then, the area multi-person large-scale behavior preliminary screening analysis is performed, for multiple person targets in each sliding window, limb large-scale movement features are extracted through a lightweight behavior recognition algorithm, the large-scale movement features include limb displacement amplitude, movement frequency and abnormal posture duration; If the same area is detected to have at least one of the large motion features in two continuous sliding windows, the area is marked as a to-be-accurately-detected area; Then, accurate individual data collection is performed, the resolution of the collection device corresponding to the to-be-accurately-detected area is temporarily increased to 720P-1080P high-definition mode, and the area is focused on to collect a video stream; YOLO algorithm is used to detect and distinguish individuals in the high-definition video stream, and combined with the extracted large motion features, the specific individual producing the abnormal large motion in the area is matched and located; For the located specific individual, the face image, body posture and voice signal thereof are extracted from the high-definition video stream, face alignment and voice noise reduction preprocessing are performed, and subsequent specific feature extraction and time series modeling steps are entered.

5. The information processing method for mental crisis screening according to claim 1, characterized in that: The determination basis of the abnormal behavior report in step S400 includes: The specific type of the abnormal feature in the multi-modal specific feature, the duration of the abnormal feature, and the cooperative occurrence of the abnormal features of each modality; the abnormal behavior report also contains key audio and video clip screenshots in the preliminary screening and accurate analysis process of the object, and the manager receives the report, manually reviews based on the screenshots and the heat map, and after the review is confirmed, starts the psychological intervention process for the student.

6. An apparatus, characterized by: The information processing method for psychological crisis screening according to any one of claims 1-5, wherein the object information in the campus is acquired and analyzed to obtain an abnormal behavior report for feedback, includes a server and a data collection module, the data collection module acquires group behavior feature information and individual feature information and sends them to the server for processing and analysis, and the server generates an abnormal behavior report.

Citation Information

Patent Citations

  • Abnormal behavior detection method, system, device and medium

    CN119785416A

  • Rail transit video intelligent analysis method, medium and system

    CN120259946A

  • Campus risk early identification and restoration method and system based on digital, intelligent and teaching conjunctions

    CN120562870A

  • Analysis method and device based on multi-modal information, equipment and medium

    CN120746783A