Multi-modal data fusion college classroom behavior real-time identification and early warning system
The real-time recognition and early warning system for classroom behavior in universities, which integrates multimodal data, solves the problem of high recognition error rate in dynamic student movement scenarios and improves the accuracy of behavior recognition.
Patent Information
- Application Number
- CN202511098890.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies cannot adapt to dynamic student movement scenarios, leading to an increase in the error rate of classroom behavior recognition and reducing the accuracy of real-time behavior recognition.
A real-time classroom behavior recognition and early warning system for universities, which adopts multimodal data fusion, acquires and integrates visual, audio, and environmental data through time node and recognition area division modules, multimodal data acquisition and target behavior recognition modules, and behavior judgment and early warning modules, and then identifies and outputs early warning information.
It improves the accuracy of real-time behavior recognition, adapts to dynamic student movement scenarios, and reduces the recognition error rate.
Smart Images

Figure CN120977010A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of educational information technology, and in particular to a multi-modal data fusion real-time recognition and early warning system for classroom behavior in colleges and universities. BACKGROUND
[0002] Real-time behavior recognition and early warning in the classroom is achieved by real-time collection and analysis of student and teacher behavior data in the classroom, combined with preset rules or machine learning models, to quickly identify abnormal behavior (such as playing with mobile phones, leaving the seat without permission, shouting, etc.) and trigger an early warning to assist teachers in maintaining classroom order and improving teaching quality.
[0003] In the current real-time behavior recognition and early warning in the classroom, behavior recognition is achieved by presetting the camera coverage area, but it cannot adapt to the dynamic movement of students, resulting in an increase in recognition error rate and a decrease in the accuracy of real-time behavior recognition. SUMMARY
[0004] The present application provides a multi-modal data fusion real-time recognition and early warning system for classroom behavior in colleges and universities, which solves the problem of the prior art that cannot adapt to the dynamic movement of students, resulting in an increase in recognition error rate and a decrease in the accuracy of real-time behavior recognition.
[0005] To achieve the above-mentioned purpose, the present application adopts a multi-modal data fusion real-time recognition and early warning system for classroom behavior in colleges and universities, which comprises a time node and recognition area division module, a multi-modal data acquisition and target behavior recognition module, and a behavior judgment and early warning module. The time node and recognition area division module is used to divide the preset time nodes according to the classroom time axis, divide the classroom space into multiple dynamic recognition areas, and trigger the data collection request for each time node area. The multi-modal data acquisition and target behavior recognition module is used to acquire the multi-modal data of the current time node, extract the multi-modal features, generate a multi-modal fusion feature vector, and identify the target behavior in the area according to the multi-modal fusion feature vector. The behavior judgment and early warning module is used to judge the abnormal situation according to the target behavior and output the early warning information.
[0006] The time node and recognition area division module comprises a time axis node preset unit and a space dynamic partition unit. The time axis node preset unit is used to divide the time unit according to the course type and dynamically adjust the node interval with the teaching link as the boundary. The space dynamic partition unit is used to real-time recognize the student gathering area and dynamically divide the classroom into multiple recognition areas.
[0007] The time node and identification area division module further comprises a coordinate system establishment unit. The coordinate system establishment unit is configured to establish a coordinate system through the classroom camera layout.
[0008] The time node and identification area division module further comprises a data acquisition trigger unit. The data acquisition trigger unit is configured to send an acquisition instruction to the sensor of the identification area when the time enters the preset node, and synchronously record the video, audio and environmental data.
[0009] The multi-modal data acquisition and target behavior identification module comprises a multi-modal data acquisition unit and a behavior identification unit. The multi-modal data acquisition unit is configured to acquire multi-modal data of the current time node area visual data, audio data and environmental data, and extract multi-modal features. The behavior identification unit is configured to fuse the feature vectors and identify the target behavior.
[0010] The multi-modal data acquisition unit comprises a visual feature extraction subunit, an audio feature extraction subunit and an environmental feature extraction subunit.
[0011] The visual feature extraction subunit is configured to acquire the current time node area visual data, frame the students and teachers in the area, and extract spatiotemporal, posture and target attribute features. The audio feature extraction subunit is configured to acquire the current time node area audio data, eliminate silent segments, and extract acoustic, prosodic and semantic features. The environmental feature extraction subunit is configured to acquire the current time node area environmental data, and extract temperature and humidity, illumination and device state features.
[0012] The behavior judgment and warning module comprises a behavior threshold adjustment unit and an abnormality warning output unit. The behavior threshold adjustment unit is configured to establish a behavior label and scene mapping table, preset a behavior threshold, and dynamically adjust the threshold according to the teaching context. The abnormality warning output unit is configured to set multiple levels of warning, judge abnormality according to the target behavior, and output warning information.
[0013] The behavior judgment and warning module further comprises a behavior continuity verification unit.
[0014] The continuity verification unit is configured to verify the behavior continuity according to the state transition.
[0015] The application is a multi-modal data fusion college classroom behavior real-time identification and early warning system, the time node and identification area division module is used for dividing the preset time node according to the classroom time axis, dividing the classroom space into multiple dynamic identification areas, and triggering the area data collection request under each time node; the multi-modal data acquisition and target behavior identification module is used for acquiring the area multi-modal data of the current time node, extracting multi-modal features, generating a multi-modal fusion feature vector, and identifying the target behavior in the area according to the multi-modal fusion feature vector; the behavior judgment and early warning module is used for judging the abnormal situation according to the target behavior and outputting the early warning information; through the above-mentioned mode, the effect of adapting to the dynamic moving scene of students and improving the accuracy of behavior real-time identification is obtained. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only show some embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0017] Figure 1 It is a structural principle diagram of the multi-modal data fusion college classroom behavior real-time identification and early warning system of the application.
[0018] Figure 2 It is a step flow chart of the multi-modal data fusion college classroom behavior real-time identification and early warning method of the application.
[0019] Figure 3 It is a step flow chart of S100 of the application.
[0020] Figure 4 It is a step flow chart of S200 of the application.
[0021] Figure 5 It is a step flow chart of S300 of the application.
[0022] Figure 6 It is a structural principle diagram of the electronic device of the application.
[0023] 400-time node and identification area division module, 401-time axis node preset unit, 402-coordinate system establishment unit, 403-space dynamic partition unit, 404-data acquisition triggering unit, 500-multimodal data acquisition and target behavior identification module, 501-visual feature extraction subunit, 502-audio feature extraction subunit, 503-environmental feature extraction subunit, 504-behavior identification unit, 600-behavior judgment and early warning module, 601-behavior threshold adjustment unit, 602-behavior continuity verification unit, 603-exceptional early warning output unit. DETAILED DESCRIPTION
[0024] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, the same numbers are used to indicate the same or similar components. The embodiments described in the following exemplary embodiments are not meant to represent all embodiments consistent with the present application.
[0025] The terminology used in the present application is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in the description of the application and the appended claims, the singular forms “a,” “an,” and “the” are intended to include plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0026] It should be understood that although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a particular order or hierarchy. These terms are used only to distinguish one from another. For example, a first information can be termed a second information, and similarly, a second information can be termed a first information, without departing from the scope of the present application. Depending on the context, the word “if’ as used herein can be interpreted as meaning “when” or “in response to determining” or “in response to a determination.”
[0027] Referring to Figure 1 The present application provides a multi-modal data fusion real-time recognition and early warning system for college classroom behavior, comprising a time node and identification area division module, a multi-modal data acquisition and target behavior identification module, and a behavior judgment and early warning module. Among them: The time node and identification area division module is used to divide preset time nodes according to the classroom time axis, divide the classroom space into multiple dynamic identification areas, and trigger the request for data collection in each time node area; The multi-modal data acquisition and target behavior recognition module is configured to acquire multi-modal data of a region at a current time node, extract multi-modal features, generate a multi-modal fusion feature vector, and recognize a target behavior in the region based on the multi-modal fusion feature vector. The behavior judgment and early warning module is configured to judge an abnormal situation based on the target behavior and output early warning information.
[0028] In the embodiment, the time node and recognition region division module divides preset time nodes according to a classroom time axis, divides a classroom space into a plurality of dynamic recognition regions, and triggers a region data acquisition request at each time node. The multi-modal data acquisition and target behavior recognition module acquires multi-modal data of a region at a current time node, extracts multi-modal features, generates a multi-modal fusion feature vector, and recognizes a target behavior in the region based on the multi-modal fusion feature vector. The behavior judgment and early warning module judges an abnormal situation based on the target behavior and outputs early warning information. In this way, the student dynamic movement scenario is adapted, and the accuracy of real-time behavior recognition is improved.
[0029] Further, the time node and recognition region division module includes a time axis node preset unit and a space dynamic partition unit. The time axis node preset unit is configured to divide time units according to a course type and dynamically adjust node intervals with teaching links as boundaries. The space dynamic partition unit is configured to identify a student gathering region in real time and dynamically divide a classroom into a plurality of recognition regions.
[0030] Further, the time node and recognition region division module further includes a coordinate system establishment unit. The coordinate system establishment unit is configured to establish a coordinate system through a classroom camera layout.
[0031] Further, the time node and recognition region division module further includes a data acquisition triggering unit. The data acquisition triggering unit is configured to send an acquisition instruction to a sensor of a recognition region when a time enters a preset node, and synchronously record video, audio, and environmental data.
[0032] In the embodiment, the time axis node preset unit divides time units according to a course type and dynamically adjusts node intervals with teaching links as boundaries. The coordinate system establishment unit establishes a coordinate system through a classroom camera layout. The space dynamic partition unit identifies a student gathering region in real time and dynamically divides a classroom into a plurality of recognition regions. The data acquisition triggering unit sends an acquisition instruction to a sensor of a recognition region when a time enters a preset node, and synchronously records video, audio, and environmental data.
[0033] Time node division rules: Basic rule: equal division according to course length (e.g. a 45-minute course is divided into 5 9-minute nodes); Dynamic adjustment rule: course type adaptation: Theory class: "knowledge point explanation-examples demonstration-class practice" as the boundary, the node interval is set to 15 / 10 / 20 minutes; Experimental class: "experiment preparation-operation demonstration-group experiment" as the boundary, the node interval is set to 10 / 15 / 20 minutes; Historical data correction: based on the behavior distribution of the same course type in the past 30 days, adjust the node to the period of intensive behavior change (e.g. the "class practice" stage of the theory course is extended by 5 minutes).
[0034] Coordinate system establishment: Take the lower left corner of the front door of the classroom as the origin (0, 0), the horizontal direction as the X-axis (0-10m), and the vertical direction as the Y-axis (0-6m) to establish a plane rectangular coordinate system; Camera layout mapping: mark the field of view boundary of 4 top corner cameras (coverage range 3m×3m) and 1 central panoramic camera (coverage 6m×6m) on the coordinate system.
[0035] Dynamic region identification and student gathering detection: real-time identification of student positions through YOLOv8 target detection model, and calculation of density heat map.
[0036] Region division strategy: If the area S of the gathering area is greater than 2m 2 (e.g. group discussion), divide it into independent identification areas; If the gathering area is dispersed (e.g. students answering questions independently), divide it into fixed areas according to the coverage of the camera (e.g. "front row area" and "back row area").
[0037] Node synchronization mechanism: When the system clock enters the preset node (e.g. 10:00:00), send acquisition instructions to the sensors in the corresponding area; Multi-modal data synchronization: Video stream: get the last 5 seconds of video stream (with timestamp) from the camera; Audio stream: get the last 5 seconds of audio stream (16kHz sampling rate) from the microphone; Environmental data: get real-time values from temperature and humidity sensors, light sensors, and projector status interfaces.
[0038] Further, the multi-modal data acquisition and target behavior recognition module includes a multi-modal data acquisition unit and a behavior recognition unit; wherein: The multi-modal data acquisition unit is configured to acquire multi-modal data of regional visual data, audio data and environmental data at a current time node respectively, and extract multi-modal features. The behavior recognition unit is configured to fuse the feature vectors and recognize a target behavior.
[0039] Further, the multi-modal data acquisition unit comprises a visual feature extraction subunit, an audio feature extraction subunit and an environmental feature extraction subunit.
[0040] Further, the visual feature extraction subunit is configured to acquire regional visual data at a current time node, frame a student and a teacher within a region, and extract spatio-temporal, posture and target attribute features respectively. The audio feature extraction subunit is configured to acquire regional audio data at a current time node, eliminate silent segments, and extract acoustic, prosodic and semantic features respectively. The environmental feature extraction subunit is configured to acquire regional environmental data at a current time node, and extract temperature and humidity, illumination and device state features respectively.
[0041] In the embodiment, the visual feature extraction subunit acquires regional visual data at a current time node, frames a student and a teacher within a region, and extracts spatio-temporal, posture and target attribute features respectively; the audio feature extraction subunit acquires regional audio data at a current time node, eliminates silent segments, and extracts acoustic, prosodic and semantic features respectively; the environmental feature extraction subunit acquires regional environmental data at a current time node, and extracts temperature and humidity, illumination and device state features respectively; and the behavior recognition unit fuses the feature vectors and recognizes a target behavior.
[0042] The target framing: a HRNet-W32 model is used to detect a student / teacher position and generate a boundary box.
[0043] The spatio-temporal feature: a 3D-CNN is used to extract spatio-temporal dynamics (such as hand raising frequency and walking trajectory) of a 5-second video segment. The posture feature: an OpenPose is used to detect key points (such as head angle and arm angle), and a posture encoding vector (such as “hand raising” encoding [1, 0, 0] and “sitting posture” encoding [0, 1, 0]) is calculated. The target attribute: clothing color and whether to carry an article (such as a book or a mobile phone) are recognized.
[0044] Audio feature extraction: Preprocessing: silent segments (energy threshold -40dB) are eliminated, and valid speech segments are retained. Feature extraction: Acoustic feature: MFCC (13 dimensions) and zero-crossing rate (ZCR) are extracted. Rhythm features: Calculate speech rate (number of phonemes per second), pitch range (F0 max -F0 min ); Semantic features: Generate speech text through Wav2Vec2.0 model, extract keywords (such as "quiet" "problem").
[0045] Environment feature extraction: Temperature and humidity: Record current values (such as temperature 25℃, humidity 60%); Light: Get illumination value (such as 300 lux) through photosensitive sensor; Device status: Read projector switch status (ON / OFF), electronic whiteboard usage flag (1 / 0).
[0046] Feature fusion: Concatenate visual (128 dimensions), audio (30 dimensions), and environmental (5 dimensions) features into a 163-dimensional vector; Cross-modal interaction through Transformer encoder to generate fusion feature vector; Behavior classification: Use pre-trained ResNet-50 model to classify fusion feature vector and output behavior label (such as "raise hand to speak" "lower head to play mobile phone" "group discussion").
[0047] Further, the behavior judgment and warning module includes a behavior threshold adjustment unit and an abnormal warning output unit; wherein: The behavior threshold adjustment unit is used to establish a behavior label and scene mapping table, preset a behavior threshold, and dynamically adjust the threshold according to the teaching context; The abnormal warning output unit is used to set multiple levels of warning, judge abnormal conditions according to target behavior, and output warning information.
[0048] Further, the behavior judgment and warning module further includes a behavior continuity verification unit.
[0049] Further, the continuity verification unit is used to verify the continuity of behavior according to the state transition condition.
[0050] In this embodiment, the behavior threshold adjustment unit establishes a behavior label and scene mapping table, presets a behavior threshold, and dynamically adjusts the threshold according to the teaching context; the abnormal warning output unit sets multiple levels of warning, judges abnormal conditions according to target behavior, and outputs warning information; the continuity verification unit verifies the continuity of behavior according to the state transition condition.
[0051] Among them, the behavior label and scene mapping, the preset rule library, such as: Scenario (Area + Course Type): Front row area_Theory class; Normal behavior tags: Raise hand, Take notes; Abnormal behavior tags: Lower head and play mobile phone, Unauthorized leaving seat; Threshold: Lower head duration>8 seconds; Scenario (Area + Course Type): Experimental area_Lab class; Normal behavior tags: Operate instruments, Record data; Abnormal behavior tags: Shouting, Not wearing safety glasses; Threshold: Volume>75 dB.
[0052] Dynamic threshold adjustment: If the projector is on (theory class scenario), the lower head threshold is shortened from 8 seconds to 5 seconds (because you need to look at the screen to take notes); If the "experimental preparation" link is detected (judged by device status), the safety glasses wearing threshold is adjusted from "none" to "must wear".
[0053] Behavior continuity verification, state transition model: Define a behavior state machine (such as "sitting posture → raising hand → speaking → sitting posture"); If the "lower head" state is detected for 10 seconds and does not transition to "taking notes" or "looking at the screen", an abnormal warning is triggered.
[0054] Multi-level warning output, warning level division: Level: Level 1; Behavior type: Minor violation (such as lowering head and playing mobile phone); Response measure: Local sound and light prompt (camera red light flashing); Level: Level 2; Behavior type: Moderate interference (such as unauthorized leaving seat); Response measure: Push message to teacher terminal + record log; Level: Level 3; Behavior type: Serious conflict (such as group shouting); Response measure: Trigger broadcast alarm.
[0055] Corresponding to the foregoing embodiment of the multi-modal data fusion real-time identification and early warning system of the college classroom behavior, the application also provides an embodiment of a multi-modal data fusion real-time identification and early warning method of a college classroom behavior.
[0056] Figure 2 is a flowchart of a multi-modal data fusion real-time identification and early warning method of a college classroom behavior according to an example embodiment. Referring to Figures 2-5 , the method comprises the following steps: S100: According to the course time axis, divide the preset time nodes, divide the classroom space into multiple dynamic identification areas, and trigger the region data collection request under each time node; S101: According to the course type, divide the time unit, and dynamically adjust the node interval with the teaching link as the boundary; S102: Establish a coordinate system through the classroom camera layout; S103: Real-time identification of student gathering area, dynamic division of classroom into multiple identification areas; S104: When the time enters the preset node, send the collection instruction to the sensor of the identification area, and synchronously record the video, audio and environment data.
[0057] S200: Obtain the area multi-modal data of the current time node, extract the multi-modal features, generate the multi-modal fusion feature vector, and identify the target behavior in the area according to the multi-modal fusion feature vector; S201: Obtain the area visual data of the current time node, frame the students and teachers in the area, and extract the spatiotemporal, posture and target attribute features respectively; S202: Obtain the area audio data of the current time node, eliminate the mute segments, and extract the acoustic, prosodic and semantic features respectively; S203: Obtain the area environment data of the current time node, and extract the temperature and humidity, illumination and device state features respectively; S204: Fuse the feature vectors and identify the target behavior.
[0058] S300: Judge the abnormal situation according to the target behavior, and output the warning information.
[0059] S301: Establish a behavior label and scene mapping table, preset a behavior threshold, and dynamically adjust the threshold according to the teaching context; S302: Verify the behavior continuity according to the state transition; S303: Set multiple levels of warning, judge the abnormal situation according to the target behavior, and output the warning information.
[0060] In the embodiment, first, the preset time nodes are divided according to the classroom time axis, the classroom space is divided into multiple dynamic identification areas, and the area data collection request under each time node is triggered; then the area multi-modal data of the current time node is obtained, the multi-modal features are extracted, the multi-modal fusion feature vector is generated, and the target behavior in the area is identified according to the multi-modal fusion feature vector; finally, the abnormal situation is judged according to the target behavior, and the warning information is output; through the above-mentioned manner, the effect of adapting to the dynamic moving scene of students and improving the accuracy of behavior real-time identification is obtained.
[0061] Correspondingly, the application also provides an electronic device, which comprises one or more processors, a memory for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the multi-modal data fusion real-time identification and early warning system for classroom behavior in colleges and universities as described above. Figure 6As shown in the figure, it is a hardware structure diagram of any data processing capable device where the multi-modal data fusion based real-time college classroom behavior recognition and early warning method of the embodiment of the present application is located. In addition to the processor, the memory and the network interface shown in the figure, any data processing capable device where the apparatus in the embodiment is located can also include other hardware according to the actual functions of the data processing capable device, which will not be described here. Figure 6 In addition to the processor, the memory and the network interface shown in the figure, any data processing capable device where the apparatus in the embodiment is located can also include other hardware according to the actual functions of the data processing capable device, which will not be described here.
[0062] Correspondingly, the present application also provides a computer readable storage medium having computer instructions stored thereon, which are executed by a processor to implement the multi-modal data fusion based real-time college classroom behavior recognition and early warning system as described above. The computer readable storage medium can be an internal storage unit of any data processing capable device, such as a hard disk or a memory. The computer readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. Further, the computer readable storage medium can include both an internal storage unit of any data processing capable device and an external storage device. The computer readable storage medium is used to store the computer program and other programs and data required by the data processing capable device, and can also be used to temporarily store data that has been output or will be output.
[0063] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The application is intended to cover any variations, uses or adaptations of the application following, in general, the principles of the application and including such departures from the present disclosure as come within known or customary practice in the art to which the application pertains.
[0064] It should be understood that the present application is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present application.
Claims
1. A real-time recognition and early warning system for classroom behavior in higher education institutions based on multimodal data fusion, characterized in that: It includes a time node and identification area segmentation module, a multimodal data acquisition and target behavior recognition module, and a behavior judgment and early warning module; among which: The time node and identification area division module is used to divide the classroom space into multiple dynamic identification areas by dividing preset time nodes according to the classroom time axis, and trigger a data collection request for each area under each time node. The multimodal data acquisition and target behavior recognition module is used to acquire regional multimodal data at the current time node, extract multimodal features, generate a multimodal fusion feature vector, and identify target behavior in the region based on the multimodal fusion feature vector. The behavior judgment and early warning module is used to judge abnormal situations based on target behavior and output early warning information.
2. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 1, characterized in that, The time node and identification region division module includes a time axis node preset unit and a spatial dynamic partitioning unit; wherein: The timeline node preset unit is used to divide time units according to course type and dynamically adjust the node interval with teaching links as boundaries. The spatial dynamic partitioning unit is used to identify student gathering areas in real time and dynamically divide the classroom into multiple identification areas.
3. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 2, characterized in that, The time node and identification region division module also includes a coordinate system establishment unit; wherein: The coordinate system establishment unit is used to establish a coordinate system based on the layout of classroom cameras.
4. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 3, characterized in that, The time node and identification area division module also includes a data acquisition triggering unit; wherein: The data acquisition triggering unit is used to send an acquisition command to the sensor in the recognition area when the time reaches a preset node, and simultaneously record video, audio and environmental data.
5. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 1, characterized in that, The multimodal data acquisition and target behavior recognition module includes a multimodal data acquisition unit and a behavior recognition unit; wherein: The multimodal data acquisition unit is used to acquire multimodal data of visual data, audio data and environmental data of the current time node area respectively, and extract multimodal features; The behavior recognition unit is used to fuse feature vectors and identify target behaviors.
6. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 5, characterized in that, The multimodal data acquisition unit includes a visual feature extraction subunit, an audio feature extraction subunit, and an environmental feature extraction subunit.
7. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 6, characterized in that, The visual feature extraction subunit is used to acquire visual data of the current time node region, define the students and teachers within the region, and extract spatiotemporal, posture and target attribute features respectively. The audio feature extraction subunit is used to acquire the audio data of the current time node region, remove silent segments, and extract acoustic, prosodic, and semantic features respectively. The environmental feature extraction subunit is used to obtain regional environmental data at the current time point and extract temperature, humidity, light intensity, and equipment status features respectively.
8. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 1, characterized in that, The behavior judgment and early warning module includes a behavior threshold adjustment unit and an abnormal early warning output unit; wherein: The behavior threshold adjustment unit is used to establish a behavior label and scene mapping table, preset behavior thresholds, and dynamically adjust the thresholds according to the teaching context. The abnormal warning output unit is used to set multi-level warnings, judge abnormal situations based on target behavior, and output warning information.
9. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 8, characterized in that, The behavior judgment and early warning module also includes a behavior continuity verification unit.
10. The multimodal data fusion-based real-time recognition and early warning system for classroom behavior in universities as described in claim 9, characterized in that, The continuity verification unit is used to verify the continuity of behavior based on the state transition situation.