Behavior recognition and abnormity early warning method, device and system applied to terrace classroom

By using multimodal fusion of audio and video data and dual-channel detection, abnormal behavior in lecture halls can be identified, solving the problem of low accuracy in existing technologies and achieving high-precision and low-latency abnormal warnings.

CN121661408APending Publication Date: 2026-03-13CHENGDU AERONAUTIC POLYTECHNIC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing methods for identifying abnormal behavior are not very accurate in lecture halls, especially in situations with large scenes, dense crowds, severe obstruction, and complex and diverse behaviors.

Method used

A multimodal data recognition method using audio and video data is adopted, which combines a cross-modal behavior recognition module, dual-channel anomaly detection, and causal knowledge graph. By fusing the spatiotemporal features of audio modality and video modality, abnormal behavior can be identified and warned.

Benefits of technology

It improves the accuracy of abnormal behavior identification, solves the problem of identity-behavior correspondence caused by severe occlusion and changing perspectives, and realizes high-precision, low-latency, and strong privacy monitoring of abnormal student behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661408A_ABST
    Figure CN121661408A_ABST
Patent Text Reader

Abstract

The invention discloses a behavior identification and abnormity early warning method, device and system applied to a terrace classroom. The method comprises the following steps: acquiring audio data and video data; positioning the audio data, and performing association matching on the audio data and people in the video data to obtain video modal spatial-temporal features and audio modal embedding features; inputting the video modal spatial-temporal features and the audio modal embedded features into a cross-modal behavior recognition module to recognize behaviors of all people in the video data to obtain behavior data of all people in the video; inputting behavior data of all people in the recognition video into an anomaly detection channel to recognize abnormal behaviors of people; and outputting a behavior abnormal result prompt according to the behavior data and the causal knowledge graph in response to the fact that the output result of the abnormal detection channel is an abnormal behavior. The problem of identity-behavior correspondence caused by serious shielding and variable view angles in a large scene can be solved, and the accuracy of abnormal behavior recognition is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of abnormal behavior recognition technology, specifically relating to a behavior recognition and abnormal warning method applied to a lecture hall. Background Technology

[0002] With social development and technological progress, the demand for analyzing abnormal behavior is increasing. Existing methods for abnormal behavior analysis employ skeleton point analysis, such as the "Abnormal Behavior Analysis Method Based on Spatial Positioning Prior and Multimodal Information Fusion" (application number 202510006576.7). This method involves: acquiring image frames from a video stream; acquiring human skeleton information from the image frames; acquiring spatial positioning information of the human body from the image frames; generating image information encoding based on the image frames; generating skeleton feature encoding based on the human skeleton information and spatial positioning information; generating a fusion sequence based on the image information encoding and skeleton feature encoding; generating abnormal behavior analysis results through an abnormal behavior analysis network based on the fusion sequence; and generating abnormal behavior warning information based on the abnormal behavior analysis results.

[0003] However, in scenarios where lecture halls are the core teaching space, there are characteristics such as large scale, dense crowds, severe obstruction, and complex and diverse behaviors. While existing abnormal behavior recognition methods can identify actions, they suffer from poor accuracy in anomaly identification. For example, "passing notes" may be harmless in a lecture, but it would be considered cheating in an exam. Summary of the Invention

[0004] To address the problem of low accuracy in identifying abnormal behavior in lecture halls using existing methods, this invention proposes a method, device, and system for behavior recognition and abnormal warning in lecture halls.

[0005] The objective of this invention is achieved through the following technical solution: The first aspect of this invention discloses a method for behavior recognition and anomaly warning applied in a lecture hall, comprising the following steps: Obtain a dataset, which includes audio data from a classroom to be monitored and video data from at least two different angles in the classroom during the same time period as the audio data; The audio data is located and associated with people in the video data to obtain video modal spatiotemporal features and audio modal embedding features; The video modal spatiotemporal features and audio modal embedding features are input into the cross-modal behavior recognition module to identify the behavior of all people in the video data, thereby obtaining the behavior data of all people in the video; The behavior data of all people in the video is input into the anomaly detection channel to identify abnormal behavior. The anomaly detection channel includes a first anomaly detection channel and a second anomaly detection channel. The first anomaly detection channel is used to match the behavior of a person with an abnormal behavior template library to identify abnormal behavior. The abnormal behavior template library contains a first behavior group and the abnormal behavior corresponding to the first behavior group. The second anomaly detection channel is used to match the behavior of a person with a normal behavior template library when no abnormal behavior is matched by the first anomaly detection channel. If no normal behavior is matched, it is judged as abnormal behavior. The normal behavior template library includes a second behavior group and the normal behavior corresponding to the second behavior group. In response to the abnormal behavior output by the anomaly detection channel, an abnormal behavior result prompt is output based on the behavior data and the causal knowledge graph. The causal indication graph includes a third behavior group and the abnormal results corresponding to the third behavior group.

[0006] The second aspect of the present invention discloses a behavior recognition and anomaly warning device for use in a lecture hall, comprising a memory and a controller connected in sequence, wherein the memory stores a computer program, and the controller is used to read the computer program and execute the behavior recognition and anomaly warning method for use in a lecture hall as described in the first aspect.

[0007] The third aspect of the present invention discloses a behavior recognition and anomaly warning system for use in a lecture hall, comprising a behavior recognition and anomaly warning device for use in a lecture hall as described in the second aspect and a prompting terminal, wherein the prompting terminal is used to signal-connect with the behavior recognition and anomaly warning device for use in a lecture hall to display the anomaly result prompt.

[0008] The beneficial effects of this invention are: This invention achieves abnormal behavior recognition based on multimodal data of audio and video data, and realizes abnormal behavior recognition based on dual channels. It can solve the problem of identity-behavior correspondence caused by severe occlusion and changing perspectives in large scenes, and effectively improve the accuracy of abnormal behavior recognition. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0012] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0013] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.

[0014] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0015] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0016] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0017] The first aspect of this invention discloses a method for behavior recognition and anomaly warning applied in a lecture hall, the method comprising steps S1 to S5. It should be noted that the step identifiers in this solution are only for ease of explanation and do not constitute a limitation on the order of steps. The order of each step is based on its verbal description and the sequential connection of each signal.

[0018] Step S1: Obtain the dataset, which includes audio data from a classroom to be monitored and video data from at least two different angles in the classroom during the same time period as the audio data.

[0019] Multiple cameras are deployed in the lecture hall to acquire full-view video data and enable the identification of abnormal human behavior throughout the classroom. Wide-angle cameras are preferred.

[0020] Audio data can be collected by a deployed microphone array to obtain acoustic audio data such as conversations and snoring in the classroom.

[0021] For example, in a lecture hall, 4-6 wide-angle cameras are installed to acquire video data. The f-angle of the wide-angle cameras... v =25 fps, resolution no less than 1080p, unified NTP time synchronization.

[0022] Install a 16-element microphone array with a sampling rate of f. a =16 kHz, array geometry has been calibrated, and the microphone coordinate system and classroom coordinate system are in agreement.

[0023] Each camera frame is timestamped. v Each audio buffer is appended with a start and end timestamp [t] a ,t a +Δt]; Both are aligned with the NTP clock and allow ≤10 ms jitter.

[0024] Step S2: Locate the audio data and associate and match the audio data with the people in the video data to obtain video modal spatiotemporal features and audio modal embedding features.

[0025] Specifically, this step uses beamforming to achieve spatial localization of audio data, that is, by calculating the time difference of sound arrival at different microphones, the location of the sound in the classroom is determined; then, spatiotemporal alignment and the Hungarian algorithm are used to achieve precise association between audio data and people. The spatiotemporal alignment here includes Network Time Protocol (NTP) synchronization and cross-correlation fine-tuning.

[0026] The essence of beamforming is to align and superimpose signals, amplifying sound from the target direction and suppressing noise from other directions. The position and orientation of the microphone array are pre-calibrated, and combined with the three-dimensional structure of the classroom, the three-dimensional coordinates of the sound source in the classroom can be calculated.

[0027] At the same time, the wide-angle camera also detected the face or body position of all students. For example, the center coordinates of student number 3 are (2.1m, 1.5m).

[0028] The Hungarian algorithm is a classic optimization algorithm that is often used in scenarios such as multimodal acquisition, beamforming, and sensor fusion to achieve one-to-one matching between different sources such as radar beams, camera detection boxes, and lidar point clouds, in order to minimize the matching cost or maximize the matching similarity and find the matching scheme with the minimum total cost.

[0029] If no reliable match can be found, such as if the voice is coming from an empty seat, mark it as "not associated".

[0030] Step S3: Input the video modal spatiotemporal features and audio modal embedding features into the cross-modal behavior recognition module to identify the behavior of all people in the video data, and obtain the behavior data of all people in the video.

[0031] The cross-modal behavior recognition module is preferably a Transformer module. The Transformer module intelligently fuses audio and video data to form a comprehensive judgment of the behavior. A small-sized Transformer is preferred. It automatically assesses whether the current video is clear and whether the sound is clear, and then decides which to trust more. For example, in dim lighting, it relies more on the sound; in quiet conditions but with the person constantly looking down, it trusts the video more.

[0032] For example: The video data characterizes "his eyes are closed and his head is down," supporting the behavior as "sleeping." The audio data characterizes "mild snoring detected," supporting the behavior as "sleeping."

[0033] But if: The video data indicates that "he is looking down". The audio data indicates that "he is reading the text aloud". The video and audio data contradict each other; perhaps the person is just engrossed in reading.

[0034] Therefore, the cross-modal behavior recognition module needs to dynamically decide which behavior to trust more.

[0035] Specifically, the AQA (Adaptive Quality-aware Gating) module is used. It evaluates the occlusion rate, sharpness, signal-to-noise ratio, and VAD stability quality of the video in real time; it converts the quality difference into a gating coefficient α using the sigmoid function; it performs weighted output on pure visual features and cross-modal fusion features; and it improves the robustness of the system under harsh conditions such as low light, occlusion, and noise.

[0036] Specifically, for video data, features such as "head down angle and hand position" are extracted using a human posture model. These features are then encoded into a string of numbers, such as [0.8, 0.2, 0.9], where each number represents the quantized value of the head down angle, the quantized value of a certain feature of the hand position, etc.

[0037] For audio data, the pre-trained sound model Wav2Vec2.0 is used to convert the sound into another string of numbers, such as [0.1, 0.7, 0.3]. These numbers represent the comprehensive quantization of different features of the sound, such as pitch, volume, timbre, etc.

[0038] The system uses an attention mechanism for fusion, calculating internally how much each action in the video is correlated with each feature in the audio. For example, "looking down" and "snoring" are highly correlated and given high weights, while "looking down" and "reading aloud" are uncorrelated and given low weights. The final output is a weighted fused vector, representing the overall judgment.

[0039] The cross-modal Transformer submodule adopts a standard cross-attention mechanism, with its Query coming from the video feature encoder and its Key and Value coming from the audio feature encoder. This module is not original to this invention, but its application to multimodal behavior recognition in lecture halls and its collaboration with the Quality Assessment Gating (AQA) constitute an essential component of the overall technical solution.

[0040] AQA is an algorithm module proposed in this patent, not part of Transformer. We use Transformer to perform attention fusion of video and audio features based on content relevance; AQA receives the fusion output of Transformer and independently calculated video / audio quality scores, and generates gating coefficients through the sigmoid function to weight the pure visual features and fused features to suppress the interference of low-quality modalities on the final discrimination.

[0041] The identified behavioral data can include looking down, moving hands, looking up, and flipping hair.

[0042] Step S4: Input the behavior data of all people in the video into the anomaly detection channel to identify abnormal human behavior. The anomaly detection channel includes a first anomaly detection channel and a second anomaly detection channel. The first anomaly detection channel is used to match human behavior with an abnormal behavior template library to identify abnormal human behavior. The abnormal behavior template library contains a first behavior group and the abnormal behavior corresponding to the first behavior group. The second anomaly detection channel is used to match human behavior with a normal behavior template library when no abnormal behavior is matched by the first anomaly detection channel. When no normal behavior is matched, it is judged as abnormal behavior. The normal behavior template library includes a second behavior group and the normal behavior corresponding to the second behavior group.

[0043] Specifically, the first anomaly detection channel uses template matching based on Mahalanobis distance and utilizes the covariance matrix to model the distribution within anomaly classes, thereby improving the discrimination accuracy.

[0044] For example, a large number of abnormal behavior samples such as "playing on mobile phones" and "sleeping" are collected, and the average characteristics and range of variation of each type of abnormal behavior are calculated. For example, when playing on mobile phones, the average length of time spent looking down is 85%, and the average hand tremor is 0.6. The range of variation can be the covariance matrix.

[0045] While ordinary distance and Euclidean distance only consider numerical differences, Mahalanobis distance takes into account the distribution characteristics of the data itself, further improving the accuracy of recognition.

[0046] The second anomaly detection channel uses a generative model trained on normal data, such as a Variational Autoencoder (VAE), to detect novel anomalies by comparing reconstruction error with deviation from the latent space. The VAE learns the inherent patterns in normal behavior data and uses these patterns to determine the reasonableness of new data. Specifically, the VAE is first trained on a large amount of normal classroom data; normal behavior can be reconstructed well by the model, while abnormal behavior shows a large reconstruction error. If the reconstruction error exceeds a dynamic threshold, such as the average of recent normal errors plus two standard deviations, it is determined to be a novel anomaly.

[0047] This step prioritizes matching against the abnormal behavior template library of the first anomaly detection channel. If a match is found, it is used first. Simultaneously, the second anomaly detection channel is also used for anomaly judgment. If the second anomaly detection channel does not match a normal phase, it is more certain that it is an anomaly. The two work together to perform a detection, improving the accuracy of the detection.

[0048] For example, if a person's behavior in a video is identified as looking down, moving their hand, or having a light on, then the identified abnormal behavior is playing with a mobile phone.

[0049] Step S5: In response to the abnormal behavior output by the abnormal detection channel, output an abnormal behavior result prompt based on the behavior data and the causal knowledge graph. The causal indication graph includes a third behavior group and the abnormal results corresponding to the third behavior group.

[0050] Taking the abnormal behavior of looking down, manual hand movements, and bright light corresponding to mobile phone use as an example, the system automatically sorts and selects the most important key evidence based on the strength of evidence, such as information gain / likelihood ratio. It then determines whether an interpretation is triggered by posterior confidence, outputting a logically clear and data-supported judgment. An abnormal result prompt could be: "Student in seat number 3 looked down 82% of the time in the past 90 seconds, and continuous micro-movements of the hand and localized bright light on the desktop were detected, leading to a high suspicion of mobile phone use."

[0051] This step can be implemented using Gemma3n. Gemma3n is a small, open-source language model from Google, with parameters such as 2B or 4B. It is optimized to run on ordinary servers and has advantages such as small parameter size, high computational efficiency, and high quality of generated natural language. It can quickly and accurately generate clear and easy-to-understand natural language interpretations with the limited computing power of local edge devices, meeting the system's real-time and low-power requirements.

[0052] For example, if you are a teaching assistant, please generate a concise explanation based on the following information: Outlier hypothesis: The student was operating the mobile phone evidence: The percentage of time spent looking down in the past 90 seconds was 82%. Detected continuous minute hand movements A highlighted area appears on the desktop. Seat number: 3 Confidence level: 91% Please output a natural language explanation, not exceeding 60 characters.

[0053] Then, when inputting into a lightweight large model (such as Gemma3n), it will output: "The student in seat number 3 looked down more than 80% of the time in the past 90 seconds, and at the same time, slight hand movements and a bright spot on the desktop were detected, which strongly suggests that the student was using a mobile phone." This process is called template-guided natural language generation (Prompt-based NLG).

[0054] Using the above method, abnormal behavior recognition is achieved based on multimodal data of audio and video data, and abnormal behavior recognition is achieved based on dual channels. This can solve the problem of identity-behavior correspondence caused by severe occlusion and changing perspectives in large scenes, and effectively improve the accuracy of abnormal behavior recognition.

[0055] This solution organically integrates five major technologies: multimodal perception, cross-modal robust fusion, dual-channel anomaly detection, causal assertion interpretation, and lightweight large-scale model on the edge. It is specifically designed for the high-challenge scenario of lecture halls and achieves intelligent monitoring of abnormal student behavior with high precision, low latency, strong privacy, and interpretability.

[0056] The second aspect of this invention discloses a behavior recognition and anomaly warning device for use in a lecture hall, comprising a memory and a controller connected in sequence. The memory stores a computer program, and the controller reads the computer program to execute the behavior recognition and anomaly warning method for use in a lecture hall described in the first aspect. Specifically, the memory may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; the controller may not be limited to using a microcontroller of the STM32F105 series. Furthermore, the computer device may also include, but is not limited to, a power supply unit, a display screen, and other necessary components.

[0057] The third aspect of the present invention discloses a behavior recognition and anomaly warning system for use in a lecture hall, comprising a behavior recognition and anomaly warning device for use in a lecture hall as described in the second aspect and a prompting terminal, wherein the prompting terminal is used to signal-connect with the behavior recognition and anomaly warning device for use in a lecture hall to display the anomaly result prompt.

[0058] The operating principles of the apparatus and system disclosed in the second and third aspects of this invention are detailed in the first aspect and will not be repeated here.

[0059] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A method for behavior recognition and anomaly warning applied in lecture halls, characterized in that, Includes the following steps: Obtain a dataset, which includes audio data from a classroom to be monitored and video data from at least two different angles in the classroom during the same time period as the audio data; The audio data is located and associated with people in the video data to obtain video modal spatiotemporal features and audio modal embedding features; The video modal spatiotemporal features and audio modal embedding features are input into the cross-modal behavior recognition module to identify the behavior of all people in the video data, thereby obtaining the behavior data of all people in the video; The behavior data of all people in the video is input into the anomaly detection channel to identify abnormal behavior. The anomaly detection channel includes a first anomaly detection channel and a second anomaly detection channel. The first anomaly detection channel is used to match the behavior of a person with an abnormal behavior template library to identify abnormal behavior. The abnormal behavior template library contains a first behavior group and the abnormal behavior corresponding to the first behavior group. The second anomaly detection channel is used to match the behavior of a person with a normal behavior template library when no abnormal behavior is matched by the first anomaly detection channel. If no normal behavior is matched, it is judged as abnormal behavior. The normal behavior template library includes a second behavior group and the normal behavior corresponding to the second behavior group. In response to the abnormal behavior output by the anomaly detection channel, an abnormal behavior result prompt is output based on the behavior data and the causal knowledge graph. The causal indication graph includes a third behavior group and the abnormal results corresponding to the third behavior group.

2. The behavior recognition and anomaly warning method applied to a lecture hall according to claim 1, characterized in that, The step of locating the audio data and associating and matching the audio data with people in the video data is as follows: Beamforming is used to achieve spatial positioning of audio data; Spatiotemporal alignment and the Hungarian algorithm are used to achieve accurate association between audio data and people, resulting in video modal spatiotemporal features and audio modal embedding features.

3. The behavior recognition and anomaly warning method applied to a lecture hall according to claim 1, characterized in that, The cross-modal behavior recognition module is a Transformer module.

4. The behavior recognition and anomaly warning method applied to a lecture hall according to claim 1, characterized in that, The first anomaly detection channel uses template matching based on Mahalanobis distance and models the distribution within anomaly classes using the covariance matrix.

5. A behavior recognition and anomaly warning device for use in a lecture hall, characterized in that, The method includes a memory and a controller connected in sequence, wherein the memory stores a computer program, and the controller is used to read the computer program and execute the behavior recognition and anomaly warning method for a tiered classroom as described in any one of claims 1-4.

6. A behavior recognition and anomaly early warning system for use in lecture halls, characterized in that, The invention includes a behavior recognition and anomaly warning device and a prompting terminal for use in a lecture hall as described in claim 5, wherein the prompting terminal is used to connect to the behavior recognition and anomaly warning device for use in a lecture hall to display the anomaly result prompt.

Citation Information

Patent Citations

  • Abnormal behavior analysis method based on spatial positioning prior and multi-modal information fusion

    CN120071209A