Classroom behavior detection method based on multi-modal large model

By adopting a multimodal large model in classroom behavior detection, combined with InsightFace, Track Anything and GroundingDINO models, the problem of insufficient generalization and personalized analysis capabilities of the detection system in the existing technology is solved, and efficient and accurate identification and statistics of classroom behavior is achieved.

CN120183048AActive Publication Date: 2025-06-20厦门工学院
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510645941.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing classroom behavior detection methods are difficult to achieve scalability and generalization when environmental changes or target types are updated, and personalized statistical analysis of student behavior is impossible.

Method used

The classroom behavior detection method based on multimodal large model is adopted, face recognition is performed through the InsightFace model, target tracking is performed by the Track Anything model, behavior detection is performed by the GroundingDINO model, and combined with the dynamic template update strategy, accurate identification and statistics of individual behavior is achieved.

Benefits of technology

It improves the generalization of the detection system, realizes the detection of unknown categories of behaviors, reduces the needs of data collection, labeling and model training, and can conduct accurate statistical analysis of individual behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183048A_ABST
    Figure CN120183048A_ABST
Patent Text Reader

Abstract

The invention discloses a classroom behavior detection method based on a multi-mode large model, and the method specifically comprises the steps: enabling a video file shot by a camera to serve as input data, enabling the video file to be directly inputted into GrondingDINO for target detection, or enabling the video file to be subjected to face recognition firstly, employing a Sub-Center ArcFace model in face recognition Insight Face, and enabling the video file to be subjected to face recognition; after the face of a specified object is recognized, a face area image or a target frame can serve as prompt information to be sent to a TrackAnything model for video target tracking, that is, a target area of the object is found in each frame of a video, then the target area is independently sent to Ground DINO for target detection, and whether the target is a behavior target to be detected or not is judged. According to the method, the behavior detection result of each person can be obtained, and then behavior statistical analysis of individuals or groups is carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and mainly relates to a classroom behavior detection method based on a multimodal large model. Background Art

[0002] Most of the previous classroom behavior detections were based on traditional closed-set object detection methods, mainly the YOLO series. A large amount of classroom data needed to be collected for object detection task annotation, and the annotation data was used to train the model. This method could achieve good average detection accuracy in a fixed environment. Once the environment changed or the target type was updated, the scalability and generalization of this method would be severely challenged. Usually, it was necessary to re-expand the dataset and retrain to meet the new application requirements.

[0003] The current classroom behavior detection does not implement the statistical analysis of the behaviors of each student and teacher. It only relies on the object detection results to summarize different types of behaviors. For example, how many people are listening to the class and how many people are playing with their mobile phones currently, and it cannot give personalized statistical results such as the concentration level of a certain student in this class.

[0004] The existing methods have insufficient coordination between multi-object tracking and behavior recognition, resulting in insufficient granularity of behavior analysis results and low recognition accuracy. They can only be limited to the detection of a limited number of fixed types of behaviors and lack generalization ability. Data collection, annotation, and model training are required, and it is difficult to meet the needs of continuous tracking and analysis of specific individual behaviors. Therefore, there is an urgent need for an efficient behavior analysis method that combines face recognition, video tracking, and object detection to achieve accurate recognition and statistics of individual behaviors. Summary of the Invention

[0005] In view of the above deficiencies, according to the first aspect of the present invention, a classroom behavior detection method based on a multimodal large model is proposed, and the specific steps are as follows: Obtain the original video through a camera, and read the original video frame by frame using a frame rate control strategy; In each frame, call the InsightFace model to recognize all faces, and compare the identity feature vectors of specific targets to obtain the target objects; Recognize the target faces in the target objects, extract the regions where the target faces are located, and use them as input prompts for the Track Anything model. The Track Anything model generates high-precision target masks based on the SAM framework of Meta AI; The Track Anything model automatically tracks the target regions in subsequent video frames according to the input prompts and outputs the bounding boxes of the targets in each frame; After individually cropping the target region images in each frame and inputting them into the GroundingDINO model, detections are made according to preset action keywords and the target detection results are output. A dynamic template update strategy is adopted. When the target is partially occluded or the viewing angle changes, the continuity of tracking is maintained through historical frame feature fusion; According to the target detection results, the action labels of each frame are stored in the classroom behavior record sequence, and the action statistical analysis of individual students or the classroom group is generated through the action analysis module.

[0006] Furthermore, the specific steps for calling the InsightFace model to recognize all faces include: Using the face detection module provided by the InsightFace model to quickly locate all face regions in the current frame image, and outputting the bounding box coordinates and key point information of each face; For each detected face region, calling the Sub-Center ArcFace model in the InsightFace model to extract feature vectors, and performing normalization operations on the feature vectors to generate face vectors representing identity features; Comparing the similarity of the face vector with the registered target identity feature vector. If the similarity is higher than the threshold, it is determined that the currently detected face is the target object.

[0007] Furthermore, the implementation process of the Track Anything specifically includes: initializing and refining using the single-frame strong segmentation ability of the SAM framework, processing frame-to-frame continuity using the temporal association ability of the XMem model, and complementing through an interactive correction mechanism.

[0008] Furthermore, the XMem model is responsible for time series tracking, constructing spatio-temporal associations based on the target mask, and quickly propagating target information in subsequent frames to solve the temporal problems of scale changes and motion blur.

[0009] Furthermore, the prompt information input into the SAM framework includes point prompts and box prompts; after receiving the point prompts and box prompts, the SAM framework generates a preliminary target mask, and uses post-processing techniques to optimize the preliminary target mask to obtain a high-precision target mask.

[0010] Furthermore, the specific steps for the Track Anything model to automatically track the target region in subsequent video frames according to the input prompts also include: the Track Anything model uses the high-precision target mask for target tracking, updates the target position of each frame and regenerates the mask, and uses optical flow estimation to predict the target motion during the tracking process.

[0011] Further, the GroundingDINO model includes: an image feature extraction module, a text feature extraction module, a feature enhancement module, a language-guided query selection module, and a cross-modal decoder module. After performing object detection on the target region image, it outputs a preliminary object detection result, maps the detection box back to the original image coordinate system, and filters out those with a confidence level higher than the set threshold to obtain the object detection result.

[0012] Further, the dynamic template update strategy specifically includes: dynamically maintaining and updating the template library according to the appearance changes of the target during the tracking process to avoid tracking errors caused by fixed templates.

[0013] According to the second aspect of the present invention, a computer program product is proposed, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above-mentioned method is implemented.

[0014] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects: (1) Combining language prompts to achieve behavior detection for unknown categories, improving the generalization of the detection system.

[0015] (2) The object detection module does not require data collection, annotation, and model training.

[0016] (3) Combining face recognition and object tracking to achieve individual behavior statistical analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the present invention. Other embodiments and many of the intended advantages of the embodiments will be readily apparent as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. Like reference numerals refer to corresponding similar components.

[0018] Figure 1 Shows a schematic structural diagram of a classroom behavior detection method based on a multimodal large model according to an embodiment of the present invention.

[0019] Figure 2 Shows a schematic diagram of Track Anything video object tracking according to an embodiment of the present invention.

[0020] Figure 3 Shows a schematic structural diagram of a computer system of an electronic device implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention and not to limit the invention. Additionally, it should be noted that for ease of description, only the parts related to the relevant invention are shown in the drawings.

[0022] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0023] Figure 1 The structural schematic diagram of a classroom behavior detection method based on a multimodal large model according to an embodiment of the present invention is shown, as Figure 1 shown: S1. Obtain the original video through a camera, and read the original video frame by frame using a frame rate control strategy; The monitoring video is collected in real time through a camera deployed in the target venue, or an imported saved original video file is used as the input source. The system performs frame-by-frame parsing and processing on the original video using a frame rate control strategy, specifically including: Sample the video stream according to the set frame rate (such as 5 frames per second, 10 frames per second, or adjusted as needed); number and timestamp each sampled frame image to ensure time continuity and synchronization for subsequent recognition and detection processing; use the frame-by-frame images as the input for subsequent face recognition, target tracking, and behavior detection modules.

[0024] S2. In each frame, call the InsightFace model to recognize all faces, and compare with the specific target identity feature vectors to obtain the target object; Use the face detection module provided by InsightFace (such as RetinaFace or YOLO-face) to quickly locate all face regions in the current frame image, and output the bounding box coordinates and key point information (such as eyes, nose tip, mouth corners, etc.) of each face.

[0025] For each detected face region, call the Sub-Center ArcFace model (pre-trained face recognition model) in InsightFace to extract feature vectors.

[0026] The model structure is based on ResNet or MobileFaceNet, and outputs a normalized face vector (generally 512-dimensional), representing the identity features of the face.

[0027] Compare the similarity between each extracted face vector and the registered target identity vector. The similarity measurement method is cosine similarity, and a threshold T is set (such as T = 0.4 - 0.6); The cosine similarity method is used to calculate the similarity, and the calculation formula is as follows: , where, represents the normalized feature vector, represents the registered target identity feature vector.

[0028] If the similarity value is higher than this threshold, then this face is determined to be the target object.

[0029] For the face recognized as the target object, record its position information (i.e., bounding box coordinates) in the current frame and the corresponding frame number, as the initialization prompt information for the subsequent target tracking module.

[0030] S3. Identify the target face in the target object, extract the area where the target face is located, as the input prompt for the Track Anything model, and the Track Anything model generates a high-precision target mask based on the SAM framework of Meta AI; The high-precision target mask is passed into the XMem model. XMem estimates the rough mask of the current frame with reference to the memory representation of historical frames, generates the rough segmentation mask of the current frame, and judges the mask quality. If the mask quality meets the requirements, it directly enters the GroundingDINO model; otherwise, it performs fine segmentation. The fine segmentation needs to call SAM again, and uses the rough mask and prompt of the current frame to generate a high-precision target mask.

[0031] The SAM framework model needs prompt information to generate the target mask. The prompt information includes point prompts and box prompts. The point prompt is a set of two-dimensional coordinate points, representing the key positions of the target (such as the face center or feature points). The box prompt is an extended form of the initial box, used to assist in segmentation.

[0032] After receiving the point prompt and box prompt, the SAM model generates a preliminary target mask. The preliminary mask may contain noise or inaccurate boundaries. Use post-processing techniques (such as conditional random field CRF or morphological operations) to optimize the preliminary mask to obtain a high-precision mask.

[0033] Based on the above SAM framework model and XMem model, the specific division of labor is as follows: SAM is responsible for single-frame precise segmentation (such as initialization, refinement, correction), supports users to generate high-quality masks through weak prompts such as clicking, and solves the problem of detailed segmentation in complex structure or occlusion scenarios.

[0034] The XMem model is responsible for time series tracking, constructs spatio-temporal associations based on the initial mask, quickly propagates target information in subsequent frames, and handles temporal challenges such as scale changes and motion blur.

[0035] S4. The Track Anything model automatically tracks the target area in subsequent video frames based on the input prompt and outputs the bounding box of the target in each frame. As Figure 2 shown, a schematic diagram of Track Anything video object tracking is shown. In the video sequence, the Track Anything model uses the target mask for object tracking. For each frame z, the target position is updated to , and the mask is regenerated ; optical flow estimation is used to predict the target motion during the tracking process; the finally output high-precision target mask is the result of multi-frame optimization.

[0036] Optical flow estimation is a technique for estimating the motion of pixel points in an image sequence. By analyzing the change in pixel intensity between adjacent frames, the motion vector of the target is inferred. The specific implementation process is as follows: Assume that I(x, y, t) represents the brightness value of the image at position (x, y) and time t. The core of optical flow estimation is to solve the motion vector v(u, v) of the pixel point, where u represents the velocity component in the horizontal direction and v represents the velocity component in the vertical direction. According to the brightness constancy assumption, the model equation is: , where = represents the gradient in the x direction, = represents the gradient in the y direction, = represents the change rate of time t; the Lucas-Kanad method or the Horn-Schunck method is introduced for constrained solution, and finally the motion vector of each pixel is obtained, thereby predicting the motion trajectory of the target.

[0037] S5. After separately cropping the target area image in each frame, it is input into the GroundingDINO model. Detection is performed according to the preset behavior keywords and the target detection results are output. A dynamic template update strategy is adopted. When the target is partially occluded or the perspective changes, the continuity of tracking is maintained through historical frame feature fusion. The GroundingDINO model is an open-set object detection model based on multi-modal fusion. Its core ability lies in supporting object detection driven by natural language text prompts. It can accurately align the objects described by humans in natural language (such as "students wearing red tops" and "laptop computers on the podium") with the visual regions in images / videos, and output the position (bounding box), category, and confidence of the objects. Different from traditional closed-set detection models (which can only recognize fixed categories defined in the training data), the GroundingDINO model fuses visual features and text semantics through a cross-modal attention mechanism, has zero-shot detection ability, and can detect objects or behaviors that have not appeared in the training data (such as "raising hands to ask questions" and "looking down at mobile phones") without retraining for new object categories, making it suitable for scenarios where the detection range needs to be flexibly expanded.

[0038] The model has the following modules: an image feature extraction module, a text feature extraction module, a feature enhancement module, a language-guided query selection module, and a cross-modal decoder module. After performing object detection on the target region image, it outputs preliminary object detection results, maps the detection boxes back to the original image coordinate system, and filters out those with a confidence higher than the set threshold to obtain the object detection results.

[0039] The core features of the model include: 1. Text-visual alignment: Combining DINO's object detection framework with Grounding technology, it aligns text descriptions and visual features through contrastive learning to achieve "image search by text" detection. 2. Zero-shot generalization: Without retraining for specific categories, it can directly detect unknown objects based on the text keywords input by users, making it suitable for flexible scenarios (such as custom behavior keyword detection). 3. High-precision positioning: Inheriting the end-to-end detection ability of the DETR series, it outputs object bounding boxes (Bounding Box), and at the same time supports multi-modal input (text + image).

[0040] Furthermore, in the object tracking scenario, a "template" usually refers to the representation of the appearance features of the object (such as visual templates, feature vectors). The core of the dynamic template update strategy is: during the tracking process, adaptively update the object template according to the real-time appearance changes of the object and historical frame information to cope with challenges such as occlusion, perspective change, and illumination change. The specific mechanism is as follows: 1) Policy objective: Solve tracking drift: When the object is partially occluded or the perspective changes, relying only on the features of the current frame may lose the true appearance of the object (such as incomplete features caused by occlusion). By fusing historical frame templates, the integrity and continuity of the features are maintained.

[0041] Balance new and old information: Avoid template obsolescence (not updated for a long time, the object appearance has changed) or introducing noise (frequent updates, incorporating occlusion / wrong features).

[0042] 2) Core mechanism ① Template representation: Feature extraction: For each frame's cropped target region (ROI), use GroundingDINO or an additional feature extractor (such as ResNet, Transformer) to generate visual feature vectors (such as RGB features, semantic features) as the temporary template for the current frame.

[0043] Multimodal fusion: If combined with text prompts (such as action keywords), the template may contain a fused representation of text embedding features (such as BERT encoding) and visual features, enhancing semantic constraints.

[0044] ② Dynamic update rules: Weighted fusion: The historical template and the current frame template are fused by weights. Example formula: New template == α Current frame features + (1 - ) Historical template. Where represents the update coefficient (such as dynamically adjusted according to the tracking confidence: reduced when occluded to reduce the influence of current noisy features).

[0045] Sliding window update: Maintain a fixed-length queue of historical frames (such as the most recent 5 frames). Discard the oldest old template during each update and incorporate the current new template to avoid excessive accumulation of outdated information in the template.

[0046] Confidence-driven: When the tracking confidence (such as detection score, IOU matching degree) is below the threshold (indicating possible occlusion or loss), suspend updating the template and continue tracking relying on the historical reliable template to avoid contaminating the template library with incorrect features.

[0047] ③ Occlusion and perspective change handling Partial occlusion: When the target features in the current frame are incomplete, supplement with the features of the unoccluded part in the historical template (such as per-channel weighting of features to retain the features of the clear regions in history).

[0048] Perspective change: Fuse the target features from different perspectives (such as front, side, back) to construct a multi-perspective robust template (similar to a "dictionary" of the target appearance), and retrieve the sub-template closest to the current perspective during matching.

[0049] ④ Anti-drift mechanism Template verification: Before each update, use the current template to perform reverse verification in adjacent frames (such as predicting the previous frame with the new template to check consistency), excluding incorrect templates introduced by misdetection.

[0050] ​Background suppression: Explicitly encode the background features around the target in the template, distinguish the target from similar distractors through contrastive learning (such as metric learning), and reduce the risk of drift.

[0051] 3) Policy advantages Robustness: Compensate for the defects of the current frame through historical frame information, and maintain tracking continuity in scenarios such as occlusion and blur.

[0052] Adaptability: Dynamically adjust the template update frequency and weight, balance "remembering history" and "adapting to new changes", and avoid the rigidity of static templates or the noise accumulation of dynamic templates.

[0053] Efficiency: Only perform feature extraction and template update on the cropped target region (ROI), without processing the entire image, reducing the computational cost.

[0054] S6. According to the target detection result, store the behavior label of each frame into the classroom behavior record sequence, and generate the behavior statistical analysis of individual students or the classroom group through the behavior analysis module.

[0055] In a specific embodiment, the input of the camera video stream data is processed in two paths. One path is for full-scene detection: directly input into the GroundingDINO model, and perform global target positioning in combination with text prompts (such as "detect all people using mobile phones"). The other path is for individual detection: first output the face target box through InsightFace face recognition, input the face target box into TrackAnything to track the individual region, input the individual region frame by frame into the GroundingDINO model, and perform local behavior judgment in combination with text prompts (such as "is student A listening attentively").

[0056] Global detection result: The positions and categories of all targets that meet the text prompts (such as the coordinates of all students "using mobile phones" in the classroom).

[0057] Individual detection result: The behavior label of each student / teacher frame by frame (such as "student B raises his hand at the 3rd minute and lowers his head at the 5th minute"), supporting the generation of analysis data such as individual concentration curves and behavior frequency statistics.

[0058] Next, refer to Figure 3 , which shows a schematic structural diagram of a computer system 300 of an electronic device suitable for implementing the embodiments of the present application. Figure 3 The shown electronic device is only an example and should not bring any limitations to the functions and usage ranges of the embodiments of the present application.

[0059] As Figure 3As shown, computer system 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the system 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0060] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.

[0061] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above-mentioned functions defined in the methods of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0062] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0063] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0064] The modules described in the embodiments of this application can be implemented in software or in hardware.

[0065] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain an original video through a camera, and perform frame-by-frame reading of the original video by adopting a frame rate control strategy; in each frame, call the InsightFace model to identify all faces, compare with a specific target identity feature vector to obtain a target object; identify the target face in the target object, extract the area where the target face is located as an input prompt for the Track Anything model, and the Track Anything model generates a high-precision target mask based on the SAM framework of MetaAI; the Track Anything model automatically tracks the target area in subsequent video frames according to the input prompt, and outputs the bounding box of the target in each frame; separately crop the target area image in each frame and input it into the GroundingDINO model, detect according to preset behavior keywords and output a target detection result, and adopt a dynamic template update strategy. When the target is partially occluded or the perspective changes, maintain the continuity of tracking through historical frame feature fusion; according to the target detection result, store the behavior label of each frame into the classroom behavior record sequence, and generate behavioral statistical analysis of individual students or the classroom group through a behavior analysis module.

[0066] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.

Claims

1. A classroom behavior detection method based on a multimodal large model, characterized in that: include: Acquire the original video through the camera, and read the original video frame by frame using a frame rate control strategy; In each frame, the InsightFace model is called to recognize all faces and compare the specific target identity feature vector to obtain the target object; Identify the target face in the target object, extract the area where the target face is located, and use it as an input prompt for the Track Anything model, which generates a high-precision target mask based on the SAM framework of Meta AI; The Track Anything model automatically tracks the target area in subsequent video frames according to the input prompts and outputs the bounding box of the target in each frame; The target area image in each frame is cropped separately and input into the GroundingDINO model. The target detection result is output based on the preset behavior keywords. The dynamic template update strategy is adopted. When the target is partially blocked or the viewing angle changes, the continuity of tracking is maintained by fusing the historical frame features. According to the target detection results, each frame behavior label is stored in the classroom behavior record sequence, and the behavior statistical analysis of individual students or classroom groups is generated through the behavior analysis module.

2. The classroom behavior detection method according to claim 1, characterized in that: The specific steps of calling the InsightFace model to recognize all faces include: The face detection module provided by the InsightFace model is used to quickly locate all face areas in the current frame image, and output the bounding box coordinates and key point information of each face; For each detected face area, calling the Sub-Center ArcFace model in the InsightFace model to extract feature vectors, normalizing the feature vectors, and generating face vectors representing identity features; The face vector is compared with the registered target identity feature vector for similarity. If the similarity is higher than a threshold, the currently detected face is determined to be the target object.

3. The classroom behavior detection method according to claim 1, characterized in that: The implementation process of Track Anything specifically includes: utilizing the single-frame strong segmentation capability of the SAM framework for initialization and refinement, utilizing the temporal association capability of the XMem model for processing inter-frame continuity, and complementing each other through an interactive correction mechanism.

4. The classroom behavior detection method according to claim 3, characterized in that: The XMem model is responsible for time series tracking, building spatiotemporal associations based on target masks, quickly propagating target information in subsequent frames, and solving the timing problems of scale changes and motion blur.

5. The classroom behavior detection method according to claim 1, characterized in that: The prompt information input by the SAM framework includes point prompts and frame prompts; after receiving the point prompts and frame prompts, the SAM framework generates a preliminary target mask, and uses post-processing technology to optimize the preliminary target mask to obtain a high-precision target mask.

6. The classroom behavior detection method according to claim 1, characterized in that: The Track Anything model automatically tracks the target area in subsequent video frames according to the input prompt. The specific steps also include: the Track Anything model uses the high-precision target mask to track the target, updates the target position of each frame and regenerates the mask, and uses optical flow estimation to predict the target motion during the tracking process.

7. The classroom behavior detection method according to claim 1, characterized in that: The GroundingDINO model includes: an image feature extraction module, a text feature extraction module, a feature enhancement module, a language-guided query selection module and a cross-modal decoder module. After performing target detection on the target area image, a preliminary target detection result is output, the detection frame is mapped back to the original image coordinate system, and the confidence level is higher than a set threshold is screened to obtain the target detection result.

8. The classroom behavior detection method according to claim 1, characterized in that: The dynamic template update strategy specifically includes: dynamically maintaining and updating the template library according to the appearance changes of the target during the tracking process to avoid tracking errors caused by fixed templates.

9. A computer program product, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

10. A computing system, characterized in that: The method comprises a processor and a memory, wherein the processor is configured to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Visual target tracking adaptive template fusion method

    CN111985375A

  • Single-target person tracking method integrating pedestrian re-identification and face detection

    CN112668483A

  • Image generation method and device, electronic equipment, storage medium and program product

    CN117315070A

  • College classroom student individual behavior tracking identification system and method

    CN117373118A

  • Intelligent supervision method based on large model and assembly line, electronic equipment and computer readable storage medium

    CN117953330A