Two-stage student standing detection method based on multiple modes

Through the multimodal two-stage detection method, combined with neural network and multimodal model, the misjudgment problem of single image recognition in complex scenarios is solved, and the accuracy and stability of student standing detection are improved.

CN120340119APending Publication Date: 2025-07-18HANGZHOU CHINGAN TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510338397.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, student standing detection based on a single image recognition method is difficult to accurately judge the student standing state in complex scenarios, and is easily affected by environmental interference or misjudgment.

Method used

A two-stage detection method based on multimodality is adopted. First, students' locations are determined through neural networks and first-stage detection is performed. IOU matching is used to determine the correspondence between face and head frames, and the standing threshold is judged by the maximum offset. Students who meet the conditions enter the second-stage detection and use multimodal model for final classification.

Benefits of technology

It significantly improves the accuracy and stability of students' standing detection, reduces the misdetection caused by bent over and straightening the waist, makes full use of the advantages of easy detection of faces near the face, and improves the ability to distinguish detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340119A_ABST
    Figure CN120340119A_ABST
Patent Text Reader

Abstract

According to the multi-mode-based two-stage student standing-up detection method provided by the invention, the accuracy and stability of student standing-up state detection are improved through two-stage processing and analysis. The method comprises the following steps of: 1) preprocessing an image into a first resolution size, and zooming the image into a second resolution size; 2) the neural network detects the positions of all students in the image; 2, 1) clustering is carried out by taking the central point of the face frame or the head frame of each student as a clustering point; 2) calculating the maximum offset; (3) if the maximum offset is larger than the first-stage standing threshold value, it shows that the first-stage filtering condition is met; 3, 1) calculating the corresponding position of the face or head position of the student on the first resolution image in the second resolution image; 2) acquiring a student human body position on the first resolution image; and 3) digging out a student human body position frame, scaling the student human body position frame to a third resolution in equal proportion, inputting the student human body position frame into a multi-modal model, and classifying student postures by the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a two-stage student standing-up detection method based on multi-modalities. Background Art

[0002] Currently, student behavior monitoring plays an important role in the field of education, and methods based on multi-modal data processing have become a research hotspot.

[0003] In terms of student standing-up detection, existing technologies mainly adopt algorithms such as object detection, key point and classification detection based on deep learning, as shown in the Chinese patent with the application number 202310990752.6 and the name of a lightweight and efficient classroom student standing-up and sitting-down detection method and system. However, relying solely on a single image recognition method has great limitations. These methods are often difficult to accurately judge the standing-up state of students in complex scenarios and are easily affected by environmental interference or misjudgment. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned deficiencies existing in the prior art, and provide a two-stage student standing-up detection method based on multi-modalities, which improves the accuracy and stability of student standing-up state detection through two-stage processing and analysis.

[0005] The technical solution adopted by the present invention to solve the above problems is as follows: A two-stage student standing-up detection method based on multi-modalities, characterized by including the following steps: Step 1, determining the position of the student's head, including the following steps: 1), installing a camera in the front of the classroom; 2), the camera captures the classroom image, preprocesses the image to the size of the first resolution, and then scales it proportionally to the size of the second resolution as the input of the neural network; 3), the neural network detects the positions of all students in the input image and distinguishes whether the students have face frames. Step 2, the first-stage detection of the student standing-up state, including the following steps: 1), for students at different distances (near and far), clustering is respectively performed with the center point of each student's face frame or head frame as the clustering point; 2), obtaining the clustering center of each student, and calculating the maximum offset in the Y direction for the points of each student's clustering center; 3), setting a set offset as the first-stage standing-up threshold. If the maximum offset is greater than the first-stage standing-up threshold, it indicates that the first-stage filtering condition is satisfied, and the students meeting the conditions enter Step 3. The students not meeting the conditions are determined to have no obvious standing-up behavior and skip Step 3; Step 3, the second-stage detection of the student standing-up state, including the following steps: 1) Obtain the positions of students' faces or heads that meet the first-stage filtering conditions, and calculate the corresponding positions of the students' faces or heads in the second-resolution image on the first-resolution image; 2) Obtain the positions of students' bodies in the first-resolution image; 3) According to the obtained positions of students' bodies, cut out the student body position frames and scale them proportionally to the size of the third resolution, and input them into the multi-modal model; 4) Use the multi-modal model to classify the standing up and sitting down behaviors of students' postures.

[0006] In step 2) of the first step of the present invention, the image is first processed into a resolution of 1920*1080, and then two copies of the processed image are copied. One copy with the resolution unchanged is used for the cropping operation in step 3) of step 2), and the other copy is scaled proportionally to a resolution of 960*512 and used as the input of the neural network.

[0007] In step 3) of the first step of the present invention, when both the head and face frames can be detected, the IOU matching method is used to match the student's head frame and face frame, so as to determine whether they are the same student and determine their positions.

[0008] The IOU matching method described in the present invention is: calculate the ratio of the intersection area to the union area between the student's head frame and face frame. This ratio is the IOU value between the two frames. When the IOU value meets the requirements, it is considered that the face frame and head frame belong to the same student and determine their positions.

[0009] In step 3) of the first step of the present invention, when only the head frame can be detected, the position of the head frame is used as the student position.

[0010] In step 3) of the first step of the present invention, a flag bit mark indicating whether there is a face frame information is assigned to each student. For students with both a face frame and a head frame, the flag bit mark = 1, and for students without a face frame but with a head frame, the flag bit mark = 0.

[0011] In step 1) of the second step of the present invention, for students with both a face frame and a head frame, the center point of the student's face frame is used as the clustering information, and for students without a face frame but with a head frame, the center point of the student's head frame is used as the clustering information.

[0012] In step 3) of the third step of the present invention, the body frame is obtained by expanding the head or face frame outward, so as to obtain the position of the student's body in the first-resolution image.

[0013] The outward expansion method of the present invention is: calculate the width wid and height hei of the student's head or face frame, and perform the following calculations: wid = xmax – xmin; hei = ymax – ymin; xmin’ = xmin - wid * n1; xmax’ = xmax + wid * n1; ymin ’= ymin - hei * n2; ymax ’= ymax + hei * n3; Where (xmin, ymin) and (xmax, ymax) are the upper left and lower right coordinates of the student's head or face frame respectively. After calculation, the upper left coordinate (xmin’, ymin’) and the lower right coordinate (xmax’, ymax’) of the human body frame are obtained. The new positions of these coordinates are the obtained student human body positions.

[0014] The multi-modal model described in the present invention is the BLIP-2 model.

[0015] Compared with the prior art, the present invention has the following advantages and effects: 1. Through two-stage process analysis, the present invention significantly improves the accuracy and stability of student standing-up detection; 2. By making full use of the advantage that nearby faces are easy to detect, the present invention can effectively distinguish the standing-up actions of students, reducing the false detection of standing up caused by the behavior of bending down and then straightening up. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of the camera installation in the embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of the IOU matching principle in the embodiment of the present invention.

[0018] Figure 3 It is a demonstration diagram of the recognized standing-up and sitting-down behaviors in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] The following further describes the present invention in detail with reference to the drawings and through embodiments. The following embodiments are explanations of the present invention and the present invention is not limited to the following embodiments.

[0020] A two-stage student standing-up detection method based on multi-modal in the embodiment of the present invention includes the following steps: Step 1: Determine the position of the student's head, including the following steps: 1). Install a camera in the front of the classroom, and the position of the camera can be adjusted according to the actual scenario; in this embodiment, the height of the camera and the schematic diagram are as Figure 1 shown.

[0021] 2) The camera first captures classroom images through a CMOS sensor, preprocesses the images to the size of the first resolution, where the first resolution is 1920*1080. Subsequently, two copies of the processed images are made. One copy with unchanged resolution is used for the matte extraction operation in step 3)2), and the other copy is scaled proportionally to the size of the second resolution as the input to the head and face recognition neural network, where the second resolution is 960*512.

[0022] 3) Based on the input from the previous step, the head and face recognition neural network detects the positions of all students in the input image and differentiates whether the students have face frames.

[0023] The head and face recognition neural network in this embodiment has very powerful performance. For students in the front row near the classroom (within eight meters), both the face frames and head frames of the students can be detected simultaneously. However, for students in the back row far from the classroom (beyond eight meters), the detection of face frames is unstable, but the detection effect of head frames for back-row students is very good.

[0024] When both the head and face frames can be detected, the neural network in this embodiment uses the IOU matching method to match the head frames and face frames of students, so as to determine whether they are the same student and determine their positions. The IOU matching method is as follows: calculate the ratio of the intersection area to the union area between the student's head frame and face frame, and this ratio is the IOU value between the two frames; through the calculation of the IOU value, the matching degree between the head frame and face frame can be obtained. When the IOU value is greater than 0.8, the embodiment considers that the face frame and head frame belong to the same student and determines their positions.

[0025] When only the head frame can be detected, the position of the head frame is used as the student's position.

[0026] The present invention assigns a flag bit mark indicating whether there is a face frame information for each student. For students with both face frames and head frames, the flag bit mark = 1; for students without face frames but with head frames, the flag bit mark = 0. Therefore, for the clustering centers of student positions in step 2)1), for students with both face frames and head frames, the embodiment uses the center point of the student's face frame as the clustering information, and for students without face frames but with only head frames, the embodiment uses the center point of the student's head frame as the clustering information.

[0027] Step 2, the first-stage detection of the student standing-up state, includes the following steps: 1) After determining the positions of students through the head and face recognition neural network in step 1)3), for students at different distances (near and far), in this embodiment, the center points of the face frames or head frames of each student are used as clustering points for clustering respectively. Using the center of the face frame as the clustering center for nearby students can reduce the probability of false detection of standing up during the process of nearby students picking up things or lying on the table and then straightening their waists.

[0028] 2) The position of each student clustering point also changes as the corresponding student position in step 1) 3) changes. Record this clustering point until the student position is lost for a certain period of time before it is cleared. In this embodiment, based on the obtained student clustering centers in the previous step, calculate the maximum offset in the Y direction for each point of the student clustering centers.

[0029] 3) In this embodiment, set a set offset as the first-stage standing-up threshold. According to the calculated maximum offset in the previous step, calculate whether the first-stage filtering condition is met. If the maximum offset is greater than the first-stage standing-up threshold, it means that the first-stage filtering condition is met. The students who meet the condition enter the second stage, that is, step 3). The students who do not meet the condition are determined to have no obvious standing-up behavior and thus skip step 3) to save performance.

[0030] Step 3, second-stage detection of the student standing-up state, includes the following steps: 1) Obtain the position of the student's face or head that meets the first-stage filtering condition through step 1), and calculate the corresponding position on the first-resolution image of the position of the student's face or head in the second-resolution image; the second-resolution image in this embodiment is scaled proportionally from the first-resolution image in step 1) 2), so the position of the student's face or head in the second-resolution image is the same as the position in the first-resolution image.

[0031] 2) Through the position of the student's face or head in the first-resolution image calculated in the previous step, expand the head or face box to obtain the human body box, so as to obtain the position of the student's human body on the first-resolution image.

[0032] The human body box can be obtained by expanding the coordinates of the student's face or head position to use the BLIP-2 model for classification. The expansion method in this embodiment is to calculate the width wid and height hei of the student's head or face box, and then offset 1.4 * wid times to the left and right in the horizontal direction, and offset 0.25 * hei times and 3.55 * hei times upward and downward in the vertical direction respectively. The calculation of the embodiment is shown below, where (xmin, ymin) and (xmax, ymax) are the upper left and lower right coordinates of the student's head or face box respectively. After calculation, the upper left coordinate (xmin’, ymin’) and lower right coordinate (xmax’, ymax’) of the human body box are obtained, and the new position of this coordinate is the obtained position of the student's human body: wid =xmax – xmin; hei =ymax – ymin; xmin’ = xmin - wid * 1.4; xmax’ = xmax + wid * 1.4; ymin ’= ymin - hei * 0.25; ymax ’= ymax + hei * 3.55.

[0033] 3), According to the student's human body position obtained in the previous step, extract the student's human body position box and scale it proportionally to the size of the third resolution to meet the input requirements of the multi-modal model. The third resolution is 224*224; due to the limited computing resources of the front-end device, the image is scaled proportionally to the size of 224*224 to reduce the excessive consumption of resources during inference. In this embodiment, the multi-modal model is the BLIP-2 model.

[0034] 4), Use the ITM image-text matching in the pre-trained BLIP-2 model to classify the standing up and sitting down behaviors of the student's posture. ITM is mainly for understanding tasks such as image classification, image retrieval, and VQA. After the two-stage classification result prediction is completed, the process ends.

[0035] In addition, it should be noted that for the specific embodiments described in this specification, the shapes and names of their components can be different. The above content described in this specification is only an example of the structure of the present invention. Any equivalent changes or simple changes made according to the structure, features, and principles described in the inventive concept of the present invention are included in the protection scope of the present invention. Those skilled in the technical field to which the present invention belongs can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, as long as they do not deviate from the structure of the present invention or exceed the scope defined by this claim book, they should belong to the protection scope of the present invention.

Claims

1. A two-stage student standing-up detection method based on multi-modalities, characterized in that: It includes the following steps: Step 1: Determine the positions of students' heads, including the following steps: 1). Install a camera in the front of the classroom; 2). The camera captures classroom images, preprocesses the images to the size of the first resolution, and then scales them proportionally to the size of the second resolution as the input of the neural network; 3). The neural network detects the positions of all students in the input image and distinguishes whether the students have face frames; Step 2: First-stage detection of students' standing-up states, including the following steps: 1). For students at different distances (near and far), the center points of the face frames or head frames of each student are used as clustering points for clustering respectively; 2). Obtain the clustering center of each student, and calculate the maximum offset in the Y direction by calculating the points of the clustering center of each student; 3). Set a set offset as the first-stage standing-up threshold. If the maximum offset is greater than the first-stage standing-up threshold, it means that the first-stage filtering condition is met. The students who meet the conditions enter Step 3, and the students who do not meet the conditions are considered to have no obvious standing-up behavior and skip Step 3; Step 3: Second-stage detection of students' standing-up states, including the following steps: 1). Obtain the positions of students' faces or heads that meet the first-stage filtering conditions, and calculate the corresponding positions on the first-resolution image of the positions of students' faces or heads in the second-resolution image; 2). Obtain the positions of students' bodies on the first-resolution image; 3). According to the obtained positions of students' bodies, cut out the student body position frames and scale them proportionally to the size of the third resolution, and input them into the multi-modal model; 4). Use the multi-modal model to classify the standing-up and sitting-down behaviors of students' postures.

2. The multi-modal based two-stage student standing-up detection method according to claim 1, characterized in that: In Step 1 2), first process the image to the resolution size of 1920*1080, and then copy the processed image twice. One copy has the same resolution and is used for the cropping operation in Step 3 2), and the other copy is scaled proportionally to the resolution size of 960*512 as the input of the neural network.

3. The multi-modal based two-stage student standing-up detection method according to claim 1, wherein: In Step 1 3), when both the head frame and the face frame can be detected, the IOU matching method is used to match the head frame and the face frame of the student, so as to determine whether they are the same student and determine their positions.

4. The multi-modal based two-stage student standing-up detection method according to claim 3, wherein: The described IOU matching method is: calculate the ratio of the intersection area to the union area between the student head frame and the face frame. This ratio is the IOU value between the two frames. When the IOU value meets the requirements, it is considered that the face frame and the head frame belong to the same student and determine their positions.

5. The multi-modal based two-stage student standing-up detection method according to claim 1, characterized in that: In Step 1 3), when only the head frame can be detected, the position of the head frame is used as the student position.

6. The multi-modal based two-stage student standing-up detection method according to claim 1, characterized in that: In Step 1 3), a flag bit mark indicating whether there is face frame information is assigned to each student. For students with both face frames and head frames, the flag bit mark = 1, and for students without face frames but with head frames, the flag bit mark = 0.

7. The multi-modal based two-stage student standing-up detection method according to claim 1, characterized in that: In Step 2 1), for students with both face frames and head frames, the center point of the face frame of the student is used as the clustering information. For students without face frames but with head frames, the center point of the head frame of the student is used as the clustering information.

8. The multi-modal-based two-stage student standing-up detection method according to claim 1, wherein: In Step 3 3), the body frame is obtained by expanding the head or face frame outward, so as to obtain the positions of students' bodies on the first-resolution image.

9. The multi-modal based two-stage student standing-up detection method according to claim 8, characterized in that: The external expansion method is as follows: calculate the width wid and height hei of the student's head or face box, and perform the following calculations: wid = xmax – xmin; hei = ymax – ymin; xmin’ = xmin - wid * n1; xmax’ = xmax + wid * n1; ymin ’= ymin - hei * n2; ymax ’= ymax + hei * n3; Among them, (xmin, ymin) and (xmax, ymax) are the upper left corner and lower right corner coordinates of the student's head or face box respectively. After calculation, the upper left corner coordinates (xmin’, ymin’) and lower right corner coordinates (xmax’, ymax’) of the human body box are obtained. The new positions of these coordinates are the obtained student human body positions.

10. The two-stage student standing-up detection method based on multi-modal according to claim 1, wherein: The multi-modal model described above is the BLIP-2 model.

Citation Information

Patent Citations

  • Face detection tracking method, dome camera head rotation control method, and dome camera

    CN108304001A

  • Two-stage pedestrian searching method combining face and appearance

    CN109635686A

  • Method and system for detecting sitting-up and sitting-down actions of students in recording and broadcasting system

    CN112597800A

  • Student standing detection method based on head detection and picture sequence

    CN112686154A

  • Lightweight and efficient classroom student standing and sitting-down detection method and system

    CN116958204A