Real-time examination room examinee abnormal behavior analysis method, system and device and medium

By real-time acquisition and frame extraction, and combining with the TPU hardware accelerated deep learning model, the problem of inability to identify candidates' abnormal behavior in the existing technology is solved, and efficient and accurate candidate behavior analysis and monitoring are achieved.

CN120340115APending Publication Date: 2025-07-18SHANDONG SAHNDA OUMASOFT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510313876.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing examination video surveillance system cannot judge the abnormal behavior of candidates in real time, and the computing resources and storage resources consume huge amounts, making it difficult to detect and deal with abnormal behaviors such as candidates' cheating during the examination process.

Method used

By collecting the examination room video stream in real time, selecting keyframes for object detection after frame extraction, combining the TPU hardware-accelerated deep learning model for abnormal behavior analysis, and performing manual secondary verification of the recognition results.

Benefits of technology

Real-time analysis and identification of candidates' abnormal behaviors is realized, the cost of calculation and manual review is reduced, the detection efficiency and accuracy are improved, and the real-time monitoring needs are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340115A_ABST
    Figure CN120340115A_ABST
Patent Text Reader

Abstract

The invention provides a real-time examination room examinee abnormal behavior analysis method, system and device and a medium, and belongs to the technical field of examination room examinee visual recognition. A camera is used for collecting RTSP video streams in an examination room in real time; performing frame extraction processing on the high-density video stream through a preset frame interval; selecting a key frame from the continuous video frames after frame extraction to perform target detection so as to accurately position the spatial position of the examinee; performing examinee abnormal behavior analysis based on the video clip after frame extraction and the examinee position; drawing an abnormal behavior key frame and a video clip according to an abnormal behavior recognition result; manual secondary verification is carried out on an abnormal behavior recognition result, so that the recognition accuracy is ensured; the deep learning model reasoning process is accelerated by utilizing TPU hardware, and the real-time behavior analysis performance is optimized. According to the invention, behavior analysis can be carried out on examinees in an examination room in real time, and the efficiency and accuracy of abnormal behavior identification can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual recognition of examinees in examination rooms, and particularly relates to a method, system, device and medium for analyzing abnormal behaviors of examinees in a real-time examination room. Background Art

[0002] Examinations are the fairest and most just way to select talents. However, with the continuous expansion of the scale of examinations, abnormal behaviors such as cheating by examinees have become potential risks that undermine the fairness and justice of examinations. In order to effectively curb such behaviors, intelligent video surveillance technology has been widely applied in standardized examination rooms to assist invigilators in conducting effective invigilation.

[0003] Existing examination video surveillance systems collect videos through cameras and can only record and save the videos, unable to judge possible abnormal problems during the examination process. It can only be left to the monitoring personnel to search for abnormal behaviors by playing back the recorded videos after or during the examination, and cannot monitor abnormal behaviors in real time during the examination process.

[0004] In the prior art, some will process the collected videos. However, when faced with a large amount of examination room video data, it will consume huge computing resources and storage resources, increasing the data processing burden and causing the examination video processing speed to be too slow or even stuck.

[0005] In a complex examination room video scenario, the position information and action behaviors of examinees are the basis for analyzing abnormal behaviors. Existing technical methods are difficult to accurately find the positions of examinees from a large number of video frames. Due to the slow calculation speed and large amount of calculation data, it is difficult to obtain the action behaviors of examinees in a timely manner and difficult to analyze the abnormal states of the examination. Summary of the Invention

[0006] The present invention provides a method for analyzing abnormal behaviors of examinees in a real-time examination room, which can analyze and feedback abnormal behaviors in the examination room in real time, reduce the labor cost of manually reviewing the monitoring to find abnormal behaviors, and effectively improve the accuracy and efficiency of detecting abnormal behaviors of examinees in the examination room.

[0007] The method includes: Step 1: Collect the video stream in the examination room in real time; Step 2: Perform frame extraction on the video stream at a preset frame interval; Step 3: Select key frames from the video frames after frame extraction for target detection to locate the positions of the examinees; Step 4: Analyze abnormal behaviors based on the video segments after frame extraction and the positions of the examinees; Step 5: According to the abnormal behavior recognition result, parse the key frames of the abnormal behavior and the original video segment; Step 6: Display the abnormal behavior recognition result information for manual secondary verification; Step 7: Use TPU hardware to accelerate the deep learning model inference process and optimize real-time behavior analysis performance.

[0008] It should be further explained that, in step 1, the video stream collected in the examination room is an RTSP video stream; Read video frames from RTSP video stream in real time and push them to the shared queue; Get video frames from the shared queue and process them.

[0009] Step 2 also includes: decoding the acquired RTSP video stream based on OpenCV, and the decoded original video frame is ,in t= 1,2,3 …M ; Set the frame extraction interval to N , the representative frame after frame extraction is Frame , the specific frame extraction method is as follows:

[0010] The video frame after extraction Frame Every T The frames are combined as a group to generate a fixed-length video frame segment; The original video corresponding to the preset frame video clip after frame extraction R to save.

[0011] It should be further explained that step 3 also includes: selecting an intermediate frame from the extracted video frames as a key frame for target detection, and the calculation method of the intermediate frame is:

[0012] Keyframe KeyFrame Input to the following YOLOx target detection model for inference:

[0013] in, Result Contains all the target information in the entire image. Result Includes: Target category Class , Confidence Score and coordinate position Box , coordinate position Box Includes: coordinates of the upper left corner of the target bounding box ) and the lower right corner coordinates ); The target detection results Result The targets of the candidates who are standing are filtered, and finally the targets of the candidates who are sitting are obtained; Store candidate goals in boxes_list list.

[0014] It should be further noted that step 4 further includes: after frame extraction in step 2 and object detection in step 3, a set of representative frames and the corresponding key frames KeyFrame of the object detection results are obtained; The consecutive representative frames and KeyFrame the object detection results boxes_list are input into the action recognition model Slowfast for action analysis; In the action analysis algorithm, the video segment is divided into two branches, namely the slow path and the fast path; Among them, the slow path further samples the video frames after frame extraction, while the input of the fast path is the T frame video after frame extraction. The formulas for the slow path and the fast path are as follows:

[0015]

[0016] The results output by the slow path and the fast path are concatenated based on the following formula to finally obtain the fused output; ) The actions predicted to be normal are filtered, and the abnormal action results are returned.

[0017] It should be further noted that step 5 further includes: according to the abnormal actions returned in step 4, rectangular boxes are used in the key frames KeyFrame to frame the positions where the abnormal candidates are located, and the types of abnormal actions are marked with text; The original video frames T corresponding to the representative frames of the abnormal action group R are saved; The key frames and the original video segments containing abnormal actions are also uniquely named based on timestamps, and the processing results are sent to the upper-layer server via an HTTP request.

[0018] It should be further noted that step 7 further includes: the used object detection model and action analysis model are trained under the GPU to obtain a file in pth format; The file in pth format is converted into the ONNX intermediate format; The network model in ONNX intermediate format is converted into bmodel format; The bmodel format is quantized into an INT8 model.

[0019] It should be further noted that step 3 further includes: extracting a preset number of pictures based on the key frames and performing annotation, marking the candidates as class 0 and the invigilators as class 1; Divide a preset number of pictures into a training set and a validation set according to a preset ratio; Based on the training set, perform multiple iterative trainings on the YOLOx network model to obtain the trained YOLOx network model; Input the preprocessed pictures into the YOLOx network model, perform downsampling in the backbone feature extraction network, compress the image size to a preset size, and at the same time expand the number of channels to a preset expansion amount; Perform feature fusion on the three effective feature layers extracted by the backbone feature extraction network, fuse the feature information of each different feature dimension, and then input the features after feature fusion into the YOLOx network model for regression prediction to finally obtain the target information in the examination room; The loss function formula of the YOLOx network model is:

[0020] Among them, represents the classification loss, represents the localization loss, represents the obj loss, The balance coefficient of the localization loss is set to 5.0, represents the number of Anchor Points classified as positive samples.

[0021] This application also provides a real-time analysis system for abnormal behaviors of examinees in the examination room. The system includes: A video acquisition module for real-time acquisition of the video stream in the examination room; A frame extraction processing module for performing frame extraction processing on the video stream at a preset frame interval; An examinee localization module for selecting key frames from the frame-extracted video frames for target detection to localize the positions of the examinees; A behavior analysis module for performing abnormal behavior analysis based on the frame-extracted video segments and the positions of the examinees; An abnormality analysis module for parsing the key frames of abnormal behaviors and the original video segments according to the abnormal behavior recognition results; A display verification module for displaying the abnormal behavior recognition result information for manual secondary verification; An analysis and reasoning module for using TPU hardware to accelerate the deep learning model reasoning process and optimize the real-time behavior analysis performance.

[0022] According to another embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the real-time analysis method for abnormal behaviors of examinees in the examination room are implemented.

[0023] According to another embodiment of the present application, a storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the real-time examination room candidate abnormal behavior analysis method are implemented.

[0024] As can be seen from the above technical solutions, the present invention has the following advantages: The real-time examination room candidate abnormal behavior analysis method provided by the present application can analyze the behaviors of candidates in the examination room. From video acquisition to abnormal behavior recognition, it can identify abnormal behaviors during the examination process. It can record the video information in the examination room as the data source for the entire abnormal behavior analysis.

[0025] Through frame extraction processing, the present application can reduce the amount of data to be processed without losing key information. Further, by reasonably selecting the frame interval, it can process the video data while ensuring that the behavior changes of candidates can be captured, thus accelerating the entire abnormal behavior analysis process. The present application also uses key frames, which contain the behavior change situations of candidates. Through target detection of the key frames, the positions of candidates can be located, and thus their behaviors can be analyzed. Compared with the existing target detection of video frames after frame extraction, only detecting the key frames can reduce the computational amount and improve the detection efficiency and accuracy at the same time.

[0026] By combining the position changes of candidates in the video segment and other image features, the present application can effectively analyze whether there are abnormal behaviors of candidates. According to different examination requirements, the judgment rules for abnormal behaviors are defined. It can adapt to the abnormal behavior analysis of various examination scenarios. By parsing the key frames and the original video segment, the process of the occurrence of abnormal behaviors can be shown. It can also restore the processes before and after the occurrence of behaviors through the original video segment.

[0027] By utilizing the parallel computing ability of the TPU hardware, the present application can improve the inference speed of each model. It can process video frames in a short time and improve the performance of abnormal behavior analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the present invention, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is a flowchart of the real-time examination room candidate abnormal behavior analysis method; Figure 2 It is a schematic diagram of the abnormal behavior analysis model; Figure 3 It is a flowchart of an embodiment of the real-time examination room candidate abnormal behavior analysis method; Figure 4 Schematic diagram of the real-time examination room candidate abnormal behavior analysis system; Figure 5 Network model diagram of YOLOx; Figure 6 Schematic diagram of an electronic device. Specific implementation manner

[0030] The real-time examination room candidate abnormal behavior analysis method provided by this application is to solve the problem that the existing examination video monitoring system collects videos through cameras and can only record and save videos, and cannot judge abnormal situations during the examination. Currently, only monitoring personnel can search for abnormal behaviors by playing back the video recordings after or during the examination, lacking the real-time monitoring of the examination process.

[0031] The deep learning method based on the TPU device in this application provides an efficient and accurate way to analyze the abnormal behaviors of candidates in the examination room. The TPU involved in this application is a tensor processing unit, which can be used as a hardware accelerator specifically optimized for deep learning and has significant advantages.

[0032] Specifically, it can collect the video stream in the examination room and then perform frame extraction on the high-density video stream; select key frames from the consecutive video frames after frame extraction for object detection to accurately locate the spatial position of the candidates. Then, perform candidate abnormal behavior analysis on the video segments and candidate positions. In order to achieve automatic analysis, key frames and video segments of abnormal behaviors can be drawn according to the abnormal behavior recognition results.

[0033] The types of abnormal behaviors involved in this application include but are not limited to: actions of candidates not following the rules during the examination, such as candidates tilting their heads left and right, candidates tilting their heads backward, candidates passing suspicious items, candidates putting their hands under the table and burying their heads, etc., which may affect the normal progress of the examination.

[0034] It should be noted that this application is for dealing with possible abnormal behaviors of candidates during the examination in a standard examination room. Some irregular actions of candidates during the examination, such as candidates tilting their heads left and right, candidates tilting their heads backward, candidates passing suspicious items, candidates putting their hands under the table and burying their heads, etc., are detected in real time, and the key frames and the original video segments are returned to the upper-layer server for manual review and confirmation.

[0035] This application combines manual secondary verification of the abnormal behavior recognition results to ensure the accuracy of recognition. This application also uses the TPU hardware to accelerate the deep learning model inference process and optimize the real-time behavior analysis performance.

[0036] The specific process of the real-time examination room candidate abnormal behavior analysis method will be described in detail below. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details.

[0037] Statements such as "an embodiment" or "some embodiments" described in the present application mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" that appear in different places in the present application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] Please refer to Figure 1 Shown is a flowchart of the real-time examination room candidate abnormal behavior analysis method in a specific embodiment. The method includes: S1: Real-time collect the video stream in the examination room.

[0040] According to the embodiments of the present application, through the Hikvision high-definition camera in the classroom, use the RTSP protocol to real-time collect the video stream, and use a dedicated collection server to receive and decode the video signal transmitted by the camera to ensure the stability and real-time of the video stream.

[0041] In some specific embodiments, high-definition cameras with network transmission functions can be reasonably arranged in the examination room to ensure coverage of all areas of the examination room. The high-definition cameras are connected to the upper-layer server or cloud server of the examination room through communication methods such as Ethernet or Wi-Fi, continuously capture the images in the examination room at 25 frames per second or 30 frames per second, and real-time transmit the video data in formats such as H.264 and H.265.

[0042] This embodiment monitors the camera in real time. If a reading failure is found, the camera is automatically reread to re-obtain the RTSP video stream to ensure that the video stream is not interrupted.

[0043] This embodiment can comprehensively record the actual situation in the examination room. Through real-time acquisition, various behaviors of candidates during the examination can be captured in a timely manner without missing key information.

[0044] Optionally, the method of collecting the video stream in the examination room can be processed based on the following method. The video sampling information of each frame is RS, and the resolution can be adjusted according to actual needs. The resolution can be set from 1080p to 720p.

[0045] 。

[0046] S2: Decimate the video stream at a preset frame interval.

[0047] In some embodiments, according to the requirements of the examination room monitoring and the frame rate of the video stream, a fixed decimation interval is set, and fixed-length video frame segments are generated by decimation. Then, preprocessing operations such as image scaling are performed on the decimated video segments. A data recording and storage mechanism is established for the extracted video frames for subsequent behavior analysis, anomaly detection, and manual verification.

[0048] As an embodiment of the present application, a fixed frame interval value can be set. For example, 1 frame is extracted every 5 frames. After receiving the real-time video stream, according to the set frame interval, the corresponding video frames are sequentially extracted from the video stream, and the extracted frames are stored in a memory buffer or a temporary file for subsequent further processing.

[0049] S3: Select key frames from the decimated video frames for object detection to locate the positions of the candidates.

[0050] In some embodiments, in order to improve the accuracy of candidate detection in a dense examination scenario, the object detection model can use the YOLOx model to extract features of each candidate in the video frame through multiple layers of CNN, and finally predict the spatial position of the candidate.

[0051] After detecting the target using the YOLOx model, it can be determined whether the state of the target is a sitting or standing state; if the detected state is a standing state, the target is automatically filtered, and only the coordinate information of the candidates in the sitting state is saved for subsequent target area annotation and drawing in the processing steps.

[0052] It should be noted that the decimated video frames are input into the pre-trained object detection model, that is, the YOLOx model. The YOLOx model recognizes and locates each object in the image, and outputs the detection results including the candidate category and the coordinate information of the candidates in the image, etc.

[0053] This application can determine the specific position of the candidate in each frame of the image according to the coordinate information.

[0054] For the selection of key frames, metrics such as pixel differences and image entropy between adjacent frames can be calculated. When the change exceeds a certain threshold, it is recognized as a key frame. It can also be determined based on rules such as time intervals. Priority is given to performing object detection on key frames to improve efficiency. For the YOLOx model to identify and locate each object in the image, the input image can be divided into multiple grid cells. Each cell predicts multiple bounding boxes and corresponding class probabilities and other information. The object detection results can be quickly obtained through the forward propagation of the convolutional neural network. The training process optimizes the model parameters through backpropagation based on a large dataset of labeled candidate image data.

[0055] Object detection can locate the candidate's position, providing spatial coordinates for subsequent analysis of their behavior. The selection of key frames further focuses on important information, reduces the amount of calculation, and improves the processing speed of the candidate status analysis process.

[0056] S4: Perform abnormal behavior analysis based on the extracted video segments and the candidate's position.

[0057] According to the embodiments of the present application, abnormal behavior analysis can be performed on the extracted video segments and the candidate's position based on the Figure 2 abnormal behavior analysis model as shown.

[0058] This embodiment can also use algorithms such as the Slowfast behavior analysis algorithm, and extract the spatial features and temporal features of video frames through parallel slow channels and fast channels respectively, so as to effectively capture the detailed information and global dynamics in the video frames.

[0059] It should be noted that the slow channel processes low-frame-rate video information and captures the global changes in the scene, which is suitable for analyzing the overall behavior patterns of candidates.

[0060] The fast channel focuses on the fast movements in the high-frame-rate video stream, ensuring that when the candidate's actions change rapidly, the details can be accurately identified to adapt to the complex examination environment.

[0061] This embodiment can also define a series of abnormal behavior rules and feature templates. For example, by judging the coordinate changes of the examinee's position in multiple consecutive frames, if it exceeds the seat area range and the duration exceeds a set threshold, such as 5 minutes, it is considered abnormal. By detecting the change frequency of the examinee's head orientation within a certain time, when it exceeds a certain set frequency value, it is determined to be abnormal, which can be achieved based on the head pose estimation algorithm. By detecting the distance between different examinee positions and information such as face orientation, when the distance is too close and the face orientation towards each other lasts for a certain time, it is determined, etc. Then, combined with the position information of the examinee and the corresponding image features in the video segment after frame extraction, it is compared and analyzed one by one according to the abnormal behavior rules and templates to determine whether there is abnormal behavior, and record information such as the time when the corresponding abnormal behavior appears and the examinees involved.

[0062] Exemplarily speaking, the Slowfast behavior analysis algorithm can be used to judge looking around, or according to the pose estimation model, by detecting the key points of the examinee's head and calculating the angles, the change of the head orientation can be analyzed.

[0063] For other postures of the examinee, corresponding analysis algorithms can be used to judge to effectively monitor the exam. In this way, it can accurately judge whether the behavior of the examinee in the examination room is abnormal and timely discover behaviors that violate the examination discipline.

[0064] S5: According to the abnormal behavior recognition result, parse the key frames of the abnormal behavior and the original video segment.

[0065] According to the embodiment of the present application, the recognized abnormal behavior result is the possible abnormal behavior of the examinee predicted based on the algorithm model. The abnormal behavior is drawn in the key frames, including drawing the position of the examinee with a red rectangle frame and marking the type of the abnormal behavior in red font.

[0066] After the abnormal behavior is recognized, according to the recorded time stamp when the abnormal behavior appears and the corresponding examinee information, find the key frames containing the abnormal behavior from the previously stored video frames after frame extraction, and at the same time, extract the corresponding video segment with the complete time range from the original video stream through the time correspondence relationship.

[0067] This embodiment can intercept the original video in the time range from a few seconds before the start of the abnormal behavior to a few seconds after the end. In this way, it is convenient for subsequent manual secondary verification and as evidence retention, etc., can show the process of the abnormal behavior occurring, and can judge the nature of the behavior and make subsequent processing decisions.

[0068] This embodiment can recognize the original video segment based on the recognition algorithm of the original video segment.

[0069] The recognition algorithm of the original video segment is:

[0070] YsF represents the original video segment obtained by the final parsing. The original video segment is the content for subsequent manual review, which contains the key frames related to the abnormal behavior and other frames within a certain time range before and after the key frames, and can provide more comprehensive information about the situation before and after the occurrence of the abnormal behavior for the manual review personnel.

[0071] YcGJ represents the key frame of the abnormal behavior detected at time t. The key frame is identified from the video frames after frame extraction and contains the key picture of the occurrence of the abnormal behavior. For example, in the examination room scenario, it may be the frame when the examinee makes a cheating action.

[0072] represents the set of all video frames from time (t - △t) to time (t + △t).

[0073] where: t is the timestamp corresponding to the key frame, which marks the time position of the key frame of the abnormal behavior in the entire video stream.

[0074] △t is a time offset used to determine the time range to be included before and after the key frame. Through this time range, the video frames within a certain time before and after the key frame can be obtained to observe the occurrence process of the abnormal behavior more comprehensively. For example, when △t is 5 seconds, it means to obtain all the video frames within 5 seconds before and after the key frame.

[0075] represents the union operation of the sets. The key frame and the set of all video frames within the time range are merged together to form the final original video segment.

[0076] It should be noted that the key frame records its corresponding timestamp during the frame extraction process. Through this timestamp, the position of the key frame can be accurately located in the original video stream. Taking the key frame timestamp as the center, a time window of ±5 seconds is set, and the video segment within this time window is extracted from the original video stream to ensure that the extracted video segment contains sufficient information about the situation before and after the occurrence of the abnormal behavior, facilitating the manual review personnel to understand the whole picture of the abnormal behavior. The key frame and the optical flow features are used to generate the condensed video segment. The optical flow features can describe the motion information between video frames. Combining with the key frame, the dynamic changes of the abnormal behavior can be captured more accurately.

[0077] In this embodiment, compared with the original video segment, the data volume of the generated video summary is compressed. During the compression process, by reasonably using the key frame and the optical flow features, the key action information of the abnormal behavior is retained, enabling the manual review personnel to view the condensed video summary in a shorter time, quickly judge the situation of the abnormal behavior, and improve the review efficiency.

[0078] S6: Display the abnormal behavior recognition result information for manual secondary verification.

[0079] The manual secondary verification involved in this embodiment is to display the abnormal behavior recognition result information through the system, which can be displayed on the computer of the verification personnel or the handheld terminal, enabling the verification personnel to review and confirm the abnormal recognition result in real time. If the abnormal video clip does have abnormalities, then it is determined that the video clip is abnormal.

[0080] If the abnormal video clip does not have abnormalities, then it is determined that the video clip is normal.

[0081] For the video clips with abnormalities, implement a hierarchical reporting mechanism, reporting from the examination site to the upper level, or uploading to the upper-layer server.

[0082] The result information of abnormal behavior recognition includes: the type of abnormal behavior, the examinees involved, the occurrence time, and the corresponding key frames and video clips of the parsed abnormal behavior, which are displayed through a visualization interface.

[0083] Optionally, this embodiment can be based on the front-end display of a web page or a desktop application. After the backend transmits the relevant data, the video information is displayed in a list form on the front-end display. At the same time, functions such as playing video clips and viewing key frame pictures are provided to facilitate manual viewing and secondary verification of the accuracy of abnormal behavior. Improve the accuracy of abnormal behavior recognition, avoid incorrect results caused by algorithm misjudgment and other reasons, and ensure the fairness of the determination of examinees' behavior.

[0084] S7: Utilize TPU hardware to accelerate the deep learning model inference process and optimize the real-time behavior analysis performance.

[0085] In this embodiment, the object detection and behavior analysis model can obtain the weight file by training under the GPU, and then convert the weight file into the intermediate format ONNX. Then convert the model into an FP16 model in the bmodel format suitable for use on the Synnovation TPU device. Finally, further quantize the FP16 model into an INT8 model for inference.

[0086] It can be seen that this embodiment can allocate the computing tasks of the YOLOx model to the TPU (Tensor Processing Unit) hardware for execution. Through the adaptation interface of the deep learning model, configure the model running environment so that the deep learning model can call the TPU for computing acceleration.

[0087] Specify the use of TPU devices during the model initialization phase. When performing operations such as object detection on the input video frames, the TPU will process a large number of tensor operations in parallel, significantly shortening the calculation time and improving the overall real-time behavior analysis speed. This allows for processing a larger number of candidate video frame data within a certain time, ensuring the timely detection and analysis of abnormal behaviors during the exam and meeting the requirements of real-time monitoring.

[0088] As an implementation manner of this embodiment, the acceleration configuration method of the TPU can be set: TPU = Tcc / Ttc.

[0089] Tcc is the time required for the terminal to perform model inference. Ttc is the time required for the TPU to perform model inference.

[0090] In this embodiment, the acceleration of the deep learning model can be executed based on the following algorithm:

[0091] I represents the input feature map, K represents the convolutional kernel, and O represents the output feature map. In the TPU, the above formula can be used to perform the convolution operation in parallel, accelerating the calculation process. This embodiment can reduce the memory access latency based on the combined convolution, batch normalization, and ReLU layers. Combining these three operations into one operation and completing it in one calculation reduces the storage of intermediate results and data transmission, thereby improving the calculation efficiency.

[0092] In an embodiment of the present invention, based on step S2, as Figure 3 shown, during the exam behavior analysis process, frame extraction is a very important step. Since the actions of candidates usually change little during the answering process and the semantics between adjacent video frames are relatively similar, frame extraction can be performed on the high-density video stream. The core purpose of frame extraction is to reduce the computational burden, improve the real-time processing efficiency, and at the same time retain the key spatio-temporal information as much as possible. The following will give a possible embodiment to non-restrictively elaborate on its specific implementation scheme.

[0093] S201: Set a fixed frame extraction interval according to the requirements of the examination room monitoring and the frame rate of the video stream, and generate a fixed-length video frame segment through frame extraction.

[0094] This embodiment can set a fixed frame extraction interval according to the specific requirements of the examination room monitoring, and then generate a fixed-length video frame segment through frame extraction. For example, if the frame rate of the video stream is 30 frames per second, the frame extraction interval can be set to extract 1 frame every 5 frames, that is, extract 1 frame every 5 frames. The fixed frame extraction interval can be set in combination with the actual monitoring accuracy requirements.

[0095] Specifically, the RTSP video stream obtained in step S1 can be decoded using OpenCV, and the original video frames after decoding are , where t= 1,2,3 …M , set the frame extraction interval to N . In this example N = 2, that is, save one frame every 2 frames, and the representative frames after frame extraction are Frame . The specific frame extraction method is as shown in formula (1): (1) S202: Take every Frame frame of the video frames after frame extraction T as a combination to generate video frame segments of a fixed length. In this example T = 32.

[0096] S203: Save the original video R corresponding to the 32-frame video segment after frame extraction. In this example R = 64.

[0097] For step S2, preprocessing operations such as image scaling are also performed on the video segments after frame extraction. Image scaling can be implemented based on interpolation algorithms. By specifying the target width and height and selecting an appropriate interpolation algorithm, each video frame after frame extraction is adjusted to an appropriate size.

[0098] Establish a data recording and storage mechanism for the extracted video frames for subsequent behavior analysis, anomaly detection, and manual verification. Here, a data recording form can be created to record the relevant key information of each extracted video frame. Exemplarily, the frame number, extraction time, frame size information, etc. can be set. Facilitate operations such as data retrieval, update, and correlation analysis, and facilitate accurate acquisition and utilization of relevant video frame data in subsequent behavior analysis, anomaly detection, and manual verification.

[0099] Combined with the specific process of step S2 above, in step S3 of this embodiment, the spatio-temporal action behavior detection algorithm requires a set of coordinate frame inputs and a set of video frame segments. The coordinate input is used to define which targets need to be analyzed for behavior, and the video segment is used for behavior analysis. Therefore, it is necessary to use an object detection algorithm to perform object detection on the key frames before behavior analysis.

[0100] The following further describes step S3.

[0101] S301: The length of the video frame sequence after frame extraction is 32. Select the middle frame as the key frame for object detection. The calculation method of the middle frame is as shown in formula (2): (2) S302: Input the key frame KeyFrame into the YOLOx object detection model for inference. During the inference process, set the confidence threshold to 0.3 and the NMS threshold to 0.3 to obtain the inference results of the key frame.

[0102] Specifically, execute as shown in formula (3): (3) where Result contains all the object information in the entire image.

[0103] Result is composed of the object category Class , confidence Score and coordinate position Box . The coordinate Box is composed of the upper left corner coordinates of the object bounding box and the lower right corner coordinates .

[0104] S303: Filter the objects set in the object detection results Result to finally obtain the sitting examinee object.

[0105] S304: Store the examinee object in the boxes_list list and input it into the subsequent behavior recognition algorithm.

[0106] In this embodiment, after combining the preservation of the original video corresponding to the 32-frame video segment after frame extraction and storing the examinee object in the R list, 1 group of representative frames boxes_list = 32 frames and the object detection results of the key frames KeyFrame corresponding to the representative frames can be obtained. T Based on step S4, input the object detection results boxes_list of 32 consecutive representative frames and KeyFrame into the behavior recognition model Slowfast for behavior analysis.

[0107] In this embodiment, within the behavior analysis algorithm, the video segment is divided into two branches, namely the slow path and the fast path. Among them, the slow path further samples the 32-frame video after frame extraction into 8 frames, and the input of the fast path is the T-frame video after frame extraction. The corresponding formulas for the slow path and the fast path are shown in formulas (4) and (5):

[0108] (4) (5) (5) Next, the results output by the slow path and the fast path are concatenated to finally obtain the fused output, as shown in formula (6): ) (6) This embodiment also filters the predicted normal behaviors and returns the abnormal behavior results.

[0109] Combined with the abnormal behaviors returned in step S4, in step S5, the relevant video frames will be immediately marked and an abnormal behavior report will be generated.

[0110] Specifically, S501: Use a red rectangular box in the key frame KeyFrame to frame the position of the abnormal candidate, and mark the type of abnormal behavior with red text.

[0111] S502: Save the original video frame R corresponding to the representative frame T of the abnormal candidate group.

[0112] S503: Return the key frame with the abnormal behavior drawn and the original video R to the upper-layer server through the HTTP service.

[0113] In order to determine the abnormal behaviors of the abnormal candidates, combined with step S6, an artificial review function is provided. The results of abnormal recognition are manually reviewed and confirmed in real time. If the abnormal video segment is indeed abnormal, the reviewed video segment is abnormal; if the abnormal video segment is not abnormal, the reviewed video segment is normal.

[0114] This embodiment reports the abnormal video segments.

[0115] Combined with step S7 involved in the above of this application, the TPU hardware can be used to accelerate the deep learning model inference process and optimize the real-time behavior analysis performance. Specifically, it can be implemented based on the following execution methods: S701: The target detection model and behavior analysis model used are trained under the GPU to obtain files in pth format.

[0116] Here, the YOLOx target detection model and behavior analysis model can be constructed based on the PyTorch module. For example, the behavior analysis model can analyze the candidate behavior sequence by constructing a convolutional neural network or a recurrent neural network, etc.

[0117] S702: Convert the files in pth format into the ONNX intermediate format.

[0118] S703: Further convert the network model in the ONNX intermediate format into the bmodel format.

[0119] S704: Further quantize the bmodel format model into an INT8 model to improve the video stream parallel processing ability.

[0120] It can be seen that after quantizing the model into an INT8 model, the calculations in the TPU can be more efficient. The TPU can utilize the integer arithmetic unit when processing INT8 data, thereby improving the parallel processing ability of the video stream. Real-time processing of high-density examination hall video streams can process more video frames in a short time and detect abnormal behaviors in a timely manner.

[0121] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0122] The following are embodiments of the real-time examination hall candidate abnormal behavior analysis system provided by the present disclosure. This system and the real-time examination hall candidate abnormal behavior analysis methods of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the real-time examination hall candidate abnormal behavior analysis system, reference can be made to the embodiments of the real-time examination hall candidate abnormal behavior analysis method.

[0123] As Figure 4 shown, the system includes: a video acquisition module for real-time acquisition of the video stream in the examination hall.

[0124] A frame extraction processing module for performing frame extraction processing on the video stream at a preset frame interval.

[0125] A candidate positioning module for selecting key frames from the frame-extracted video frames for object detection to locate the positions of the candidates.

[0126] A behavior analysis module for performing abnormal behavior analysis based on the frame-extracted video segments and the candidate positions.

[0127] An abnormality parsing module for parsing the key frames of the abnormal behavior and the original video segments according to the abnormal behavior recognition results.

[0128] A display verification module for displaying the abnormal behavior recognition result information for manual secondary verification.

[0129] An analysis and inference module for using the TPU hardware to accelerate the deep learning model inference process and optimize the real-time behavior analysis performance.

[0130] In connection with the real-time examination hall candidate abnormal behavior analysis system involved in the present application, the following modules are further included: A time judgment module: It can start the system during the examination, thereby effectively avoiding false alarms in the periods before and after the official examination, ensuring that the behavior analysis is only carried out during the examination period, and improving the accuracy and efficiency of the algorithm.

[0131] Configure the producer module and the consumer module. In this embodiment, the video frame processing process is optimized by introducing the producer module and the consumer module.

[0132] The producer module is used to read video frames in real time from the RTSP video stream and push them to the shared queue, while the consumer module fetches the video frames from the queue and performs subsequent processing, significantly reducing the real-time video stream reading latency and improving the processing efficiency.

[0133] The timestamp naming and data feedback module is used to uniquely name each group of key frames containing abnormal behaviors and the original video segments based on timestamps, and send the processing results to the upper-layer server via an HTTP request for further storage and analysis.

[0134] In the system of this embodiment, based on the above modules, the target detection and behavior analysis models are trained on the GPU to obtain weight files, then the weight files are converted into the intermediate format ONNX, and further the models are converted into the FP16 model in the bmodel format suitable for use on the SOPHGO TPU device, and finally quantized into the INT8 model for inference to ensure that the behavior analysis can be performed in real time.

[0135] Among them, for the target detection model involved in the candidate positioning module, 6635 pictures in a real examination room can be selected for annotation, with candidates marked as class 0 and invigilators marked as class 1. The 6635 pictures are divided into a training set of 5308 pictures and a validation set of 1327 pictures according to the ratio of 8:2.

[0136] The target detection model is trained by iterating 300 times on the training set of pictures. For example, Figure 5 is the network model diagram of YOLOx. The preprocessed pictures with a size of 640x640x3 are input into the target detection model. The pictures are continuously downsampled in the backbone feature extraction network, compressing the image size to the minimum while expanding the number of channels to the maximum.

[0137] Then, in the feature fusion part, the three effective feature layers extracted by the backbone feature extraction network are fused. Through this operation, the feature information of different feature dimensions can be fused, which is beneficial to the detection of small, medium, and large targets. Finally, the features after feature fusion are input into the YOLO-head for regression prediction, and the target information in the examination room is finally obtained.

[0138] The loss function of YOLOx is shown in formula (1): (1) Among them, represents the classification loss, represents the localization loss, represents the obj loss, The balance coefficient representing the positioning loss is set to 5.0. It represents the number of Anchor Points classified as positive samples.

[0139] As Figure 6 As shown, the present application also provides an electronic device, including a display module 103, a memory 102, a processor 101, and a computer program stored on the memory and executable on the processor 101. When the processor 101 executes the program, it implements the steps of the real-time analysis method for abnormal behaviors of examinees in the examination room.

[0140] In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers.

[0141] The electronic device can also use the Synnovation SE8-288 TPU micro-server, which has powerful video encoding and decoding capabilities, and uses TPU device hardware acceleration for the object detection model and the behavior analysis model, greatly improving the computational efficiency and speed during the model inference process. The hardware acceleration enables the system to process and recognize the behaviors of multiple examinees in real time, especially in a large-scale examination room environment, and can ensure the simultaneous analysis of multiple videos.

[0142] In the embodiments of the present application, the processor 101 can be implemented by using at least one of an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, and an electronic unit designed to execute the functions described herein. In some cases, such an implementation can be implemented in a controller. For a software implementation, an implementation of a process or function can be implemented with a separate software module that allows the execution of at least one function or operation. The software code can be implemented by a software application (or program) written in any suitable programming language. The software code can be stored in the memory and executed by the controller.

[0143] The display module 103 is used to display information input by the user or information provided to the user. The display module 103 may include a display panel, and the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0144] The memory 102 can be used to store software programs and various data. The memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0145] This application also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the real-time examination room candidate abnormal behavior analysis method are implemented.

[0146] The storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0147] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for analyzing abnormal behaviors of examinees in a real-time examination room, characterized in that, The method includes: Step 1: Collect the video stream in the examination room in real time; Step 2: Perform frame extraction on the video stream at a preset frame interval; Step 3: Select key frames from the extracted video frames for object detection to locate the positions of the examinees; Step 4: Perform abnormal behavior analysis based on the extracted video segments and the positions of the examinees; Step 5: According to the abnormal behavior recognition result, parse the key frames of the abnormal behavior and the original video segments; Step 6: Display the abnormal behavior recognition result information for manual secondary verification; Step 7: Utilize TPU hardware to accelerate the inference process of the deep learning model and optimize the real-time behavior analysis performance.

2. The real-time abnormal behavior analysis method of examinees in the examination room according to claim 1, wherein in Step 1, the video stream collected in the examination room is an RTSP video stream; Read video frames from the RTSP video stream in real time and push the video frames to the shared queue; Obtain video frames from the shared queue and process them; Step 2 further includes: decoding the acquired RTSP video stream based on OpenCV, and the decoded original video frames are , where t= 1,2,3 …M ; Set the frame extraction interval to N , and the representative frame after frame extraction is Frame . The specific frame extraction method is as follows: The video frames after frame extraction Frame Every T frame is taken as a combination to generate video frame segments of a fixed length; Save the original video corresponding to the preset frame video segment after frame extraction. R Perform saving.

3. The real-time abnormal behavior analysis method of examinees in the examination room according to claim 2, wherein Step 3 further includes: Select the middle frame from the extracted video frames as the key frame for object detection, and the calculation method of the middle frame is: Input the key frame KeyFrame into the following YOLOx object detection model for inference and set the confidence level; Among them, Result contains all the target information in the entire image; Result including: the category of the target Class , confidence Score and coordinate position Box , the coordinate position Box including: the upper left corner coordinates of the target bounding box ) and the lower right corner coordinates ); Filter the standing objects in the target detection results Result to finally obtain the sitting examinee objects; Store the examinee's target in boxes_list the list.

4. The real-time abnormal behavior analysis method of examinees in the examination room according to claim 3, wherein Step 4 further includes: After frame extraction in Step 2 and object detection in Step 3, obtain a set of representative frames and the object detection results of the key frame KeyFrame corresponding to the representative frames; Input the object detection results boxes_list of consecutive representative frames and KeyFrame into the behavior recognition model Slowfast for behavior analysis; In the behavior analysis algorithm, the video segment is divided into two branches, namely the slow path and the fast path; Among them, the slow path further samples the extracted video frames, and the input of the fast path is the T-frame video after frame extraction. The formulas of the slow path and the fast path are as follows: Stitch the results output by the slow path and the fast path based on the following formula to finally obtain the fused output; ) Filter the behaviors predicted to be normal and return the abnormal behavior results.

5. The real-time abnormal behavior analysis method of examinees in the examination room according to claim 1 or 2, wherein Step 5 further includes: According to the abnormal behavior returned in Step 4, use a rectangular box in the key frame KeyFrame to frame the position of the abnormal examinee and mark the type of abnormal behavior with text; Save the original video frame R corresponding to the representative frame T of the abnormal behavior group; Also, perform unique naming based on the time stamp for the key frames and the original video segments containing abnormal behaviors, and send the processing results to the upper-layer server through an HTTP request.

6. The real-time abnormal behavior analysis method of examinees in the examination room according to claim 1, wherein Step 7 further includes: the target detection model and the behavior analysis model used are trained under GPU to obtain a file in pth format; Convert the file in pth format into an ONNX intermediate format; Convert the network model in ONNX intermediate format into bmodel format; Quantize the bmodel format into an INT8 model.

7. The real-time examination room candidate abnormal behavior analysis method according to claim 1, characterized in that Step 3 further includes: extracting a preset number of pictures based on key frames and performing annotation, marking candidates as class 0, and marking invigilators as class 1; Divide the preset number of pictures into a training set and a validation set according to a preset ratio; Perform multiple iterative trainings on the YOLOx network model based on the training set to obtain a trained YOLOx network model; Input the preprocessed picture into the YOLOx network model, perform downsampling in the backbone feature extraction network, compress the image size to a preset size, and at the same time expand the number of channels to a preset expansion amount; Perform feature fusion on the three effective feature layers extracted by the backbone feature extraction network, fuse the feature information of each different feature dimension, and then input the feature after feature fusion into the YOLOx network model for regression prediction to finally obtain the target information in the examination room; The loss function formula of the YOLOx network model is: Among them, represents the classification loss, represents the localization loss, represents the obj loss, represents that the balance coefficient of the localization loss is set to 5.0, represents the number of Anchor Points classified as positive samples.

8. A real-time analysis system for abnormal behaviors of examinees in an examination hall, characterized in that, The system is used to implement the real-time examination room candidate abnormal behavior analysis method described in any one of claims 1 to 7; the system includes: A video acquisition module for real-time acquisition of the video stream in the examination room; A frame extraction processing module for performing frame extraction processing on the video stream at a preset frame interval; A candidate positioning module for selecting key frames from the frame-extracted video frames for target detection to locate the position of the candidate; A behavior analysis module for performing abnormal behavior analysis based on the frame-extracted video segment and the candidate position; An abnormal parsing module for parsing the abnormal behavior key frames and the original video segment according to the abnormal behavior recognition result; A display verification module for displaying the abnormal behavior recognition result information for manual secondary verification; An analysis and inference module for using TPU hardware to accelerate the deep learning model inference process and optimize the real-time behavior analysis performance.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the real-time examination room candidate abnormal behavior analysis method described in any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the real-time examination room candidate abnormal behavior analysis method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Behavior identification alarm method, system and device of semiconductor intelligent factory and medium

    CN120953928A

  • Cloud-edge collaborative abnormal behavior character recognition method, device and system

    CN121121872A