Real-time end-to-end human behavior detection system based on deep spatial-temporal feature fusion and detection method thereof
A real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, combined with YOLO network and temporal buffer, enables real-time proctoring of offline examinations for multiple participants. This solves the problem of existing technologies being unable to accurately judge complex movement changes, thus improving the efficiency and accuracy of proctoring.
Patent Information
- Application Number
- CN202510915845.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-31
AI Technical Summary
Existing online examination systems cannot meet the real-time requirements of offline examinations with multiple participants, cannot accurately judge the complex changes in candidates' actions, have high computational complexity, and cannot detect cheating behavior in a timely manner during the examination.
A real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion is adopted. Through a feature extraction module, a detection head, a temporal buffer, and a behavior recognition module, combined with YOLO's Backbone network and a target/key point detection head, behavior analysis is performed using a temporal buffer and spatiotemporal graph convolution to achieve real-time proctoring for multi-person examinations.
It can accurately detect abnormal behavior of candidates in real time during multi-person exams, reducing the workload of invigilators, improving the fairness and efficiency of the exam, and reducing labor costs.
Smart Images

Figure CN120877327A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, specifically relating to a real-time end-to-end human behavior detection system and method based on deep spatiotemporal feature fusion. Background Technology
[0002] Currently, examinations are mainly divided into two categories: online examinations and offline examinations. Online examinations are proctored using multi-camera setups and microphones, with each candidate in a relatively independent environment. Some online examination systems, such as ExamStar, are widely used for this type of individual online examination. In these systems, the proctoring process relies heavily on computer technology. Before the exam begins, facial recognition technology is typically used for identity verification to ensure the candidate is present. During the exam, multi-camera setups record the candidate's facial expressions and body movements to prevent cheating behaviors such as checking their phone or communicating with others; microphones capture ambient sound to ensure a quiet environment and prevent cheating via voice; for online exams where electronic devices are used to answer questions, screen recording is also implemented, and candidates are prohibited from switching screens or copying and pasting to prevent cheating through web searches or sending exam content elsewhere during the exam. After the exam, proctors can review the entire process using the system-generated proctoring report and exam recording.
[0003] These types of online exams have many advantages. First, the exam location is no longer restricted; candidates can take the exam anywhere with an internet connection. Second, automated proctoring reduces reliance on human proctors, lowering labor costs, as proctors can use the online system to proctor multiple candidates simultaneously, thus saving significant human resources. Third, the online system has strong cheat detection capabilities, accurately determining whether candidates are cheating, thereby improving the fairness and credibility of the exam.
[0004] However, most exams are still conducted offline. In physical exam rooms, invigilation typically relies on the collaborative efforts of invigilators, in-room cameras, and proctors. This method has several drawbacks: firstly, it requires significant human and material resources; exam venues must provide the space and pay invigilators, resulting in high costs; secondly, human invigilation is affected by the invigilator's condition, making it difficult for them to maintain full focus throughout the exam, inevitably leading to fatigue or oversight. Existing online exam systems cannot overcome these shortcomings; video analysis functions for individual candidates cannot meet the real-time requirements of multiple-person exams, and their effectiveness may be inferior to human invigilation.
[0005] An existing technical solution is a cheating detection system based on YOLO V8. YOLO (You Only Look Once) is a very popular object detection algorithm that can complete the detection process in a single forward propagation without complex region proposal processes, and provides high accuracy. This solution is implemented in two stages: the first stage uses YOLO V8 for object detection, identifying all examinees in the monitoring screen; the second stage performs behavior classification, extracting image patches containing individual examinees using ROI (Region of Interest) technology, and using a classification network to determine their current behavior category, such as normal, standing, looking back, etc., thereby determining whether cheating has occurred.
[0006] This approach has some problems. Although the YOLO algorithm performs well in object detection, it mainly processes single-frame images and does not consider temporal information, which may prevent it from accurately judging the examinee's behavior. On the one hand, some rapid actions, such as looking up or down, may cause image blurring, leading to misjudgment or missed judgment; on the other hand, examinees may perform a series of actions during the exam, such as looking down at the exam paper, then looking up at their surroundings, and then looking down again to continue answering questions. The single-frame-based detection mode cannot capture such complex changes in movement.
[0007] Traditional action recognition methods can effectively utilize temporal information, but their time complexity is often too high. Methods based on dual-stream networks separate spatial and temporal channels (pixels and optical flow), which can improve action representation capabilities, but the computation of optical flow requires a significant amount of time. Methods based on 3D convolution (such as C3D and I3D) directly model spatiotemporal relationships through 3D convolution, effectively capturing short-term action features, but their computational complexity increases cubically with the length of the video sequence, leading to excessive memory consumption and insufficient real-time performance. Attention-based methods (such as spatiotemporal attention) can focus on keyframes and key regions, but the attention mechanism has high time complexity and requires massive amounts of training data to avoid overfitting, making practical implementation difficult.
[0008] Existing online examination systems have several significant shortcomings. First, they can only analyze individual candidates at a time, proving inadequate for scenarios involving multiple candidates in offline examinations, which limits their practical value, given that most exams are conducted offline. Second, existing online examination systems typically require a large number of consecutive video frames to analyze candidate behavior, often requiring analysis reports to be generated after the exam, making real-time monitoring impossible. This prevents timely alerts to potential cheating during the exam, and post-exam review significantly increases the workload for exam administrators. Summary of the Invention
[0009] This invention provides a real-time end-to-end human behavior detection system and method based on deep spatiotemporal feature fusion, which can solve the problem of proctoring multiple examinees more efficiently and reliably, that is, it can analyze the examination process of multiple examinees in real time at the same time.
[0010] This invention is achieved through the following technical solution: A real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, the system comprising a feature extraction module, a detection head, a temporal buffer, and a behavior recognition module; The feature extraction module is connected to both the target detection head and the human pose key point detection head. The target detection head is connected to the ROI extraction module, and the ROI extraction module is connected to the behavior recognition module. The human posture key point detection head is connected to the ROI-guided key point extraction module, which is connected to the behavior recognition module. The key point extraction module is also connected to a temporal buffer. The human posture key point detection head is used to detect the key point positions of the head, left shoulder, right shoulder, left elbow, right elbow, left wrist, and right wrist.
[0011] Furthermore, the ROI extraction module extracts pixel blocks based on the position information detected by the target detection head, and then inputs the pixel blocks into the behavior recognition module for abnormal behavior information recognition.
[0012] Furthermore, the RIO-guided key point extraction module uses the human skeleton nodes detected by the human posture key point detection head as key points and inputs them into the temporal buffer. The temporal buffer annotates the key points with time sequence to form temporal key points, and then inputs the temporal key points into the behavior recognition module for abnormal behavior information recognition.
[0013] A detection method for a real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, the method using the aforementioned real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, the detection method comprising: acquiring a monitoring screen, the feature extraction module receiving a single frame image from the video stream, the feature extraction module extracting image features from the single frame image, and the image features being respectively input to a target detection head and a human posture key point detection head; The target detection head extracts target location information from image features and inputs it into the ROI extraction module; The human pose key point detection head extracts key points containing human skeleton nodes from image features and inputs them into the ROI-guided key point module. The behavior recognition module takes the target pixel block and the sequence of temporal key points at the current moment as input and outputs the behavior information of the target examinee, realizing real-time end-to-end human behavior detection based on deep spatiotemporal feature fusion.
[0014] Furthermore, the ROI extraction module cropped out sub-regions containing only a single examinee and standardized the size of the image blocks to 32x32. The ROI-guided keypoint module categorizes all keypoints according to the candidates they belong to.
[0015] Furthermore, the time-series buffer specifically stores the key points of human skeletal nodes. When the time-series buffer processes frame T, it stores the key points of frames T-1, T-2, ..., T-31. When it switches to processing frame T+1, the contents of the buffer become the key points of frames T, T-1, ..., T-30.
[0016] Furthermore, the behavior recognition module consists of two information processing paths. One path processes the input target pixel block through multiple 2D convolutional layers and max pooling layers to finally output a one-dimensional feature vector. Another approach is to extract spatiotemporal action features from the input temporal keypoint sequence through multiple spatiotemporal-graph convolutions (ST-GCN), thus obtaining a one-dimensional feature vector. The two feature vectors are concatted to fuse information from the pixel domain and the time domain, and then fed into a fully connected (FC) system for behavior classification.
[0017] Furthermore, each ST-GCN contains a graph convolutional GCN that captures spatial pose features based on adjacent keypoints and a 1D convolutional TCN in the temporal dimension that fuses pose features from multiple time points to obtain temporal action features.
[0018] An electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the image processing method described above.
[0019] A computer program product includes a computer program or computer instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions to cause the computer device to perform the image processing method described above.
[0020] The beneficial effects of this invention are: This invention combines YOLO's Backbone network with a target / key point detection head to quickly detect the position and pose key points of multiple examinees, without being limited by the size of the examination room, which is especially important for large-scale examinations.
[0021] This invention, through the use of temporal key point caching and spatiotemporal graph convolution, can accurately capture changes in examinees' actions while maintaining low complexity. Combined with pixel-domain features extracted by multi-layer convolution and spatiotemporal feature fusion, it enables precise behavioral analysis. This allows for real-time feedback of the analysis results to invigilators, enabling them to promptly detect and address potential cheating. This improves the efficiency and accuracy of invigilation, reduces reliance on manual invigilation, and thus lowers labor costs.
[0022] The present invention has strong environmental adaptability. Even in examination room environments with obstructions, it can achieve accurate monitoring of examinee behavior through advanced computer vision technology.
[0023] The solution proposed in this invention not only improves the efficiency of invigilation and reduces the workload of invigilators, but also provides candidates with a fairer and more impartial examination environment. Through these innovations, this invention brings meaningful progress to the field of examination invigilation. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the structure of the present invention.
[0025] Figure 2 This is a schematic diagram of the YOLO Backbone network structure of the present invention.
[0026] Figure 3 This is a schematic diagram of the detection head network structure of the present invention.
[0027] Figure 4 This is a schematic diagram of the behavior recognition network structure of the present invention. Detailed Implementation
[0028] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.
[0029] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0030] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0031] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0032] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0033] Implementation Method 1 This embodiment provides a real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, such as... Figure 1 As shown, the system includes a feature extraction module, a detection head, a temporal buffer, and a behavior recognition module; The feature extraction module is connected to both the target detection head and the human pose key point detection head. The target detection head is connected to the ROI extraction module, and the ROI extraction module is connected to the behavior recognition module. The human posture key point detection head is connected to the ROI-guided key point extraction module, which is connected to the behavior recognition module. The key point extraction module is also connected to a temporal buffer, which caches key points extracted from past frame images (7 points per person, so the key points are the coordinates of dozens or hundreds of points). These key points are sent to the ROI-guided key point module, which uses the position information obtained from the ROI module. Since there is one examinee in each bounding box, each key point is mapped to each examinee. The human posture key point detection head is used to detect the key point positions of the head, left shoulder, right shoulder, left elbow, right elbow, left wrist, and right wrist.
[0034] Furthermore, the ROI extraction module extracts pixel blocks based on the position information detected by the target detection head, and then inputs the pixel blocks into the behavior recognition module for abnormal behavior information recognition.
[0035] Furthermore, the ROI-guided key point extraction module takes the human skeletal nodes detected by the human posture key point detection head as key points, namely the key point positions of the head, left shoulder, right shoulder, left elbow, right elbow, left wrist and right wrist, and inputs them into the temporal buffer. The temporal buffer marks the key points with time sequence to form temporal key points, and then inputs the temporal key points and the behavior recognition module to identify abnormal behavior information.
[0036] Furthermore, such as Figure 4 As shown, the image blocks of the behavior recognition module extract pixel domain feature information through multiple convolutional and pooling layers, and the key point sequence of the behavior recognition module extracts spatiotemporal action features through multiple spatiotemporal graph convolution ST-GCN. The two features are fused through a Concat operation and then classified through a fully connected layer to output the abnormal behavior category.
[0037] Feature extraction module. It uses the existing YOLO model's backbone network, with the corresponding network structure as follows: Figure 2 As shown, this module consists of convolutional layers, C3K2, C2PSA, and other network layers, and its function is as a core feature extractor. After inputting a single frame image from the video stream, this module can extract image features, making it a crucial module in the detection stage.
[0038] Detection heads. These include target detection heads and human pose keypoint detection heads, with a network structure as follows: Figure 3 As shown, it mainly consists of convolutional layers and has strong multi-scale prediction capabilities, accurately identifying objects of different sizes in an image. The detection head takes the extracted image features as input and outputs the examinee's location information and key points. The location information is the detection box of the examinee's area, and the key points are the skeletal nodes of the human body. Since the examination process only focuses on the examinee's upper body, a total of 7 key points are detected: head, left shoulder, right shoulder, left elbow, right elbow, left wrist, and right wrist.
[0039] The Region of Interest (ROI) module is used to crop out sub-regions containing only a single examinee after obtaining the location information. Upsampling / downsampling techniques are then used to standardize the image patch size to 32x32. Since the obtained keypoints are the coordinates of all keypoints in the entire frame and cannot be used directly, an ROI-guided keypoint module is introduced. This module maps each keypoint to its corresponding target region based on the location information, effectively classifying all keypoints according to the examinee they belong to.
[0040] Temporal buffer. The temporal buffer stores key points of human pose extracted from the previous 31 frames. Since human behavior changes are a continuous process, the ongoing action cannot be fully confirmed by the image at a single moment; therefore, previous pose information is needed to assist in the analysis.
[0041] Behavior recognition module. Network structure as follows: Figure 4 As shown, the module takes the target pixel block and the temporal keypoint sequence at the current moment as input and outputs the behavioral information of the target examinee. Internally, the module has two pathways: the input 32x32 pixel block undergoes multiple 2D convolutional layers and max-pooling layers, ultimately outputting a one-dimensional feature vector; while the temporal keypoint sequence is processed through multiple ST-GCN (Spatiotemporal-Graph Convolutional) layers to extract spatiotemporal action features, also resulting in a one-dimensional feature vector. Finally, the two feature vectors are concatenated to fuse information from the pixel and temporal domains, and then fed into a fully connected (FC) layer for behavior classification.
[0042] YOLO's Backbone network provides powerful feature extraction capabilities, enabling subsequent detection heads to accurately acquire the examinee's location information and key point coordinates. A temporal buffer stores historical key points, and the 31-frame buffer size does not impose significant memory pressure. The key innovation of this invention lies in its action recognition module, which comprehensively considers both pixel-domain and temporal key point domains. Unlike traditional action recognition networks such as two-stream networks and 3D convolutions, this structure combining graph convolutions and traditional convolutions maintains high accuracy while keeping computationally low, thus meeting real-time requirements.
[0043] In summary, this invention can accurately detect abnormal behavior of examinees in multi-person examination environments, thereby reducing the workload of invigilators.
[0044] Implementation Method 2 To overcome the shortcomings of the YOLO algorithm in utilizing temporal information and avoid the high time complexity of traditional action recognition methods, this scheme introduces a temporal buffer and innovatively presents a low-complexity, high-efficiency action recognition module. This scheme still uses the YOLO network for object detection, but adds an additional detection head after the backbone network to obtain the coordinates of key points in human pose. The action recognition module not only uses pixel blocks of a single examinee for pixel-domain feature extraction (spatial features), but also uses spatiotemporal-graph convolution to analyze key points across multiple frames to capture changes in the examinee's movements (temporal features), and fuses spatiotemporal features to ensure recognition accuracy. Since the temporal information utilized is only key points, the temporal buffer faces less memory pressure, and the number of key points is far less than the number of pixels, resulting in low computational complexity for graph convolution, which meets real-time requirements.
[0045] This embodiment provides a detection method for a real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, such as... Figure 1 As shown, the method uses a real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion as described in Embodiment 1. The detection method includes: acquiring a monitoring screen; the feature extraction module receiving a single-frame image from the video stream; the feature extraction module extracting image features from the single-frame image; and the image features being input to the target detection head and the human posture key point detection head, respectively. The target detection head extracts the target location information from the image features and inputs it into the ROI extraction module; the target detection head only needs to obtain the location of the target (i.e., the bounding box), and the pixel block is extracted from the original image by the ROI module based on the obtained location information; The human pose key point detection head extracts key points containing human skeleton nodes from image features and inputs them into the ROI-guided key point module. The behavior recognition module takes the target pixel block and the sequence of temporal key points at the current moment as input and outputs the behavior information of the target examinee, realizing real-time end-to-end human behavior detection based on deep spatiotemporal feature fusion.
[0046] 1) Feature extraction module. The Backbone network of the existing YOLO model is used, and the corresponding network structure is as follows: Figure 2 As shown, this module consists of convolutional layers, C3K2, C2PSA, and other network layers, and its function is as a core feature extractor. After inputting a single frame image from the video stream, this module can extract image features, making it a crucial module in the detection stage.
[0047] 2) Detection Head. This includes a target detection head and a human pose keypoint detection head. The network structure is as follows: Figure 3 As shown. It mainly consists of convolutional layers and has strong multi-scale prediction capabilities, accurately identifying objects of different sizes in an image. The detection head takes the image features extracted in step 1) as input and outputs the examinee's location information and key points. The location information is the detection box of the examinee's area, and the key points are the skeletal nodes of the human body. Since the examination process only focuses on the examinee's upper body, a total of 7 key points are detected: head, left shoulder, right shoulder, left elbow, right elbow, left wrist, and right wrist.
[0048] 3) ROI Module (Region of Interest). After obtaining the location information, the ROI module crops out a sub-region containing only a single examinee, and uses upsampling / downsampling techniques to unify the image patch size to 32x32. Since the keypoints obtained in 2) are the coordinates of all keypoints in the entire frame and cannot be used directly, an ROI-guided keypoint module is introduced. Based on the location information, each keypoint is mapped to its corresponding target region, that is, all keypoints are classified according to the examinee they belong to.
[0049] 3) Timing Buffer. The timing buffer stores the human pose key points obtained from 3) for the previous 31 frames. Since human behavior changes are a continuous process, the ongoing action cannot be fully confirmed by the image at a single moment; therefore, previous pose information is needed to assist in the analysis. The contents of the timing buffer are continuously replaced. When processing frame T, the buffer stores the key points of frames T-1, T-2, ..., T-31. When processing frame T+1, the contents of the buffer become the key points of frames T, T-1, ..., T-30.
[0050] 4) Behavior recognition module. The network structure is as follows: Figure 4 As shown, the module takes the target pixel block and the temporal keypoint sequence at the current moment as input and outputs the behavioral information of the target examinee. Internally, the module has two pathways: the input 32x32 pixel block undergoes multiple 2D convolutional layers and max-pooling layers, ultimately outputting a one-dimensional feature vector; while the temporal keypoint sequence is processed through multiple ST-GCN (Spatiotemporal-Graph Convolutional) layers to extract spatiotemporal action features, also resulting in a one-dimensional feature vector. Finally, the two feature vectors are concatenated to fuse information from the pixel and temporal domains, and then fed into a fully connected (FC) layer for behavior classification.
[0051] Each ST-GCN contains a GCN (Graph Convolutional Network) that captures spatial pose features based on adjacent keypoints, and a TCN (Temporal Transformation Network) that fuses pose features from multiple time points to obtain temporal motion features. The structure of multiple ST-GCNs cascaded together provides the ability to recognize high-dimensional complex actions.
[0052] YOLO's Backbone network provides powerful feature extraction capabilities, enabling subsequent detection heads to accurately acquire the examinee's location information and key point coordinates. A temporal buffer stores historical key points, and the 31-frame buffer size does not impose significant memory pressure. The key innovation of this invention lies in its action recognition module, which comprehensively considers both pixel-domain and temporal key point domains. Unlike traditional action recognition networks such as two-stream networks and 3D convolutions, this structure combining graph convolutions and traditional convolutions maintains high accuracy while keeping computationally low, thus meeting real-time requirements.
[0053] First, by using the YOLO algorithm, this invention can simultaneously detect multiple examinees, capturing the position and posture information of all examinees in multi-exam scenarios. Second, this invention introduces a temporal feature cache to cache key points of human posture from multiple consecutive frames, thereby utilizing temporal information. Finally, the action recognition module uses the pixel domain information at the current moment and the cached key points to more accurately capture and analyze changes in examinee behavior. This method effectively utilizes the temporal nature of actions, improving the accuracy and reliability of recognition. Therefore, in scenarios such as offline multi-exam examinations where real-time monitoring of multiple examinee behavior is required, this improved algorithm can play a better role, helping invigilators to promptly detect and handle abnormal examinee behavior, ensuring the fairness and impartiality of the examination.
[0054] Implementation Method 3 This invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor. The memory stores software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory and processor are connected via a bus. Specifically, the processor implements any step in Embodiment 1 by running the computer program stored in the memory.
[0055] It should be understood that, in the embodiments of the present invention, the processor may be a Central Processing Unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0056] Memory may include read-only memory, flash memory, and random access memory, and provides instructions and data to the processor. Some or all of the memory may also include non-volatile random access memory.
[0057] It should be understood that if the integrated modules / units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods described above can also be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0058] The above description of the disclosed embodiments enables those skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0059] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0060] It should be noted that the methods and detailed examples provided in the above embodiments can be incorporated into the apparatus and devices provided in the embodiments for mutual reference, and will not be repeated here.
[0061] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0062] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units described above is merely a logical functional division, and in actual implementation, it can be divided in other ways. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0063] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, characterized in that, The system includes a feature extraction module, a detection head, a temporal buffer, and a behavior recognition module; The feature extraction module is connected to both the target detection head and the human pose key point detection head. The target detection head is connected to the ROI extraction module, and the ROI extraction module is connected to the behavior recognition module. The human posture key point detection head is connected to the ROI-guided key point extraction module, which is connected to the behavior recognition module. The key point extraction module is also connected to a temporal buffer. The human posture key point detection head is used to detect the key point positions of the head, left shoulder, right shoulder, left elbow, right elbow, left wrist, and right wrist.
2. The detection system according to claim 1, characterized in that, The ROI extraction module extracts pixel blocks based on the location information detected by the target detection head, and then inputs the pixel blocks into the behavior recognition module for abnormal behavior information recognition.
3. The detection system according to claim 1, characterized in that, The ROI-guided key point extraction module uses the human skeleton nodes detected by the human posture key point detection head as key points and inputs them into the temporal buffer. The temporal buffer marks the key points with time sequence to form temporal key points, and then inputs the temporal key points into the behavior recognition module for abnormal behavior information recognition.
4. A detection method for a real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion, characterized in that, The method uses the real-time end-to-end human behavior detection system based on deep spatiotemporal feature fusion as described in any one of claims 1-3. The detection method includes: acquiring a monitoring screen; the feature extraction module receiving a single-frame image from the video stream; the feature extraction module extracting image features from the single-frame image; and the image features being input to the target detection head and the human posture key point detection head, respectively. The target detection head extracts target location information from image features and inputs it into the ROI extraction module; The human pose key point detection head extracts key points containing human skeleton nodes from image features and inputs them into the ROI-guided key point module. The behavior recognition module takes the target pixel block and the sequence of temporal key points at the current moment as input and outputs the behavior information of the target examinee, realizing real-time end-to-end human behavior detection based on deep spatiotemporal feature fusion.
5. The detection method according to claim 4, characterized in that, The ROI extraction module cropped out a sub-region containing only a single examinee and standardized the size of the image block to 32x32. The ROI-guided keypoint module categorizes all keypoints according to the candidates they belong to.
6. The detection method according to claim 5, characterized in that, Specifically, the timing buffer stores key point information of human skeletal nodes. When the timing buffer processes frame T, it stores key point information of frames T-1, T-2, ..., T-31. When it switches to processing frame T+1, the content in the buffer becomes key point information of frames T, T-1, ..., T-30.
7. The detection method according to claim 4, characterized in that, The behavior recognition module consists of two information processing paths. One path outputs a one-dimensional feature vector by passing multiple 2D convolutional layers and max pooling layers through the input target pixel block. Another approach is to extract spatiotemporal action features from the input temporal keypoint sequence through multiple spatiotemporal-graph convolutions (ST-GCN), thus obtaining a one-dimensional feature vector. The two feature vectors are concatted to fuse information from the pixel domain and the time domain, and then fed into a fully connected (FC) system for behavior classification.
8. The detection method according to claim 7, characterized in that, Each ST-GCN contains a graph convolutional GCN that captures spatial pose features based on adjacent keypoints, and a 1D convolutional TCN that fuses pose features from multiple time points to obtain temporal motion features.
9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the image processing method as described in any one of claims 4 to 8.
10. A computer program product, comprising a computer program or computer instructions, characterized in that, The computer program or the computer instructions are stored in a computer-readable storage medium, the processor of the computer device reads the computer program or the computer instructions from the computer-readable storage medium, and the processor executes the computer program or the computer instructions, causing the computer device to perform the image processing method as described in any one of claims 4 to 8.