Personnel behavior detection method and system

Through the dual-branch convolutional neural network combined with deep learning model, the problem of low recognition accuracy in complex environments is solved, and accurate personnel detection is achieved in the case of insufficient light, occlusion and blind spots in the field of view is improved, and the reliability and security of the security system are improved.

CN120544271APending Publication Date: 2025-08-26CHINA TOWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510627782.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Traditional personnel detection methods have problems with low recognition accuracy and high missed detection rates in complex environments, especially in the absence of light, occlusion and blind spots in the field of view, which are difficult to achieve accurate identification and detection.

Method used

The dual-branch convolutional neural network is used to combine a deep learning model to extract the spatial detail features of the image through the convolutional neural network branches, and the deep learning model branches perform long-distance dependence modeling to achieve accurate identification and detection of personnel behavior.

Benefits of technology

In complex environments, the personnel detection capabilities are effectively improved, ensuring that personnel behavior can be accurately identified under blind spots in the field of vision and equipment occlusion conditions, real-time security protection is provided, false alarm rates and missed rates are reduced, and the reliability of the security system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544271A_ABST
    Figure CN120544271A_ABST
Patent Text Reader

Abstract

The invention discloses a personnel behavior detection method and system, and the method comprises the steps: monitoring personnel in a target region in real time through monitoring equipment, generating a monitoring video in real time, and extracting an original image of a video frame; performing personnel behavior information detection on the original image by using a double-branch convolutional neural network to generate a double-branch detection result; and judging whether the person has a retention or abnormal behavior based on a double-branch detection result, and if so, triggering an alarm mechanism. According to the invention, personnel behaviors can be accurately identified and detected in a complex environment, especially in a visual field blind area and equipment shielding condition, and the personnel detection capability is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for detecting human behavior. Background Art

[0002] In today's rapidly developing information and intelligence environment, personnel detection has become a key technology in security, intelligent monitoring and other fields. Especially in important facilities and high-risk places such as data centers, computer rooms, warehouses, etc., it is crucial to ensure the accurate identification and timely positioning of personnel.

[0003] Traditional personnel detection methods mostly rely on simple motion detection or image recognition technology. These methods often face problems such as insufficient lighting, occlusion, and blind spots in complex environments, resulting in low recognition accuracy and high missed detection rates, affecting the efficiency and effectiveness of security management.

[0004] Therefore, there is an urgent need for a detection method for personnel behavior to accurately identify and detect personnel behavior, so as to achieve early warning of dangerous behaviors of staff, avoid causing dangerous situations, and ensure that dangerous accidents are cut off from the source. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for detecting human behavior, which can accurately identify and detect human behavior in complex environments, especially in blind spots and equipment occlusion conditions, thereby effectively improving the human detection capability.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] In a first aspect, an embodiment of the present invention provides a method for detecting human behavior, comprising:

[0008] Use surveillance equipment to monitor people in the target area in real time, generate surveillance videos in real time and extract the original images of video frames;

[0009] Use a dual-branch convolutional neural network to detect human behavior information in the original image and generate dual-branch detection results;

[0010] Based on the dual-branch detection results, it is determined whether the personnel are staying or engaging in abnormal behavior. If so, an alarm mechanism is triggered.

[0011] In a second aspect, an embodiment of the present invention provides a human behavior detection system, the system comprising:

[0012] A monitoring unit, configured to monitor personnel in a target area in real time using monitoring equipment, generate monitoring videos, and extract original images of video frames;

[0013] An algorithm processing unit, configured to detect human behavior information from the original image using a dual-branch convolutional neural network and generate a dual-branch detection result;

[0014] The alarm intelligent connection unit is used to determine whether the personnel are detained or have abnormal behavior based on the dual-branch detection results. If so, the alarm mechanism is triggered to send an alarm message.

[0015] The technical effects and advantages of the present invention are as follows: The method of the present invention effectively improves the ability to detect people in complex environments, especially in blind spots and equipment occlusion conditions, by integrating the advantages of convolutional neural networks (CNNs) and deep learning models (SwinTransformer models). By utilizing the convolutional neural network branches to extract the spatial detail features of the image, and then using the long-distance dependency modeling capabilities of the deep learning model branches, the ability to extract global features at long distances is comprehensively enhanced. Through this dual-branch network structure, the present invention can more accurately capture personnel information at different scales and spatial levels, ensuring that real-time and effective security protection can be provided in different scenarios.

[0016] The system of the present invention can be used to detect the identities of people in the tower machine room and warn of dangerous behaviors, avoid dangerous situations, and ensure that the occurrence of dangerous accidents is cut off at the source; it is effective for the identification of people in the tower machine room and the regulation of their behavior, and provides optional remote operation and management functions, thereby providing effective protection for the safe operation of the tower machine room; it can effectively identify people and their behaviors in real-time videos, improving the recognition effect of convolutional neural networks on people and the normative constraints on people's behavior.

[0017] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 This is a flow chart of a method for detecting human behavior according to an embodiment of the present invention;

[0020] Figure 2 A flowchart of generating a dual-branch detection result in an embodiment of the present invention;

[0021] Figure 3 Flowchart of a convolutional neural network branch in an embodiment of the present invention;

[0022] Figure 4 This is a flowchart of the deep learning model branching in an embodiment of the present invention;

[0023] Figure 5 The figure is a structural diagram of a personnel behavior detection system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0025] To address the deficiencies of the prior art, the present invention discloses a method for detecting human behavior. Figure 1 As shown, the following steps are included:

[0026] Step S1: Using monitoring equipment to monitor people in the target area in real time, generating monitoring videos and extracting original images of video frames;

[0027] Step S2: Use a dual-branch convolutional neural network to detect the behavior information of the original image and generate a dual-branch detection result;

[0028] Step S3: Based on the dual-branch detection results, determine whether the personnel have any detention or abnormal behavior, and if so, trigger the alarm mechanism.

[0029] In some specific embodiments, step S1: using monitoring equipment to monitor people in a target area in real time, generating a monitoring video and extracting original images of the video frames.

[0030] Among them, the monitoring equipment includes a camera group, which includes a visible light camera and an infrared camera; the visible light camera is responsible for capturing images or videos of people under normal lighting, while the infrared camera is responsible for capturing images or videos of people in low light or night environments.

[0031] This step mainly involves installing a camera group in the target area, using the camera group to monitor the personnel in the target area in real time, generating surveillance video and sending it to the cloud management platform, so that the cloud management platform can call and view the surveillance video; extracting the original image of the video frame based on the generated surveillance video and sending it to the algorithm processing unit for image analysis and personnel detection; that is, mainly using a dual-branch convolutional neural network to fuse the multimodal information from the two shooting devices and perform personnel recognition and detection to improve the accuracy and stability of personnel detection, especially in complex situations such as lighting changes, obstructions and blind spots, it can still accurately identify and track the target personnel.

[0032] In some specific embodiments, step S2: using a dual-branch convolutional neural network to detect the behavior information of the original image, and generating a dual-branch detection result; Figure 2 As shown, including the following:

[0033] The algorithm processing unit includes an input subunit, a feature extraction subunit and an output subunit; wherein the feature extraction subunit includes a dual-branch convolutional neural network.

[0034] Step S21: The input subunit in the algorithm processing unit preprocesses the original image using a bilinear interpolation method to reduce its original size, and then performs normalization to obtain a preprocessed original image. The bilinear interpolation method reduces the image to 1 / 4 of the original image's original size. The input subunit then uses the preprocessed original image as an input image, along with the original image, into a two-branch convolutional neural network.

[0035] Step S22: The dual-branch convolutional neural network uses the convolutional neural network branch and the deep learning model branch to perform parallel feature extraction on the input image and the original image respectively to obtain output features of the dual branches; specifically, the steps include:

[0036] The dual-branch convolutional neural network includes a convolutional neural network branch and a deep learning model branch; the convolutional neural network branch uses a multi-stage convolution structure to extract local detail features in the input image, and at the same time, the deep learning model branch uses a window self-attention mechanism to extract global context features in the original image (that is, the following steps a and b are performed simultaneously), thereby obtaining the output features of the dual branches.

[0037] Step a: The convolutional neural network branch uses a multi-stage convolution structure to extract local detail features in the input image, specifically including:

[0038] like Figure 3As shown, the convolutional neural network branch includes four stages, including stage 1, stage 2, stage 3 and stage 4. An input image of size (4, 1024, 1024) is input to stage 1 of the convolutional neural network branch.

[0039] (1) In Phase 1:

[0040] The convolutional neural network branch uses a conversion layer to preprocess the input image to increase the number of channels and reduce the size of the input image, enhance the ability to capture detailed features of the input image, and generate a first feature map.

[0041] Exemplarily, the conversion layer includes a convolution module layer and a maximum pooling layer connected in series, and the convolution module layer includes two first convolution layers, a first normalization layer, and an activation function. The convolutional neural network branch first uses two 5×5 first convolution layers with a channel number of 64 and a step size of 2 to perform a downsampling convolution operation on the input image to reduce the size of the input image while increasing the number of channels to obtain the convolved input image; the convolution input image is subjected to the first normalization layer and activation function processing using BN (Batch Normalization) and ReLU (Rectified Linear Unit) to obtain a preliminary feature map, and then the maximum pooling layer is used for dimensionality reduction and feature extraction to generate a first feature map of size (64, 256, 256).

[0042] In this example, a relatively large transformation layer is replaced by a concatenated convolutional module layer and a maximum pooling layer to capture more details of the input image, reduce the size of the input image, and stack the number of channels to facilitate subsequent feature extraction. Larger transformation layers usually involve higher computational costs, so by decomposing them into multiple smaller layers, the amount of computation can be reduced while maintaining similar performance. The gradual reduction in the input image size helps the dual-branch convolutional neural network focus on more abstract features, while the stacked number of channels provides more feature representation capabilities.

[0043] (2) During Stages 2 to 4:

[0044] In stages 2, 3, and 4, the convolutional neural network branch uses an improved reverse residual block structure to perform convolution processing on the first feature map in sequence, so that the size of the first feature map gradually decreases, and at the same time, the detailed features in the first feature map are extracted to obtain the local features of the input image.

[0045] Exemplarily, in stage 2, stage 3, and stage 4, respectively, in these three stages, ordinary convolutional layers with different numbers of layers are used to sequentially perform downsampling processing and channel superposition on the first feature map, gradually reducing the size of the first feature map and extracting features to generate a fourth feature map;

[0046] Among them, each ordinary convolution layer is replaced by an alternating combination of a 3×3 depthwise separable convolution layer and a 1×1 pointwise convolution layer with a stride of 2, and the deep separable convolution layer is separated into two parallel 1×3 horizontal convolution layers and a 3×1 vertical convolution layer of different sizes; the convolutional neural network branch uses the depthwise separable convolution layer to perform a convolution operation on the current feature map, and performs channel superposition through the point-by-point convolution layer.

[0047] In this example, the present invention decomposes the traditional ordinary convolution operation into two simpler operations: a depthwise separable convolution layer and a pointwise convolution layer, thereby realizing low-rank branches of parallel operations, which greatly reduces the number of parameters in the convolution calculation, thereby making the dual-branch convolutional neural network model more lightweight, speeding up the inference speed of the convolutional neural network, and effectively capturing the detailed features in the feature image while reducing the amount of calculation. The decomposed deep separable convolution layer enables the convolutional neural network branch to achieve more efficient convolution calculations and further reduces the amount of calculation; at the same time, this decomposition method allows the convolutional neural network branch to perform parallel calculations, effectively improving the operation speed. These convolution layers are fused using a 1×1 pointwise convolution layer, so that image features can be superimposed and optimized during the parallel calculation process.

[0048] like Figure 3 As shown in the figure, the first feature map of size (64, 256, 256) generated in stage 1 is input into stage 2, and the second feature map of size (128, 128, 128) is generated after convolution processing of three layers of ordinary convolutional layers; the second feature map generated in stage 2 is input into stage 3 for the same convolution processing of three layers of ordinary convolutional layers to generate a third feature map of size (256, 64, 64); the third feature map is then input into stage 4, and the fourth feature map of size (513, 32, 32) is generated after convolution processing of six layers of ordinary convolutional layers, thus obtaining the local detail features of the input image.

[0049] In this step, the feature image is further processed in detail using stages 2 to 4 to complete the processing of personnel information confirmation and behavior recognition features in the feature image; as the size of the feature image gradually decreases, the convolutional neural network can handle more complex scenes and realize the recognition of complex behavior patterns through more accurate feature extraction; by combining the refined features of each stage, the convolutional neural network branch can finally output an accurate analysis of the behavior of people in the feature image.

[0050] Step b: The deep learning model branch uses the window self-attention mechanism to extract global context features from the original image, specifically including:

[0051] The deep learning model branch includes multiple stages and each stage has the same structure, such as Figure 4 As shown in the figure, each stage includes a patch embedding layer, a multi-head self-attention module and a patch merging layer; wherein, the feature image size of each stage is reduced by 2 times, the number of channels is increased by 2 times, and the size of the feature image obtained in each stage is shifted by 2C×H / 2×W / 2.

[0052] Taking the first stage as an example, the specific process of each stage is as follows:

[0053] (1) The first stage of patch embedding layer:

[0054] The deep learning model inputs the original image of size (4, 1024, 1024) into the patch embedding layer, which divides the original image into multiple non-overlapping patches and maps each patch into a low-dimensional vector through linear projection to generate a patch embedding sequence; the patch embedding layer inputs the patch embedding sequence into the window-based multi-head self-attention module.

[0055] (2) Multi-head self-attention module in the first stage:

[0056] The multi-head self-attention module divides all patch embedding sequences into multiple non-overlapping windows according to spatial positions, and performs self-attention calculation in each window to capture the interaction relationship of the patches in the window, obtaining the feature results of each window, that is, obtaining the patch feature map of each window; the multi-head self-attention module inputs the feature results of each window into the patch merging layer.

[0057] Exemplarily, the multi-head self-attention module includes a local window multi-head self-attention submodule and a shift window multi-head self-attention submodule; in the multi-head self-attention module, the local window multi-head self-attention submodule and the shift window multi-head self-attention submodule are alternately arranged, and a second normalization layer (i.e., LayerNorm layer) is provided before the local window multi-head self-attention submodule and the shift window multi-head self-attention submodule to standardize features and thereby enhance the stability of the multi-head self-attention module structure; a multi-layer perceptron is provided after the local window multi-head self-attention submodule and the shift window multi-head self-attention submodule, which enables the deep learning model branch to better understand the relationship between features at different positions in the original image;

[0058] Therefore, the patch embedding sequence is first normalized by the second normalization layer and then input into the local window multi-head self-attention submodule (W-MSA) to obtain the normalized patch embedding sequence; the normalized patch embedding sequence is then divided into non-overlapping local windows by the local window multi-head self-attention submodule (each local window includes multiple patches, specifically 64 patches of size 4×4), and the patch embedding sequences of different local windows are obtained; and the relative position information is introduced to perform multi-head self-attention calculation in each local window to capture the interaction relationship of the patches in the local window and generate an updated patch feature sequence in the window;

[0059] The local window multi-head self-attention submodule performs window merging and residual connection processing on the patch feature sequence in the window, and then inputs it into the multi-layer perceptron and the second normalization layer in sequence for processing to obtain the processed patch feature map; the shifted window multi-head self-attention submodule performs cyclic shift on the processed patch feature map so that adjacent windows cover different areas in the next layer to obtain the shifted patch feature map; the shifted window multi-head self-attention submodule divides the window on the shifted patch feature map, and fuses the cross-window features by calculating the mask self-attention, and outputs the feature results of each window; the feature results of each window are then processed by the multi-layer perceptron and input into the patch merging layer.

[0060] The core component of the deep learning model is the multi-head self-attention module, which is mainly used to calculate self-attention and can adaptively learn the relationship between different regions in the original image. As shown below, the local window multi-head self-attention submodule introduces relative position information to perform multi-head self-attention calculations within each local window:

[0061] Using different parameter matrices, the patches in each local window are mapped to different subspaces: Q = XW Q , K=XW K 、V=XW V , where Q represents the query matrix, K represents the key matrix, V represents the value matrix, X represents the patch, and W Q 、W K 、W V Both represent parameter matrices;

[0062] Split Q, K, and V according to the number of heads, and then introduce relative position information to calculate the attention weight of each head: Where QK T It represents the feature of interest obtained by calculating the similarity between Q and K, B represents the relative position deviation, that is, the relative position between patches within a single window; d represents the feature dimension of the patch within each window.

[0063] In this step, the cross-window fusion of patch features is achieved through a combination of window division and shifted window. This design can greatly reduce the computational complexity of the dual-branch neural network model. In addition, this step sets up the alternating use of the local window multi-head self-attention submodule (W-MSA) and the shifted window multi-head self-attention submodule (SW-MSA). This design can achieve the boundary of the window in the original layer network through the self-attention calculation in the new window divided by the next layer of network, thereby providing communication connections between local windows and capturing global context information. Among them, the calculation process of the multi-head self-attention module is as follows:

[0064]

[0065] Where, Represents the output features of the W-MSA submodule, W-MSA represents the local window multi-head self-attention submodule, MSA represents the multi-head attention mechanism, and LayerNorm represents the second normalization layer before the multi-head self-attention submodule; Z W +1 represents the self-attention calculation in the new window divided by the latter layer, which provides connections across local windows across the boundaries of the previous window in the previous layer; M represents the output features of the first multilayer perceptron in each stage, Represents the output feature results of the SW-MSA submodule, Z M+1 Represents the output features of the second multilayer perceptron in each stage.

[0066] (3) The first stage of patch merging layer:

[0067] After the multi-head self-attention module inputs the feature results of each window into the patch merging layer, the patch merging layer merges and splices adjacent patches to generate a (8, 512, 512) first merged feature map and outputs it to the next stage.

[0068] Exemplarily, the patch merging layer includes a patch partitioning sublayer and a linear embedding sublayer. The feature result of each window represents the patch feature map of each window; the patch partitioning sublayer partitions the patch feature map of each window, and splices the adjacent patch feature maps after partitioning to enhance local context information; the patch partitioning sublayer then inputs the spliced ​​patch feature map into the linear embedding sublayer, and the linear embedding sublayer linearly projects the spliced ​​patch feature map to reduce the channel dimension and generate a (8, 512, 512) first merged feature map. The patch merging layer is mainly used to downsample the feature results of each window to reduce the resolution of the feature map and further increase the receptive field. At the same time, it enables the deep learning module to better understand the relationship between features at different positions in the original image and improve the perception of spatial structure.

[0069] When the first phase is over:

[0070] The first merged feature map then undergoes the same processing in subsequent stages, gradually reducing the size of the original image, adding channels, and gradually reducing the feature resolution and extracting features. This generates a second merged feature map of (16, 256, 256), a third merged feature map of (32, 128, 128), and a fourth merged feature map of (64, 64, 64), respectively, generating the global context features of the original image. These four extremes form a hierarchical structure, gradually reducing feature resolution and increasing the receptive field to represent the interaction of long-range features to obtain global information about the original image.

[0071] Step S23: performing feature concatenation and feature complementary enhancement on the output features of the dual branches to obtain dual-branch detection results.

[0072] For example, the present invention utilizes a complementary feature enhancement module to combine the output features of a convolutional neural network branch and a deep learning model branch. This module combines the local detail features of the input image with the global contextual features of the original image. Parallel computing further enhances feature refinement and enhancement, ultimately yielding a dual-branch detection result. The output subunit then performs classification output, localization output, and confidence scoring on the dual-branch detection results.

[0073] In the present invention, CNN and Swin Transformer are important components of a dual-branch convolutional neural network, which is divided into two parallel branches; the CNN branch (i.e., the convolutional neural network branch) is responsible for extracting local features and capturing details of the input image; and the Swin Transformer branch (i.e., the deep learning model branch) is responsible for modeling long-distance dependencies to comprehensively enhance the accurate detection of human targets in the original image, especially when dealing with occluded, long-distance, or dynamically changing targets. The global modeling capability of Swin Transformer can significantly improve the detection accuracy; then, by refining and enhancing the output features of the two branches, the detection accuracy of small targets such as stranded personnel and abnormal behaviors can be further improved, providing reliable real-time monitoring and security protection for computer rooms or data centers, and ultimately obtaining dual-branch detection results.

[0074] Step S3: Based on the dual-branch detection results, determine whether the personnel have any detention or abnormal behavior, and if so, trigger the alarm mechanism.

[0075] For example, the alarm intelligent connection unit is used to quickly analyze and identify the dual-branch detection results and their classification outputs, positioning outputs, and confidence scores, and accurately detect potential dangerous behaviors of people in the image (such as climbing over fences, illegally entering restricted areas, etc.). Once the alarm intelligent connection unit detects that a person in the image has engaged in dangerous behavior, it quickly triggers an alarm to notify the relevant personnel, and at the same time, sends it to the management personnel to record the behavior. This mechanism not only improves the response speed of personnel detection, but also reduces false alarms triggered by misjudgments, ensuring that management personnel can receive accurate warnings in a timely manner and quickly take measures to deal with possible safety hazards. Among them, general alarm methods include sound reminders and mobile phone notifications. Among them, the alarm information of the alarm intelligent connection unit is mainly received through the cloud management platform and sent to the manager account for timely early warning notification, and the alarm event is analyzed based on the alarm information, and the alarm information is stored in the database.

[0076] The cloud management platform provides an intuitive operation interface for remote operation and maintenance and management personnel, and supports real-time monitoring. The surveillance videos generated by all camera groups can be uploaded to the cloud management platform for processing in real time, that is, the surveillance video images can be viewed remotely in real time through the cloud management platform; the cloud management platform also supports remote configuration and location of monitoring equipment (including but not limited to camera groups), such as camera group angle adjustment and device restart, and the cloud management platform can also monitor the device status of the camera group in real time to ensure the safe and stable operation between each module.

[0077] In the present invention, through the cloud management platform, users can fully grasp the personnel dynamics of the computer room or data center, realize remote viewing and management, and improve overall security and operational efficiency.

[0078] In this embodiment, the present invention effectively improves the ability to detect people in complex environments, especially under blind spots and equipment occlusion conditions, by integrating the advantages of convolutional neural networks (i.e., CNN) and deep learning models (i.e., SwinTransformer models). By utilizing the convolutional neural network branches to extract the spatial detail features of the image, and then through the long-distance dependency modeling capabilities of the deep learning model branches, the ability to extract global features at long distances is comprehensively enhanced. That is, through this dual-branch network structure, the present invention can more accurately capture personnel information at different scales and spatial levels, ensuring that real-time and effective security protection can be provided in different scenarios.

[0079] The present invention uses a dual-branch convolutional neural network to efficiently fuse multimodal information from two camera devices and perform personnel identification and detection to improve the accuracy and stability of personnel detection, especially in complex situations such as lighting changes, obstructions, and blind spots in the field of view, it can still accurately identify and track target personnel. The present invention not only has efficient real-time processing capabilities, but also can issue alarms in time when abnormal situations such as personnel entry, exit, and detention occur to ensure the timeliness and accuracy of security management; compared with traditional methods, it significantly improves the accuracy and response speed of personnel detection, reduces the false alarm rate and missed alarm rate, and enhances the reliability of the security system. In addition, the method of the present invention performs particularly well in dealing with occlusion or blind spot problems in complex environments, and can effectively handle small target detection to ensure that personnel safety monitoring is not interfered with.

[0080] The embodiment of the present invention discloses a personnel behavior detection system, such as Figure 5 As shown, the system includes:

[0081] A monitoring unit, which is used to monitor personnel in a target area in real time using monitoring equipment, generate monitoring videos in real time, and extract original images of video frames;

[0082] An algorithm processing unit, configured to detect human behavior information from the original image using a dual-branch convolutional neural network and generate a dual-branch detection result;

[0083] The alarm intelligent connection unit is used to determine whether the personnel are detained or have abnormal behavior based on the dual-branch detection results. If so, the alarm mechanism is triggered to send an alarm message.

[0084] Exemplarily, the monitoring equipment includes a camera group, which is mainly formed by installing a camera group in the target area to form a monitoring unit, and at the same time using the camera group to monitor the personnel in the target area in real time, generate monitoring video and send it to the cloud management platform, so that the cloud management platform can call and view the monitoring video; and the monitoring unit also extracts the original image of the video frame based on the generated monitoring video and sends it to the algorithm processing unit, so that the algorithm processing unit can perform image analysis and personnel detection.

[0085] The algorithm processing unit includes an input subunit, a feature extraction subunit, and an output subunit; wherein the feature extraction subunit includes a dual-branch convolutional neural network. The input subunit performs image preprocessing on the original image by a bilinear interpolation method to reduce the original size, and then performs normalization processing to obtain the preprocessed original image; the input subunit then uses the preprocessed original image as the input image and inputs it together with the original image into the feature extraction subunit (i.e., the dual-branch convolutional neural network); the feature extraction subunit extracts features from the input image and the original image respectively, correspondingly generates local detail features of the input image and global context features of the original image, and performs feature splicing and feature complementary enhancement on the local detail features and the global context features, finally obtaining the dual-branch detection results and inputting them into the output subunit; the output subunit performs classification output, positioning output, and confidence scoring on the dual-branch detection results.

[0086] The algorithm processing unit outputs the unit results to the alarm quality unit, and the alarm intelligent connection unit quickly analyzes and identifies the dual-branch detection results and their classification output, positioning output and confidence score, and accurately detects potential dangerous behaviors of people in the image (such as climbing over fences, illegally entering restricted areas, etc.). Once the alarm intelligent connection unit detects dangerous behaviors of people in the image, it quickly triggers an alarm to notify relevant personnel.

[0087] In some specific embodiments, the system further includes a cloud management platform.

[0088] The cloud management platform is used to remotely view surveillance video images and receive alarm information, and analyze alarm events based on the alarm information; it is also used to monitor the device status of the camera group in real time, and support remote configuration and maintenance of the camera group.

[0089] For example, the cloud management platform serves as the management core of the entire system, providing centralized control and data analysis functions. The cloud management platform provides an intuitive operating interface for remote operation and maintenance and management personnel, and supports real-time monitoring. The monitoring videos generated by all camera groups can be uploaded to the cloud management platform in real time for processing, that is, the monitoring video images can be viewed remotely and in real time through the cloud management platform. The cloud management platform also supports remote configuration and location of monitoring equipment (including but not limited to camera groups), such as camera group angle adjustment and device restart, and the cloud management platform can also monitor the device status of the camera group in real time to ensure the safety and stable operation of each module. In the present invention, through the cloud management platform, users can fully grasp the personnel dynamics of the computer room or data center, realize remote viewing and management, and improve overall security and operational efficiency.

[0090] In some specific embodiments, redundant design and permission management functions can be added to the above-mentioned detection system to ensure that the detection system can still work normally under different network levels or hardware failures, and prevent unauthorized data access and operation; by adding cloud storage, the cloud management platform can query historical data and perform equipment statistical analysis to help managers make decisions; thereby realizing the system's efficient data storage and query capabilities, ensuring the security of operations and audit functions.

[0091] In this embodiment, the system provided by the present invention can effectively identify and regulate the behavior of personnel in the tower machine room, and provide optional remote operation and maintenance and management functions; it provides effective protection for the safety of the tower machine room, can detect the safety of the personnel and behavior in the machine room in real time, avoid dangerous abnormal accidents, and provide intelligent protection for the normal operation of the tower; it can effectively identify people and their behaviors in real-time videos, improve the convolutional neural network's recognition effect on people and the normative constraints on people's behavior.

[0092] Regarding the system in the above embodiment, the specific manner in which each unit module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0093] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for detecting human behavior, characterized in that: include: Use surveillance equipment to monitor people in the target area in real time, generate surveillance videos in real time and extract the original images of video frames; Use a dual-branch convolutional neural network to detect human behavior information in the original image and generate dual-branch detection results; Based on the dual-branch detection results, it is determined whether the personnel are staying or engaging in abnormal behavior. If so, an alarm mechanism is triggered.

2. A personnel behavior detection method according to claim 1, characterized in that: A dual-branch convolutional neural network is used to detect human behavior information in the original image and generate dual-branch detection results, including: The dual-branch convolutional neural network includes a parallel convolutional neural network branch and a deep learning model branch; Performing image preprocessing on the original image to reduce the original size to obtain a preprocessed original image, and using the preprocessed original image as an input image; Inputting the input image and the original image into a dual-branch convolutional neural network, the dual-branch convolutional neural network uses a convolutional neural network branch and a deep learning model branch to perform parallel feature extraction on the input image and the original image respectively, to obtain dual-branch output features; The output features of the two branches are spliced ​​and enhanced complementaryly to obtain the two-branch detection results.

3. A method for detecting human behavior according to claim 2, characterized in that: The dual-branch convolutional neural network uses a convolutional neural network branch and a deep learning model branch to perform parallel feature extraction on the input image and the original image respectively, and obtains output features of the dual branches, including: The convolutional neural network branch uses a multi-stage convolution structure to extract local detail features in the input image, while the deep learning model branch uses a window self-attention mechanism to extract global context features in the original image, thereby obtaining the output features of the dual branches.

4. A method for detecting human behavior according to claim 2, characterized in that: The convolutional neural network branch uses a multi-stage convolution structure to extract local detail features in the input image, including: The convolutional neural network branch includes four stages, including stage 1, stage 2, stage 3 and stage 4; In stage 1, the convolutional neural network branch preprocesses the input image using a conversion layer to reduce the size of the input image and generate a first feature map; The convolutional neural network branch uses an improved reverse residual block structure to perform convolution processing on the first feature map in stages 2, 3, and 4, so that the size of the first feature map gradually decreases, and at the same time extracts the detail features in the first feature map to obtain local detail features of the input image.

5. A method for detecting human behavior according to claim 4, characterized in that: The input image is preprocessed using a conversion layer to reduce the size of the input image and generate a first feature map, including: The conversion layer includes a convolution module layer and a maximum pooling layer connected in series, and the convolution module layer includes a first convolution layer, a first normalization layer and an activation function; In stage 1, the convolutional neural network branch uses the first convolutional layer to perform convolution processing on the input image to downsample and increase the number of channels to obtain the convolved input image; The convolved input image is processed by the first normalization layer and activation function, and then passes through the maximum pooling layer for dimensionality reduction and feature extraction to generate the first feature map.

6. A method for detecting human behavior according to claim 4, characterized in that: The improved reverse residual block structure is used to perform convolution processing on the first feature map in sequence, so that the size of the first feature map gradually becomes smaller, and the detail features in the first feature map are extracted at the same time to obtain the local detail features of the input image, including: In stages 2, 3, and 4, different numbers of ordinary convolutional layers are used to sequentially downsample and superimpose the first feature map, gradually reducing its size and extracting features to generate a fourth feature map. Among them, each ordinary convolution layer is replaced by an alternating combination of a 3×3 depth-wise separable convolution layer and a 1×1 point-wise convolution layer with a stride of 2, and the deep separable convolution layer is separated into two parallel 1×3 horizontal convolution layers and a 3×1 vertical convolution layer of different sizes; The convolutional neural network branch uses a depth-wise separable convolution layer to perform a convolution operation on the current feature map, and performs channel number superposition through the point-by-point convolution layer.

7. A method for detecting human behavior according to claim 3, characterized in that: The deep learning model branch uses the window self-attention mechanism to extract global context features from the original image, including: The deep learning model branch includes multiple stages and each stage has the same structure, each stage includes a patch embedding layer, a multi-head self-attention module and a patch merging layer; In the first stage, the deep learning model inputs the original image into a patch embedding layer, which segments the original image into multiple non-overlapping patches and generates a patch embedding sequence by mapping each patch into a low-dimensional vector through linear projection; The patch embedding layer inputs the patch embedding sequence into the window-based multi-head self-attention module, which divides all patch embedding sequences into multiple non-overlapping windows according to spatial positions and performs self-attention calculation in each window to obtain the feature results of each window; The multi-head self-attention module inputs the feature results of each window into the patch merging layer, which merges and splices adjacent patches to generate a first merged feature map and outputs it to the next stage; The first merged feature map is processed in the same way in the following multiple stages to finally generate the global context feature of the original image.

8. A method for detecting human behavior according to claim 7, characterized in that: The multi-head self-attention module includes a local window multi-head self-attention submodule and a shift window multi-head self-attention submodule; In the multi-head self-attention module, the local window multi-head self-attention sub-module and the shift window multi-head self-attention sub-module are alternately arranged, and a second normalization layer is arranged before the local window multi-head self-attention sub-module and the shift window multi-head self-attention sub-module, and a multi-layer perceptron is arranged after them.

9. A method for detecting human behavior according to claim 8, characterized in that: The multi-head self-attention module divides all patch embedding sequences into multiple non-overlapping windows according to spatial positions, and performs self-attention calculations in each window to obtain the feature results of each window, including: The patch embedding sequence is normalized by a second normalization layer and then input into the local window multi-head self-attention submodule; The local window multi-head self-attention submodule divides the normalized patch embedding sequence into non-overlapping local windows, and introduces relative position information to perform multi-head self-attention calculation in each window to generate an updated patch feature sequence in the window; wherein each local window includes multiple patches; The local window multi-head self-attention submodule performs window merging and residual connection processing on the patch feature sequence in the window, and sequentially inputs the processed patch feature map into the multi-layer perceptron and the second normalization layer for processing; The shifted window multi-head self-attention submodule cyclically shifts the processed patch feature map to obtain a shifted patch feature map; the shifted window multi-head self-attention submodule divides the window on the shifted patch feature map, and fuses cross-window features by calculating masked self-attention, and outputs a feature result for each window; The feature results of each window are processed by the multi-layer perceptron and input into the patch merging layer.

10. A method for detecting human behavior according to claim 9, characterized in that: The multi-head self-attention module inputs the feature results of each window into the patch merging layer, which merges and splices adjacent patches to generate a first merged feature map and outputs it to the next stage, including: The patch merging layer includes a patch partitioning sublayer and a linear embedding sublayer, and the feature result of each window represents a patch feature map of each window; The patch feature map of each window is partitioned through the patch partitioning sublayer, and the adjacent patch feature maps after partitioning are spliced ​​to enhance the local context information; The patch partitioning sublayer inputs the spliced ​​patch feature map into the linear embedding sublayer, and the linear embedding sublayer performs linear projection on the spliced ​​patch feature map to reduce the channel dimension and generate a first merged feature map.

11. A personnel behavior detection system, characterized in that: The system comprises: A monitoring unit, which is used to monitor personnel in a target area in real time using monitoring equipment, generate monitoring videos in real time, and extract original images of video frames; An algorithm processing unit, configured to detect human behavior information from the original image using a dual-branch convolutional neural network and generate a dual-branch detection result; The alarm intelligent connection unit is used to determine whether the personnel are detained or have abnormal behavior based on the dual-branch detection results. If so, the alarm mechanism is triggered to send an alarm message.

12. A personnel behavior detection system according to claim 11, characterized in that: The system also includes a cloud management platform, The cloud management platform is used to remotely view surveillance video images and receive alarm information, and analyze alarm events based on the alarm information; it is also used to monitor the device status of the camera group in real time and support remote configuration and maintenance of the camera group; Among them, the monitoring equipment includes a camera group.