A behavior detection method, device, equipment and storage medium thereof

By improving the feature extraction and fusion module and the target tracking module, and combining them with the Yolov4 algorithm, the problems of high computational load and low detection accuracy in the existing technology have been solved. This enables real-time safety supervision of staff during the 5G base station survey and acceptance process, improves the accuracy and robustness of detection, and ensures the safety of the project implementation process.

CN116778214BActive Publication Date: 2026-04-07CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing real-time security detection methods are computationally intensive and cannot meet real-time requirements, and the accuracy of detection results is low, especially in the robustness and accuracy of moving target detection in video sequences.

Method used

By embedding the feature information of the template image into the image to be searched through the improved feature extraction and fusion module, and combining the target tracking module and the behavior safety classification module, the YOLOv4 algorithm is used for real-time target tracking and dangerous behavior determination. Anti-aliasing pooling, feature channel selection enhancement and multi-scale fusion techniques are used to improve detection accuracy.

Benefits of technology

Without compromising real-time performance, the accuracy and robustness of detection have been improved, enabling timely detection of dangerous behaviors and issuing warnings to ensure the safety of staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778214B_ABST
    Figure CN116778214B_ABST
Patent Text Reader

Abstract

This disclosure provides a behavior detection method, apparatus, device, and storage medium. The method includes: embedding feature information of a template image into feature information of a search image using an improved feature extraction and fusion module to obtain fused feature information; wherein the template image is an image taken by a worker before the implementation of a project, and the search image is each frame of a real-time monitoring video of the worker; determining the region of interest (ROI) of the worker in the search image based on the fused feature information using a target tracking module; and determining whether the worker's behavior includes dangerous behavior based on the ROI using a behavior safety classification module. Thus, the method provided by this application is more conducive to target localization, thereby enabling hazard warnings to relevant workers; and it also improves the accuracy and robustness of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image target detection, and in particular to a behavior detection method, apparatus, device and storage medium thereof. Background Technology

[0002] Real-time security detection methods are a major research hotspot in the fields of communication engineering security and image target detection. To improve the accuracy of real-time security detection, information from the work site can be obtained through video surveillance, and feature fusion can be performed on video images and template images. This allows for accurate and real-time detection, enabling timely detection of dangerous behaviors and issuing warnings, thereby improving the safety of on-site personnel.

[0003] Current real-time security detection methods cannot meet real-time requirements due to their large computational demands, and the accuracy of the detection results is also low. Summary of the Invention

[0004] In view of this, embodiments of this application provide a behavior detection method, apparatus, device, and storage medium thereof.

[0005] The technical solution of this application is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a behavior detection method, the method comprising: embedding feature information of a template image into feature information of a search image through an improved feature extraction and fusion module to obtain fused feature information; wherein the template image is an image taken by a worker before the implementation of a project, and the search image is each frame of a real-time monitoring video of the worker; determining the region of interest of the worker in the search image based on the fused feature information through a target tracking module; and determining whether the worker's behavior includes dangerous behavior based on the region of interest of the worker through a behavior safety classification module.

[0007] Secondly, embodiments of this application also provide a behavior detection device, the device comprising:

[0008] An improved feature extraction and fusion module is used to embed the feature information of a template image into the feature information of the image to be searched to obtain fused feature information; wherein, the template image is an image taken by the staff before the project is implemented, and the image to be searched is each frame of the real-time monitoring video of the staff.

[0009] The target tracking module is used to determine the region of interest of the staff member in the image to be searched based on the fused feature information;

[0010] The behavior safety classification module is used to determine whether the behavior of the staff member includes dangerous behavior based on the staff member's region of interest.

[0011] Thirdly, embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the steps in the above-described method.

[0012] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.

[0013] In this embodiment, firstly, an improved feature extraction and fusion module embeds the feature information of the template image into the feature information of the image to be searched, obtaining fused feature information. The template image is an image taken by workers before the project is implemented, and the image to be searched is each frame of the workers' real-time monitoring video. Secondly, a target tracking module determines the region of interest (ROI) of the workers in the image to be searched based on the fused feature information. Finally, a behavior safety classification module determines whether the workers' behavior includes dangerous behavior based on the ROI. As can be seen, by fusing the features of the image taken before implementation as the template image and the image to be detected, it is more conducive to target localization, thus enabling hazard warnings to relevant workers. Furthermore, it solves the problems of target-background imbalance and network sensitivity to input changes, thereby improving detection accuracy and robustness without affecting real-time performance.

[0014] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0016] Figure 1 A schematic diagram illustrating the implementation process of a behavior detection method provided in an embodiment of this application;

[0017] Figure 2 A schematic diagram illustrating the overall implementation process of a behavior detection method provided in this application embodiment;

[0018] Figure 3 A schematic diagram illustrating the implementation process of an improved feature extraction and fusion module provided in an embodiment of this application;

[0019] Figure 4 A comparative diagram showing the improvement of the maximum pooling operation before and after the implementation of this application.

[0020] Figure 5 A schematic diagram illustrating the implementation process of a behavior safety classification module provided in this application embodiment;

[0021] Figure 6 for Figure 5 A schematic diagram illustrating the implementation process of the improved global context module;

[0022] Figure 7 for Figure 6 A schematic diagram illustrating the implementation process of the mid-channel attention module;

[0023] Figure 8 This is a schematic diagram of the composition structure of a behavior detection device provided in an embodiment of this application;

[0024] Figure 9 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation

[0025] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.

[0026] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0027] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. It should also be noted that the terms "first, second, third" used in the embodiments of this application are merely for distinguishing similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0028] Real-time security detection methods are a major research hotspot in the fields of communication engineering security and image target detection. With the increasing sophistication of modern image acquisition and digital video technologies, to improve the accuracy of real-time security detection, information from the work site can be obtained through video surveillance. Feature fusion can then be performed on video images and template images, enabling accurate and real-time detection to promptly identify dangerous behaviors and issue warnings, greatly improving the safety of on-site personnel during surveying and acceptance processes.

[0029] In recent years, research on deep learning-based object detection and tracking algorithms has been booming. Mature algorithms not only free up significant human and material resources but also offer advantages in accuracy and processing speed. Most current real-time security detection methods are based on the OPENPOSE network. While this network boasts good robustness and accuracy, its computational complexity is prohibitive and cannot meet real-time requirements. Meanwhile, existing image recognition models based on the YOLOv3 object detection algorithm take frame-by-frame images as input, but the output results are numerous and complex. The extremely high training difficulty of these models leads to low detection accuracy. Although existing algorithms are relatively mature, their robustness, accuracy, and real-time performance in detecting moving targets in video sequences still need improvement.

[0030] The Faster R-CNN algorithm, which uses a convolutional neural network to extract features from template images, performs the following steps: First, features are extracted from the template image using shared convolutional layers. Second, the extracted features are fed into a Region Proposal Network (RPN), which generates bounding boxes, specifies the location of regions of interest (ROIs), and performs a first refinement of the RPI bounding boxes. Then, the RPI pooling layer selects features corresponding to each RPI from the feature map based on the RPN output, and sets the dimension to a fixed value. Finally, fully connected layers are used to classify the bounding boxes and perform a second refinement of the target bounding boxes.

[0031] Therefore, this application provides a behavior detection method, which can be executed by the processor of a computer device. The computer device can refer to a server, laptop, tablet, desktop computer, or other device with data processing capabilities. This method fuses feature information from a template image and a search image, determines the region of interest (ROI) of the worker in the search image using a target tracking module, and inputs the ROI into a behavior safety classification module to determine whether the worker's behavior includes dangerous behavior. By fusing features from a pre-action image as the template image and the image to be detected, it is more conducive to target localization, thus enabling hazard warnings to relevant workers. Furthermore, it solves the problems of target-background imbalance and network sensitivity to input changes, thereby improving detection accuracy and robustness without affecting real-time performance.

[0032] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0033] In view of this, embodiments of this application provide a behavior detection method, referring to... Figure 1 The method may include steps S101 to S103, wherein:

[0034] Step S101: Through the improved feature extraction and fusion module, the feature information of the template image is embedded into the feature information of the image to be searched to obtain fused feature information; wherein, the template image is an image taken by the staff before the implementation of the project, and the image to be searched is each frame of the real-time monitoring video of the staff.

[0035] Here, the template image is an image of a worker wearing personal protective equipment taken with a mobile terminal device before the project is implemented; the image to be searched is a video image of each frame of the worker's implementation process taken with a camera during the project.

[0036] Step S102: Using the target tracking module, the region of interest of the staff in the image to be searched is determined based on the fused feature information.

[0037] Here, the target tracking module can employ target tracking algorithms, such as the YOLOv4 algorithm, during implementation. Based on the fused feature information obtained from the improved feature fusion module, the YOLOv4 algorithm can more accurately regress the area where the worker is located in each frame of the image.

[0038] Step S103: Using the behavior safety classification module, determine whether the employee's behavior includes dangerous behavior based on the employee's region of interest.

[0039] Here, the behavior safety classification module performs binary classification on the region of interest of the staff in each frame of the regressed image. When the output result of the behavior safety classification module is greater than the set safety threshold, it is judged as dangerous behavior; otherwise, it is judged as safe behavior.

[0040] In this embodiment, firstly, the feature information of the template image is embedded into the feature information of the search image through an improved feature extraction and fusion module to obtain fused feature information. The template image is an image taken by the staff before the project is implemented, and the search image is each frame of the real-time monitoring video of the staff. Secondly, the target tracking module determines the region of interest (ROI) of the staff in the search image based on the fused feature information. Finally, the behavior safety classification module determines whether the staff's behavior includes dangerous behavior based on the ROI. As can be seen from the above, because target tracking in video has high real-time requirements, the target tracking module in this embodiment uses the YOLOv4 algorithm. Compared to the Faster R-CNN algorithm, the YOLOv4 algorithm has the highest accuracy among real-time target detection algorithms, achieving the best balance between accuracy and speed.

[0041] The following will refer to Figures 2 to 7 This application provides a detailed description of a behavior detection method based on its embodiments.

[0042] A behavior detection method is provided based on steps S101 to S103. To facilitate understanding of the embodiments of this application, a specific scenario is used as an example below, combined with... Figure 2 This paper introduces the overall implementation process of a behavior detection method provided in the embodiments of this application.

[0043] The overall implementation process includes two parts: pre-implementation and implementation. Pre-implementation primarily involves determining whether the wearing of safety equipment is compliant, while implementation mainly involves real-time monitoring of staff and determining whether their actions are safe. For example... Figure 2 As shown, the overall implementation process can include the following two aspects:

[0044] a) Before implementation, the personal safety equipment wearing rationality detection model 201 is used to extract features from the template image and output the results. The output results are compared with the detection boxes with a set threshold. If the output results are greater than the detection boxes with a set threshold, it is determined that the personal safety equipment is worn in compliance and subsequent work can be carried out. Otherwise, a reminder is issued and subsequent work cannot be carried out. The detection boxes with a set threshold can include: safety helmets, safety clothing, safety shoes, eye protection glasses, anti-static wrist straps, and other safety equipment.

[0045] b) In implementation, firstly, the improved feature extraction and fusion module in the worker target tracking model 202 embeds the feature information of the template image into the feature information of the image to be searched to obtain fused feature information; wherein, the template image is an image taken of the worker before the implementation of the project, and the image to be searched is a frame of an image in the real-time monitoring video of the worker; through the target tracking module, based on the fused feature information, the region of interest of the worker in the image to be searched is determined; through the behavior safety classification module 203, based on the region of interest of the worker, it is determined whether the worker's behavior includes dangerous behavior.

[0046] refer to Figure 3 The improved feature extraction and fusion module may include: a shared convolutional layer 301, an anti-aliasing pooling module 302, a feature channel selection enhancement module 303, and an attention module 304. Through the improved feature extraction and fusion module, the feature information of the template image is embedded into the feature information of the image to be searched to obtain fused feature information. This can be achieved through steps S201 to S204, wherein:

[0047] Step S201: Feature extraction is performed on the template image and the image to be searched through a shared convolutional layer to obtain the global feature map of the template image and the global feature map of the image to be searched.

[0048] Here, the template image is an image of a worker wearing personal protective equipment taken with a mobile terminal device before the project is implemented; the image to be searched is a video image of each frame of the worker's implementation process taken with a camera on a fixed device during the project implementation.

[0049] Step S202: The global feature map of the target image is downsampled using the anti-aliasing pooling module to obtain the downsampled feature map of the target image;

[0050] Here, the target image may include either a template image or the image to be searched. That is, the processing operation performed on the template image by the anti-aliasing pooling module is the same as the processing operation performed on the image to be searched. The anti-aliasing pooling module can improve the impact of image target offset caused by shooting and other reasons on the network output.

[0051] Step S203: The feature map of the downsampled target image is recalibrated by the feature channel selection enhancement module to obtain the feature information of the target image;

[0052] Here, the feature channel selection enhancement module can address issues such as the large difference in the proportion of workers and background in the target image, leading to higher detection difficulty. The purpose of feature channel selection enhancement is to assign different weights to the features of different channels in the feature map, thereby increasing the network's focus on useful information (i.e., the worker target).

[0053] Step S204: The feature information of the template image and the feature information of the image to be searched are fused through the attention module to obtain fused feature information; wherein, the fused feature information contains useful information in the template image.

[0054] Here, integrating feature information can better identify the target location personnel and improve the accuracy of subsequent target tracking.

[0055] refer to Figure 4 The global feature map of the target image is downsampled by the anti-aliasing pooling module 302 to obtain the downsampled feature map of the target image. This can be achieved through steps S211 and S212, wherein:

[0056] Step S211: Select the maximum pixel value from the global feature map of the target image to obtain the feature map of the maximum pixel value of the target image;

[0057] Step S212: Downsample the feature map of the maximum value of the target image to obtain the downsampled feature map of the target image.

[0058] Here, during the feature extraction process of the image to be searched, the staff need to move during the project, and the image to be searched and the template image will also be shifted due to shooting methods and other reasons. The shift of the target in these images will aggravate the impact on the network output. The reason is that the convolutional neural network does not satisfy the sampling theorem during downsampling and ignores the signal aliasing. Therefore, even a small displacement will completely change the network output.

[0059] Because max pooling lacks anti-aliasing capabilities, even small changes in the input can have a significant impact on the network output. To reduce the influence of target offset, an improved max pooling operation is employed (see reference). Figure 4 b); where the improved max pooling operation is based on the max pooling operation (see reference). Figure 4 a) A low-pass filter is introduced between the two steps of selecting the dense maximum value and downsampling; the low-pass filter is used to eliminate the situation where the sampling theorem is not satisfied due to high-frequency signals.

[0060] Continue to refer to Figure 3 The feature map of the downsampled target image is recalibrated by the feature channel selection enhancement module 303 to obtain the feature information of the target image. This can be achieved through steps S221 to S224, wherein:

[0061] Step S221: The feature map of the target image downsampled is processed by convolution and pooling through two convolutional layers and a first global pooling layer respectively;

[0062] Step S222: Reduce the dimensionality of the features after convolution and pooling through the first fully connected layer;

[0063] Step S223: The features after dimensionality reduction are restored to their original dimensions by sequentially passing through the first activation layer and the second fully connected layer to obtain the feature vector of the target image;

[0064] Step S224: Multiply the feature vector of the target image with the feature vector of the target image after passing through two convolutional layers, and superimpose it with the downsampled feature map to obtain the feature information of the target image.

[0065] Since the tracked worker target occupies a small portion of the overall image, the feature channel selection enhancement module is used to address the foreground (worker) to background imbalance problem in segmentation. The structure of the feature channel selection enhancement module is as follows: Figure 3 As shown in the dashed box, the structure consists of two convolutional layers, one global pooling layer, and two fully connected layers, containing two branches. The first branch uses direct connections of the residual structure, and the second branch is used to recalibrate the downsampled feature map after the two convolutions.

[0066] The feature map downsampled from the target image is sequentially convolved through two convolutional layers. This increases the receptive field of the feature map and adds non-linear features through non-linear activation layers. After the two convolutional processes, the feature map is recalibrated. Assuming the size of the downsampled feature map is c×w×h, after a global pooling layer, each two-dimensional downsampled feature map becomes a real number. This real number, to some extent, possesses a global receptive field, so the downsampled feature map becomes c×1×1, achieving a global distribution of the response across the feature map channels. This also allows subsequent fully connected layers to acquire a global receptive field. Next, the first fully connected layer reduces the feature dimension after convolution and pooling to (c / 8)×1×1. After the first sigmoid activation layer, a second fully connected layer restores the reduced-dimensional features to their original dimensions, resulting in the feature vector of the target image. Using two fully connected layers, compared to a single fully connected layer, allows for the addition of more non-linear features, thus fitting complex correlations between channels, significantly reducing the number of connections, and improving computational speed. Then, the feature vector of the target image is multiplied by the feature vector of the target image after passing through two convolutional layers, and the feature maps of the two branches are superimposed to obtain the feature information of the target image. This process keeps the network in an optimal state, so that the network performance does not decrease with increasing depth.

[0067] Continue to refer to Figure 3The improved feature extraction and fusion module includes, in sequence: a shared convolutional layer 301, an initial anti-aliasing pooling module 302, a feature channel selection enhancement module 303, and a three-layer repetitive processing module. Each repetitive processing module includes: an anti-aliasing pooling module, a feature channel selection enhancement module, and an attention module. The improved feature extraction and fusion module also includes a multi-scale fusion module 305.

[0068] The feature map with reduced resolution is processed by a multi-scale fusion module through convolutional, deconvolutional and convolutional layers to obtain a feature map with increased resolution. The feature map with reduced resolution is obtained by the anti-aliasing pooling module and the feature channel selection enhancement module in the third layer repetitive processing module.

[0069] The feature maps with increased resolution are superimposed on the feature maps of the first and second processing modules to obtain feature maps with different resolutions. The result of two multi-scale fusions of the feature maps with increased resolution is input into the initial anti-aliasing pooling module.

[0070] It should be noted that in the process of feature extraction of target images, if the same receptive field size is used, the convolutional neural network model is very likely to lose attention to the foreground information of the target region of interest after the network is deepened. Therefore, a multi-scale module that restores the feature map scale step by step is added below the backbone feature extraction network to achieve multi-scale fusion.

[0071] The multi-scale fusion module, consisting of convolutional layers followed by deconvolutional layers and then another convolutional layer, adds non-linearity while restoring the original feature maps. Each feature map at a different resolution is magnified by a factor of two after passing through the multi-scale fusion module, and then superimposed with the corresponding resolution feature map from the feature extraction backbone to obtain a fused feature map of different resolutions. Through this connection, each layer's feature map incorporates features of different resolutions, including high-dimensional semantic information and low-dimensional texture information, thus achieving the fusion of features at different resolutions and allowing the network to balance its attention to different resolution sizes.

[0072] refer to Figure 5 The behavior safety classification module includes: a residual network 501, an improved global context module 502, a second global pooling layer 503, a fully connected layer 504, and a normalization layer 505. Based on the worker's region of interest, the behavior safety classification module determines whether the worker's behavior includes dangerous behavior, which can be achieved through steps S301 to S303, wherein:

[0073] Step S301: Semantic feature extraction is performed on the region of interest of the staff through a residual network to obtain the input feature map;

[0074] Step S302: The input feature map is processed by global information extraction through the improved global context module to obtain the output feature map;

[0075] Step S303: The output feature map is sequentially processed by pooling, concatenating and normalizing through the second global pooling layer to obtain the confidence level of whether the worker's behavior includes dangerous behavior.

[0076] Because real-time target tracking in video requires high precision, the target tracking module can employ the YOLOv4 algorithm. This algorithm boasts the highest accuracy among real-time target detection algorithms, achieving an optimal balance between accuracy and speed. Through the target tracking module, based on the fused feature information obtained from the improved feature fusion module, the YOLOv4 algorithm can more accurately regress the region of interest of the worker in each frame of the image. Then, the region of interest of each frame is input into the behavior safety classification module for binary classification. When the output of the behavior safety classification module exceeds a set safety threshold, it is judged as a dangerous behavior, at which point the worker is alerted, improving safety during the project implementation process.

[0077] refer to Figure 6 The improved global context module performs global information extraction processing on the input feature map to obtain the output feature map, which can be achieved through steps S401 to S405, wherein:

[0078] Step S401: Perform feature compression processing on the input feature map through the channel attention module;

[0079] Here, the dimensions of the input feature map are C×H×W.

[0080] Step S402: The compressed input feature map is transposed and normalized sequentially through the first convolution operation and normalization operation.

[0081] Here, the input features, after transpose and normalization, have a dimension of HW×1.

[0082] Step S403: Multiply the input feature map with the normalized features, and reduce the feature dimension through the second convolution operation;

[0083] Here, the feature dimension obtained by multiplying the input feature map with the normalized feature is C×1×1, and then the feature dimension is reduced to (C / r)×1×1 through the second convolution operation.

[0084] Step S404: The reduced-dimensional features are restored to their original dimensions by sequentially passing an activation layer and a third convolution operation.

[0085] Here, the reduced-dimensional features are restored to their original dimensions C×1×1 by sequentially passing an activation layer and a third convolution operation.

[0086] Step S405: The input feature map is superimposed with the features after the third convolution operation reduces the dimension to obtain the output feature map.

[0087] Here, the input feature map is superimposed with the features after the third convolution operation reduces the dimension, resulting in an output feature map with dimensions C×H×W.

[0088] Since the global context information simulated by the nonlocal module is almost identical for different query positions, a large amount of redundant information exists in the huge attention graph. Here, an improved global context module is used to simplify the generation of the attention graph. It directly uses convolution to generate a global attention graph that is independent of the query position and shares it across all positions. This simplified nonlocal module reduces computational complexity while maintaining accuracy.

[0089] The improved global context module can obtain more global information, that is, the output feature map references features from all locations. In contrast, the input feature map only calculates the local area covered by the convolution kernel, that is, it only considers the correlation of a small patch of image pixels within the receptive field, and lacks a grasp of global features.

[0090] refer to Figure 7 The input feature map is compressed using a channel attention module to obtain the output features, which can be achieved through steps S501 to S503, wherein:

[0091] Step S501: After multiplying the input feature map with the transposed feature map, the channel attention map is obtained by normalization.

[0092] Step S502: Multiply the channel attention map with the input feature map to obtain the feature map after enhanced feature representation;

[0093] Step S503: Multiply the enhanced feature map by a coefficient and superimpose it with the input feature map to obtain the output feature.

[0094] Introducing a channel attention module before the first convolution of the improved global context module allows the global attention map to acquire more global information. Similar to a self-attention mechanism, the channel attention module captures the channel dependencies between any two channel feature maps and updates the value of each channel using a weighted sum of all channels. Therefore, compared to the input, the weights of different channels in the enhanced feature map are recalibrated, strengthening interdependent feature channels and improving the feature representation of semantic features.

[0095] The operation of the channel attention module can be represented by the following formula (1):

[0096]

[0097] This formula represents the influence of the i-th channel on the j-th channel, and is used in formula (2) for the output features below.

[0098]

[0099] Among them, CA ji This represents the influence of the i-th channel on the j-th channel. In and Out represent the input and output features, respectively, and C is the number of channels. γ is a parameter that is learned during the network's learning process, with an initial value of 0.

[0100] In deep convolutional neural networks, each feature channel is considered a response of a specific class, and these responses are interconnected. The channel attention module utilizes spatial information from all relevant locations to construct feature channel correlations, thereby optimizing semantically specific feature representations. Here, the channel attention module is placed before the first convolutional operation of the global context module, and the channel weights are weighted using a network-learnable parameter, thus improving the global context module.

[0101] China's 5G development has entered a phase of rapid acceleration, leading to a continuous speedup in 5G base station construction. Safety is paramount during surveying and acceptance work, especially since this work often involves working at heights and is typically carried out by only one engineer. The engineer's personal safety measures rely solely on self-inspection, and dangerous actions during the work process are often unannounced. Therefore, strengthening safety supervision during the work process is crucial to eliminating potential safety hazards. With the increasing sophistication of modern image acquisition and digital video technologies, information about the work site can be obtained through video surveillance, including whether personal safety protection measures are compliant and whether dangerous operations are detected in the video. Accurate and real-time detection can promptly identify dangerous actions and issue warnings, greatly improving the safety of on-site personnel during surveying and acceptance.

[0102] In recent years, research on deep learning-based object detection and tracking algorithms has been booming. Mature algorithms not only free up significant human and material resources but also offer advantages in accuracy and processing speed. While existing algorithms are relatively mature, their robustness, accuracy, and real-time performance in detecting moving targets in video sequences still need improvement. Existing technologies mainly suffer from the following shortcomings: Related technology 1 only involves detecting safety equipment during implementation, lacking a pre-implementation safety inspection process. Related technology 2 only involves detecting the wearing of safety helmets and seat belts, lacking information on the wearing of other safety equipment.

[0103] In summary, the embodiments of this application provide a method for personnel safety detection and real-time monitoring before and during the implementation of 5G base stations. This is achieved through an improved target detection and classification network, which improves the accuracy and generalization of model detection while ensuring the real-time performance of the algorithm, making it applicable to complex and ever-changing application scenarios.

[0104] With the deepening of the national informatization strategy and the continuous advancement of urban network construction across the country, fifth-generation communication technology (5G, 5G) is being used. th The construction of 5G base stations has also accelerated. By the end of last year, my country had accumulated over 400,000 5G base stations, bringing it one step closer to the commercial application of 5G standalone networking. While accelerating 5G network construction, ensuring safety during the construction process is paramount. Since 5G base station surveying and acceptance work is typically carried out by a single engineer, and high-altitude operations are extremely dangerous, this application aims to provide a real-time safety detection method for 5G base station surveying and acceptance personnel based on deep learning. This method addresses the deficiencies and shortcomings of existing technologies, thereby ensuring the safety of engineers during operations.

[0105] To address the aforementioned issues, this application employs a deep learning-based approach to monitor the safety of workers during the inspection and acceptance of 5G base stations in real time. A comprehensive process is designed to determine compliance before and during implementation. Upon obtaining the determination results, timely alarms and reminders can be issued for dangerous behaviors by workers, thereby improving the safety of the project implementation. Based on the needs of actual application scenarios, the accuracy of safety equipment wearing compliance checks before implementation is crucial. During implementation, real-time tracking of workers and assessment of the compliance of safety equipment wearing and the safety of their actions are required, demanding high real-time performance. Therefore, this application adopts different methods before and during implementation to better ensure the safety of construction personnel.

[0106] refer to Figure 2 The overall methodology mainly consists of two parts based on the implementation process: pre-implementation and implementation.

[0107] 1. Before project implementation, take a photo of the worker wearing personal protective equipment (PPE) using a mobile terminal device and upload it. The photo will be used to determine compliance of PPE wearing by a PPE compliance detection model, including whether the worker is wearing a safety helmet, safety clothing, safety shoes, goggles, and anti-static wrist strap, and whether the equipment is worn correctly. If the output of the PPE wearing compliance detection model shows that the detection frame contains all of the above safety equipment, the worker is deemed to be wearing PPE compliantly and can proceed with the subsequent work; otherwise, non-compliance information will be output and the worker will be reminded.

[0108] 2. During project implementation, cameras are used to film workers during the process. To better track workers, this application proposes an improved feature extraction and fusion module. Based on the Transformer attention module, the extracted image features from pre-construction images are embedded into the frame-by-frame image features of the video to be searched. Anti-aliasing pooling, feature channel selection enhancement, and multi-scale modules are used to improve the feature extraction stage. The safety monitoring process involves first extracting features using the improved feature extraction and fusion module and then determining the worker's region of interest using the target tracking module. The region of interest is then input into a behavior safety classification module for categorizing safe and dangerous behaviors. Non-compliant use of personal protective equipment, working outside the designated work area, and using tools near edges during high-altitude operations are classified as dangerous behaviors. When the output of the behavior safety classification module exceeds a set safety threshold, it is determined to be a dangerous behavior, and an alarm is issued to alert workers who have violated regulations.

[0109] The target detection, tracking, and classification modules of the two-part design will be described in detail below.

[0110] Part 1: Compliance Testing of Personal Safety Equipment Wearing

[0111] For the task of verifying the compliance of personal safety equipment wearing before project implementation, compared with the safety inspection during implementation, the accuracy requirements of the results are higher and the real-time requirements are lower. Considering this, the embodiment of this application is based on the Faster R-CNN algorithm and improves it in the feature extraction stage. The specific improvement method can be seen in the feature extraction branch in the feature extraction fusion module. The input part of the network is preprocessed by enhancement (here, grayscale stretching is selected). The overall process of the Faster R-CNN algorithm includes: First, using a shared convolutional layer to extract features from the whole image to obtain a global feature map; Second, the extracted features are fed into the Region Proposal Network (RPN). The RPN generates the bounding boxes to be detected, specifies the location of the region of interest, and performs the first correction on the bounding boxes of the region of interest; Then, the region of interest pooling layer selects the features corresponding to each region of interest on the feature map according to the output of the RPN and sets the dimension to a fixed value; Finally, a fully connected layer is used to classify the detection boxes and perform a second correction on the target bounding boxes.

[0112] The Faster R-CNN algorithm introduces the RPN network, which can perform region proposal faster than other algorithms in the R-CNN series. The RPN network is also a convolutional neural network; its improved training speed is due to parameter sharing with the subsequent detection network. This algorithm uses the RPN network to select candidate regions faster and better, and then classifies and identifies the target based on the multiple proposed candidate boxes.

[0113] The RPN network structure is shown in the figure below. For each anchor point on the global feature map extracted by the shared convolutional layer, anchor boxes with different scales and aspect ratios are generated. These anchor boxes are pre-defined, with k = 9, meaning there are 9 types of rectangles with lengths of 128, 256, and 512, and aspect ratios of 2:1, 1:1, and 1:2. These anchor boxes are processed through a sliding window (3x3 convolution) to obtain 256-dimensional features, which are then input into two network layers (fully connected layers) to obtain classification results, i.e., whether the anchor box features belong to the foreground and their coordinate positions. Since these regions of interest have different lengths and scales, a region of interest pooling layer is used to obtain a uniform size. The full-image features extracted using the shared convolutional layer are not only used by the RPN network to generate detection boxes, but also share parameters with subsequent classification and regression modules, thus improving the training speed of Faster R-CNN.

[0114] Part Two: Real-time Safety Monitoring of Survey and Acceptance Activities

[0115] To monitor the behavior of staff during the inspection and acceptance of 5G base stations in real time and ensure the safety of project implementation, this application proposes an improved target tracking module that can more accurately regress the area where the staff is located in each frame of image. Then, the region of interest of each frame of image is input into the behavior safety classification module for binary classification. When the output result of the behavior safety classification module is greater than the set safety threshold, it is judged as dangerous behavior, and the staff can be alerted to improve the safety of the project implementation process.

[0116] 2.1 Staff Target Tracking Module:

[0117] The target tracking in this application embodiment has high real-time requirements, so the YOLOv4 algorithm is used instead of the Faster RCNN method. The YOLOv4 algorithm has the highest accuracy among real-time target detection algorithms, achieving the best balance between accuracy and speed. In order to further improve the accuracy of target tracking, this application embodiment proposes an improved feature extraction and fusion module to better achieve the accuracy of target tracking.

[0118] refer to Figure 3 This application embodiment uses images taken by workers before project implementation as template images, and each frame of real-time monitoring video as the search image. These two images are fed into an improved feature extraction and fusion module to extract fused features. To embed worker information from the template image into the search image, this application embodiment uses a Transformer approach to integrate the template feature information into the search feature. The fused features generated by this operation are more effective in identifying and locating workers, improving the accuracy of subsequent target tracking. Furthermore, considering the impact of image target offset caused by shooting and other factors on network output, this application embodiment incorporates anti-aliasing pooling into the improved feature extraction and fusion module for improvement. Also, considering the significant difference in the proportion of workers to the background, which increases the detection difficulty, this application proposes to address the target-background imbalance problem in the image through feature channel selective enhancement and multi-scale fusion, further improving the accuracy of security detection. The improvements will be described in detail below.

[0119] 2.1.1 Transformer Attention Module: Following the overall safety supervision process for staff inspecting and accepting 5G base stations as proposed in this application embodiment, an image needs to be taken before project implementation to verify the compliance of personal safety equipment use. This image also plays another crucial role in the subsequent safety monitoring during project implementation: serving as a template image. Considering the interference from various factors such as video shooting angle and outdoor lighting when tracking staff in videos, fusing the features of the template image with the features of the video image to be searched can significantly improve the accuracy of target localization and avoid interference from many uncertainties in actual application scenarios.

[0120] refer to Figure 3In this embodiment, the Transformer attention module 304 is used to embed the information of the staff in the template image into the image to be searched. Transformer has achieved great success in the field of natural language processing and is now widely used in the field of image processing. The principles of Transformer will not be detailed here. The specific implementation of feature fusion is as follows: taking the first Transformer module as an example, the template image and the image to be searched each extract features through their respective branches. After passing through the feature selection enhancement modules on their respective branches, the feature map is a tensor of dimension H×W×C. These feature maps are then stacked together to form a tensor of size 2×H×W×C, which is then fed into the Transformer for information fusion. The output at this point already considers the useful information in the template image. The fused features are then fed into the two feature extraction backbones and added to the multi-scale fused feature map as the input for the next stage.

[0121] 2.1.2 Anti-aliasing Pooling Module: During feature extraction from video frame images, the movement of personnel during the process, along with the shifting of video frame images and pre-processed images due to shooting methods, exacerbates the network output. This is because convolutional neural networks (CNNs) do not satisfy the sampling theorem during downsampling, neglecting signal aliasing. Therefore, even small displacements can drastically alter the network output. In signal processing, there are two methods to address this problem: one is to increase the sampling frequency, setting the stride to 1 in CNN processing (but a stride of 1 is the limit); the other is to use low-pass filtering for anti-aliasing before downsampling. Aliasing refers to the sawtooth effect caused by the sampling frequency not satisfying the sampling theorem. Anti-aliasing filters eliminate this phenomenon by first using low-pass filtering and then downsampling, thus eliminating the high-frequency signal-induced non-sampling theorem violation. In CNNs, average pooling is equivalent to performing a box filter operation before downsampling, reducing high-frequency influences and maintaining a certain degree of translation invariance. Research has shown that max pooling achieves better results in extracting significant features; however, max pooling lacks anti-aliasing capabilities, so even small changes in input can have a significant impact on the network output. To reduce the impact of target offset, this application employs anti-aliasing pooling to optimize downsampling operations in the improved feature extraction and fusion module.

[0122] refer to Figure 4a. Max pooling can be viewed as two steps: the first step is dense maximum selection, implemented using a sliding window with a step size of 1, which exhibits translation invariance; the second step is downsampling. Due to the relatively low sampling frequency, high-frequency components are retained during sampling, violating translation invariance. To ensure max pooling satisfies the downsampling theorem, refer to... Figure 4 b. A low-pass filter is introduced between the dense maximum selection and the downsampling operation so that the translation invariance can be preserved to the maximum extent during downsampling. This process is represented by Equation (3):

[0123]

[0124] In the formula, the improved max pooling layer with a pooling window of k×k and a stride of s can be regarded as a dense maximum selection Max pooling layer with a pooling window of k×k. k Adding a low-pass filter Blur with a kernel size of m×m m and Subsampl with step size s s The last two steps are combined into a fuzzy low-pass filter, Blurpool, with a kernel size of m×m and a step size of s. m,s .

[0125] The specific parameters and settings are as follows: Replace the max pooling layer with a pooling window of 2×2 and a stride of 2 with an anti-aliasing max pooling operation. Specifically, first pass through a max pooling layer with a pooling window of 2×2 and a stride of 1, then perform a 2D convolution with a stride of 2 and a kernel set to... By integrating the fuzzy low-pass filter into the existing convolutional module in this way, the improved max pooling operation has translation invariance during feature extraction, which improves the accuracy and robustness of subsequent target tracking.

[0126] 2.1.3 Feature Channel Selection Enhancement Module: Generally, convolution operations aggregate information across feature dimensions. When the feature map undergoes further processing, features from different channels have the same weight, meaning the network pays equal attention to each dimension of the feature map. Since the tracked worker target occupies a small portion of the entire image, this becomes unreasonable as the network deepens. Therefore, a feature channel selection enhancement module is added to the improved feature extraction and fusion module to address the foreground (worker) and background imbalance problem in the segmentation issue.

[0127] refer to Figure 3 The feature channel selection enhancement module 303 consists of two convolutional layers, one global pooling layer, and two fully connected layers, and contains two branches. The first branch uses direct connection of residual structure, i.e., "short cut", and the second branch is used to recalibrate the downsampled feature map during the convolution operation.

[0128] The downsampled feature map of the target image is first fed into two convolutional layers to increase the receptive field and add non-linear features through a non-linear activation layer. After the two convolutions, feature recalibration is performed on this feature map. Assuming the size of the downsampled feature map of the target image is c×w×h, after a global pooling layer, each two-dimensional feature map becomes a real number, which, to some extent, has a global receptive field. This feature map then becomes c×1×1, achieving a global distribution of the response across the feature map channels, and enabling subsequent fully connected layers to obtain a global receptive field. Next, two consecutive fully connected layers are applied. First, the feature dimension is reduced to (c / 8)×1×1, and after an activation layer, it is raised back to the original dimension through another fully connected layer. Using two fully connected layers, compared to a single fully connected layer, can add more non-linear features, thereby fitting the complex correlations between channels, significantly reducing the number of connections and improving computational speed. Then, an activation layer using the sigmoid function normalizes the weights to the range [0, 1], thus obtaining the score for recalibrating the feature map downsampled from the target image. This structure can be viewed as similar to the gate mechanism in a recurrent neural network, where the feature vector of the target image quantifies the correlation between the feature channels of the downsampled feature map. The feature vector of the target image is then multiplied back into the features resulting from the two convolutions to weight the recalibrated score onto the features of each channel. Finally, the feature maps of the two branches are superimposed, ensuring the network remains in an optimal state and its performance does not decrease with increasing depth.

[0129] 2.1.4 Multi-scale fusion module: In the feature extraction process, if the same receptive field size is used, the convolutional neural network model is very likely to lose attention to the foreground information of the target region of interest after the network is deepened. Therefore, a multi-scale module that restores the feature map scale step by step is added below the backbone feature extraction network to achieve multi-scale fusion.

[0130] Traditional approaches include image pyramids and feature layering. Image pyramid structures are computationally expensive, while feature layering directly teaches different layers of the network the same information. Considering these two points, in this embodiment, the multi-scale fusion module 305 is connected by convolutional layers-deconvolutional layers-convolutional layers, which can restore the original feature maps while adding non-linear characteristics. After passing through the multi-scale fusion module, the feature maps of each resolution are scaled by a factor of two and superimposed on the corresponding resolution feature maps on the feature extraction backbone. Through this connection, the feature maps of each layer fuse features of different resolutions, including high-dimensional semantic information and low-dimensional texture information, thereby achieving the fusion of feature maps of different resolutions and allowing the network to balance its attention to different resolution sizes. At the same time, since this method only adds extra cross-layer connections to the original network, it adds almost no extra time and computation in practical applications.

[0131] 2.2, Behavioral Safety Classification Module:

[0132] After the target tracking module regresses the region of interest of the staff, it is input into the improved global context module. The overall structure is referenced. Figure 5 The semantic features obtained from the region of interest of the staff through the residual network are then processed by the improved global context module, global pooling layer, fully connected layer and Softmax to obtain a probability value. When the probability value is greater than the set threshold, it will be judged as dangerous behavior, otherwise it will be judged as safe behavior.

[0133] For different query positions, the global context information simulated by the nonlocal module is almost identical. Therefore, there is a large amount of redundant information in the huge HW×HW attention graph. The improved global context module simplifies the generation of the attention graph by directly generating a global attention graph that is independent of the query position using 1×1 convolutions and sharing it across all positions. The simplified nonlocal module reduces computational complexity while maintaining accuracy.

[0134] The improvements to the global context module will be described in detail below.

[0135] refer to Figure 6 A channel attention module is introduced before the first convolution of the improved global context module, and its structure is referenced. Figure 7 This method borrows the idea of ​​spatial grouping enhancement modules to enable the global attention map to acquire more global information. The input feature map is multiplied by the transpose of the input feature map, and then subjected to a Softmax operation to obtain a C×C channel attention map. The channel attention map is multiplied by the input feature map to enhance the feature representation. The enhanced feature map is multiplied by a coefficient γ and then added to the original feature map to obtain the output feature.

[0136] The above operation can be represented by the following formula (1):

[0137]

[0138] This formula represents the influence of the i-th channel on the j-th channel, and is used in formula (2) for the output features below.

[0139]

[0140] Among them, CA ji This represents the influence of the i-th channel on the j-th channel. In and Out represent the input and output features, respectively, and C is the number of channels. γ is a parameter that is learned during the network's learning process, with an initial value of 0.

[0141] In deep convolutional neural networks, each feature channel is considered a response of a specific class, and these responses are interconnected. The channel attention module utilizes spatial information from all relevant locations to construct feature channel correlations, thereby optimizing semantically specific feature representations. Here, the channel attention module is placed before the first convolutional operation of the global context module, and the channel weights are weighted using a network-learnable parameter, thus improving the global context module.

[0142] It can be seen from the above embodiments that:

[0143] 1) This application embodiment designs a method for the 5G base station survey and acceptance stages, which is to detect the compliance of safety equipment wearing before the project is implemented and to track and supervise the behavior of staff in real time during the implementation process. Through accurate and real-time detection, dangers can be detected in time and alarms can be issued to staff who violate the rules, which greatly improves the safety of staff.

[0144] 2) This application proposes an improved two-stage real-time monitoring method for survey and acceptance activities. The first stage tracks the area where personnel are located. An improved feature extraction and fusion module is proposed for the target tracking module. This improved module includes: firstly, using a Transformer attention module to embed template features into the search features, which facilitates the localization of personnel; secondly, adding anti-aliasing pooling to reduce the impact of offset on the network output; and finally, addressing the target-background imbalance problem in the image through feature channel selective enhancement and multi-scale fusion. These improvements in the feature extraction stage enhance the accuracy and robustness of real-time tracking. The second stage uses a behavior safety classification module (an improved classification network) for behavior classification. The accuracy of behavior classification is improved by adding an improved global context module to the behavior safety classification module.

[0145] Compared with related technologies, the embodiments of this application have the following advantages:

[0146] 1) Related technologies only involve testing the wearing of safety helmets and safety belts by workers. The personal safety equipment wearing compliance test proposed in this application involves whether safety helmets, safety clothing, safety shoes, goggles, and anti-static wrist straps are worn correctly. Since 5G base station survey and acceptance scenarios include, but are not limited to, high-altitude operation scenarios, testing only safety helmets and safety belts is insufficient.

[0147] 2) This application proposes to divide safety testing into pre-implementation and implementation processes. Before implementation, subsequent work can only be carried out if the personal safety equipment wearing test is compliant, thereby preventing danger.

[0148] 3) In the improved feature extraction and fusion module, the Transformer attention mechanism is used to fuse the template image features into the search image features. The feature channel selection enhancement module, multi-scale fusion module and anti-aliasing pooling module are also used for improvement, thereby solving the problems of target and background imbalance and network sensitivity to input changes. In this way, the accuracy and robustness of detection can be improved without affecting real-time performance.

[0149] 4) The behavior detection methods in related technologies cannot track construction workers in videos. This means that if a safety wearing issue occurs, only all construction workers can be alerted, not just the specific workers involved. This application proposes a feature fusion method in the target tracking module, using images captured before implementation as template images and the images to be detected. This is more conducive to target localization and allows for hazard alerts to relevant workers.

[0150] 5) Related technologies only involve testing personnel safety equipment during implementation, lacking a testing process before project implementation. Timely detection of non-compliance with personal protective equipment before implementation is crucial. This application's embodiment divides safety testing into pre-implementation and implementation phases. Subsequent work can only proceed after the personal safety equipment is tested and found to be compliant, thus preventing potential hazards.

[0151] Based on the foregoing embodiments, this application provides a behavior detection device, which includes the modules included, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0152] Figure 8 This is a schematic diagram of the composition structure of a behavior detection device provided in an embodiment of this application, as shown below. Figure 8 As shown, the behavior detection device 800 includes:

[0153] The improved feature extraction and fusion module 810 is used to embed the feature information of the template image into the feature information of the image to be searched to obtain fused feature information; wherein, the template image is an image taken by the staff before the implementation of the project, and the image to be searched is each frame of the real-time monitoring video of the staff.

[0154] The target tracking module 820 is used to determine the region of interest of the staff member in the image to be searched based on the fused feature information;

[0155] The behavior safety classification module 830 is used to determine whether the behavior of the staff member includes dangerous behavior based on the staff member's region of interest.

[0156] In some embodiments, the improved feature extraction and fusion module includes: a shared convolutional layer, an anti-aliasing pooling module, a feature channel selection enhancement module, and an attention module;

[0157] The shared convolutional layer is used to extract features from the template image and the image to be searched, respectively, to obtain the global feature map of the template image and the global feature map of the image to be searched.

[0158] The anti-aliasing pooling module is used to downsample the global feature map of the target image to obtain a downsampled feature map of the target image; the target image includes a template image or an image to be searched.

[0159] The feature channel selection enhancement module is used to recalibrate the feature map downsampled from the target image to obtain the feature information of the target image;

[0160] The attention module is used to fuse the feature information of the template image and the feature information of the image to be searched to obtain the fused feature information; wherein, the fused feature information contains useful information from the template image.

[0161] In some embodiments, the anti-aliasing pooling module is further configured to select the maximum pixel value from the global feature map of the target image to obtain a feature map of the maximum pixel value of the target image; and to downsample the feature map of the maximum pixel value of the target image to obtain a downsampled feature map of the target image.

[0162] In some embodiments, the feature channel selection enhancement module is further configured to: sequentially perform convolution and pooling processing on the feature map of the downsampled target image through two convolutional layers and a first global pooling layer; reduce the dimensionality of the features after convolution and pooling processing through a first fully connected layer; restore the dimensionality of the reduced features to their original dimensions by sequentially passing through a first activation layer and a second fully connected layer to obtain the feature vector of the target image; multiply the feature vector of the target image with the features of the target image after passing through the two convolutional layers, and superimpose it with the downsampled feature map to obtain the feature information of the target image.

[0163] In some embodiments, the improved feature extraction and fusion module sequentially includes: a shared convolutional layer, an initial anti-aliasing pooling module, the feature channel selection enhancement module, and a three-layer repetitive processing module, wherein each layer of the repetitive processing module includes: an anti-aliasing pooling module, the feature channel selection enhancement module, and an attention module; the improved feature extraction and fusion module further includes a multi-scale fusion module;

[0164] The multi-scale fusion module is used to sequentially perform convolutional, deconvolutional, and convolutional operations on the reduced-resolution feature map to obtain a magnified feature map.

[0165] The reduced-resolution feature map is obtained by the anti-aliasing pooling module and the feature channel selection enhancement module in the third-layer repetitive processing module.

[0166] The feature maps with magnified resolution are superimposed on the feature maps of the first and second processing modules to obtain feature maps with different resolutions; the result of two multi-scale fusions of the magnified resolution feature maps is input into the initial anti-aliasing pooling module.

[0167] In some embodiments, the behavioral safety classification module includes: a residual network, an improved global context module, a second global pooling layer, a fully connected layer, and a normalization layer;

[0168] A residual network is used to extract semantic features from the region of interest of the staff member to obtain an input feature map;

[0169] An improved global context module is used to perform global information extraction processing on the input feature map to obtain an output feature map;

[0170] The second global pooling layer is used to sequentially perform pooling, concatenation, and normalization processing on the output feature map to obtain the confidence level of whether the worker's behavior includes dangerous behavior.

[0171] In some embodiments, the improved global context module is further configured to perform feature compression processing on the input feature map through a channel attention module; transpose and normalize the compressed input feature map sequentially through a first convolution operation and a normalization operation; multiply the input feature map with the normalized feature, and reduce the feature dimension through a second convolution operation; restore the reduced-dimensional feature to its original dimension sequentially through an activation layer and a third convolution operation; and superimpose the input feature map with the reduced-dimensional feature obtained by the third convolution operation to obtain the output feature map.

[0172] In some embodiments, the channel attention module is configured to: multiply the input feature map with a transposed feature map and then normalize the result to obtain a channel attention map; multiply the channel attention map with the input feature map to obtain a feature map with enhanced feature representation; multiply the feature map with enhanced feature representation by a coefficient and then superimpose it with the input feature map to obtain the output feature.

[0173] It should be noted that, in the embodiments of this application, if the above-described behavior detection method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0174] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0175] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.

[0176] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0177] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0178] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0179] It should be noted that, Figure 9 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this application, such as... Figure 9 As shown, the hardware entities of the computer device 900 include: a processor 901, a communication interface 902, and a memory 903, wherein: the processor 901 typically controls the overall operation of the computer device 900. The communication interface 902 enables the computer device to communicate with other terminals or servers via a network.

[0180] The memory 903 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 901 and various modules in the computer device 900. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 901, the communication interface 902, and the memory 903 can be performed via bus 904.

[0181] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0182] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0183] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0184] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0185] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0186] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0187] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0188] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A behavior detection method, characterized in that, The method includes: By using an improved feature extraction and fusion module, the feature information of the template image is embedded into the feature information of the image to be searched to obtain fused feature information; wherein, the template image is an image taken by the staff before the project is implemented, and the image to be searched is each frame of the real-time monitoring video of the staff. The target tracking module determines the region of interest of the staff member in the image to be searched based on the fused feature information. The behavioral safety classification module determines whether the employee's behavior includes dangerous behavior based on the employee's region of interest. The improved feature extraction and fusion module includes: a shared convolutional layer, an anti-aliasing pooling module, a feature channel selection enhancement module, and an attention module. The improved feature extraction and fusion module embeds the feature information of the template image into the feature information of the image to be searched to obtain fused feature information, including: The shared convolutional layer is used to extract features from the template image and the image to be searched, respectively, to obtain the global feature map of the template image and the global feature map of the image to be searched; The anti-aliasing pooling module downsamples the global feature map of the target image to obtain a downsampled feature map of the target image; the target image includes the template image or the image to be searched. The feature map of the downsampled target image is recalibrated by the feature channel selection enhancement module to obtain the feature information of the target image. The attention module fuses the feature information of the template image and the feature information of the image to be searched to obtain the fused feature information; wherein, the fused feature information contains useful information from the template image; The step of recalibrating the feature map downsampled from the target image through the feature channel selection enhancement module to obtain the feature information of the target image includes: The feature map of the target image downsampled is sequentially processed by two convolutional layers and one first global pooling layer, respectively, through convolution and pooling. The first fully connected layer reduces the dimensionality of the features processed by convolution and pooling. The features of the target image are restored to their original dimensions by sequentially passing through the first activation layer and the second fully connected layer; The feature vector of the target image is multiplied with the feature vector of the target image after passing through the two convolutional layers, and then superimposed with the downsampled feature map to obtain the feature information of the target image.

2. The method according to claim 1, characterized in that, The anti-aliasing pooling module downsamples the global feature map of the target image to obtain a downsampled feature map of the target image, including: The maximum pixel value is selected from the global feature map of the target image to obtain the feature map of the maximum pixel value of the target image; The feature map of the maximum value of the target image is downsampled to obtain the downsampled feature map of the target image.

3. The method according to claim 1, characterized in that, The improved feature extraction and fusion module sequentially includes: a shared convolutional layer, an initial anti-aliasing pooling module, the feature channel selection enhancement module, and a three-layer repetitive processing module, wherein each layer of the repetitive processing module includes: an anti-aliasing pooling module, the feature channel selection enhancement module, and an attention module; the improved feature extraction and fusion module also includes a multi-scale fusion module; After performing convolutional, deconvolutional, and convolutional layer operations on the reduced-resolution feature map through the multi-scale fusion module, a larger-resolution feature map is obtained. The reduced-resolution feature map is obtained by the anti-aliasing pooling module and the feature channel selection enhancement module in the third-layer repetitive processing module. The feature maps with magnified resolution are superimposed on the feature maps of the first and second processing modules to obtain feature maps with different resolutions; the result of two multi-scale fusions of the magnified resolution feature maps is input into the initial anti-aliasing pooling module.

4. The method according to any one of claims 1 to 3, characterized in that, The behavioral safety classification module includes: a residual network, an improved global context module, a second global pooling layer, a fully connected layer, and a normalization layer; The behavioral safety classification module determines, based on the worker's region of interest, whether the worker's behavior includes dangerous behavior, including: Semantic feature extraction is performed on the region of interest of the staff member using a residual network to obtain an input feature map; The input feature map is processed by a global information extraction module to obtain an output feature map. The output feature map is sequentially processed by the second global pooling layer, including pooling, concatenation, and normalization, to obtain the confidence level of whether the worker's behavior includes dangerous behavior.

5. The method according to claim 4, characterized in that, The input feature map is processed by a modified global context module to extract global information, resulting in an output feature map, including: The input feature map is compressed using a channel attention module. The compressed input feature map is transposed and normalized sequentially through the first convolution operation and the normalization operation. The input feature map is multiplied with the normalized feature, and the feature dimension is reduced by a second convolution operation; The reduced-dimensional features are restored to their original dimensions by sequentially passing through an activation layer and a third convolution operation. The input feature map is superimposed with the features after the third convolution operation reduces the dimension to obtain the output feature map.

6. The method according to claim 5, characterized in that, The input feature map is compressed using a channel attention module to obtain output features, including: After multiplying the input feature map with the transposed feature map, the channel attention map is obtained by normalization. Multiply the channel attention map with the input feature map to obtain the feature map after enhanced feature representation; The enhanced feature map is multiplied by a coefficient and then superimposed with the input feature map to obtain the output feature.

7. A behavior detection device, characterized in that, The device includes: An improved feature extraction and fusion module is used to embed the feature information of a template image into the feature information of the image to be searched to obtain fused feature information; wherein, the template image is an image taken by the staff before the project is implemented, and the image to be searched is each frame of the real-time monitoring video of the staff. The target tracking module is used to determine the region of interest of the staff member in the image to be searched based on the fused feature information; The behavior safety classification module is used to determine whether the worker's behavior includes dangerous behavior based on the worker's region of interest. The improved feature extraction and fusion module includes: a shared convolutional layer, an anti-aliasing pooling module, a feature channel selection enhancement module, and an attention module; The shared convolutional layer is used to extract features from the template image and the image to be searched, respectively, to obtain the global feature map of the template image and the global feature map of the image to be searched; The anti-aliasing pooling module is used to downsample the global feature map of the target image to obtain a downsampled feature map of the target image; the target image includes the template image or the image to be searched. The feature channel selection enhancement module is used to recalibrate the feature map downsampled from the target image to obtain the feature information of the target image; The attention module is used to fuse the feature information of the template image and the feature information of the image to be searched to obtain the fused feature information; wherein, the fused feature information contains useful information in the template image; The feature channel selection enhancement module is further configured to: sequentially perform convolution and pooling processing on the downsampled feature map of the target image through two convolutional layers and a first global pooling layer; reduce the dimensionality of the features after convolution and pooling processing through a first fully connected layer; restore the dimensionality-reduced features to their original dimensions through a first activation layer and a second fully connected layer to obtain the feature vector of the target image; multiply the feature vector of the target image with the features of the target image after passing through the two convolutional layers, and superimpose it with the downsampled feature map to obtain the feature information of the target image.

8. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • RGB-D feature target tracking method based on twin network

    CN112785624A

  • Behavior detection method and device, electronic equipment and storage medium

    CN114067248A