Behavior detection method and related equipment
Through the method of detecting non-safe behavior in the mine environment, the feature extraction network is redesigned, and the partial convolution and adaptive fine-grained channel attention mechanism is introduced, which solves the problem of redundant calculation and feature extraction difficulties, and achieves efficient detection accuracy.
Patent Information
- Application Number
- CN202411904506.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-27
AI Technical Summary
The mine environment is complex, the image quality is degraded, the noise interference is large, and the feature extraction is difficult. The feature extraction module of the existing Transformer model has problems such as redundant calculation, large floating-point calculation volume and high latency.
The feature extraction network of RT-DETR-R18 was redesigned, partial convolution (PConv) technology was introduced to replace the parameter-intensive 3×3 convolution layer, and the adaptive fine-grained channel attention mechanism was fused after the PConv module.
It significantly reduces the number of parameters and calculation burden of the model, improves the detection accuracy, and can effectively complete non-safe behavior detection tasks in the mine environment.
Smart Images

Figure CN120047993A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of behavior detection technology, and in particular, to a behavior detection method and related devices. Background Art
[0002] Currently, the internal environment of mines is complex. Not only is the space narrow, cluttered, and there are a large number of floating dust particles, but also the internal lighting is uneven, with large shadow areas. All of the above objective environmental conditions will lead to a significant decline in image quality, large noise interference, and difficult feature extraction. If the size of the detection target is small, the available feature information is limited, and problems such as feature loss are likely to occur during detection. Currently, due to the limitations of hardware conditions such as small memory and low computing power, the detection models deployed on mine non-safe behavior detection platforms should be lightweight. However, there are a large number of redundant calculations in the feature extraction module of the Real-Time Detection Transformer (RT-DETR) model, with a large amount of floating-point operations and high latency. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose a behavior detection method and related devices.
[0004] Based on the above purpose, this application provides a behavior detection method, including:
[0005] Obtain preprocessed behavior data;
[0006] Perform at least one convolution process on the behavior data to obtain a first feature map;
[0007] Perform convolution processing on the first feature map using a first convolutional residual block and update the mask to obtain a second feature map;
[0008] Perform convolution processing on the second feature map using a second convolutional residual block, and based on the adaptive fine-grained channel attention mechanism, obtain a third feature map;
[0009] Input the third feature map into a hybrid encoder to obtain the final detection result.
[0010] In a possible implementation, the behavior data is preprocessed through the following steps:
[0011] Obtain a behavior video;
[0012] Based on the behavior video, extract video frames therefrom to obtain behavior images;
[0013] Convert the behavior images into deep learning tensors to obtain the behavior data.
[0014] In a possible implementation, the step of performing convolution processing on the first feature map by using the first convolutional residual block and updating the mask to obtain a second feature map includes:
[0015] In the first convolutional residual block, perform a first convolution processing on the first feature map to obtain a fourth feature map;
[0016] Perform partial convolution processing on the fourth feature map to obtain a fifth feature map;
[0017] Perform batch normalization processing on the fifth feature map to obtain a sixth feature map;
[0018] Use an activation function to process the sixth feature map to obtain a seventh feature map;
[0019] Based on the seventh feature map, obtain the second feature map.
[0020] In a possible implementation, the method further includes:
[0021] Perform a second convolution processing on the first feature map to obtain an eighth feature map;
[0022] Perform average pooling processing on the eighth feature map to obtain a ninth feature map;
[0023] The step of obtaining the second feature map based on the seventh feature map includes:
[0024] Based on the seventh feature map and the ninth feature map, obtain the second feature map.
[0025] In a possible implementation, the step of performing convolution processing on the second feature map by using the second convolutional residual block and obtaining a third feature map based on the adaptive fine-grained channel attention mechanism includes:
[0026] In the second convolutional residual block, perform a first convolution processing on the second feature map to obtain a tenth feature map;
[0027] Perform partial convolution processing on the tenth feature map to obtain an eleventh feature map;
[0028] Perform batch normalization processing on the eleventh feature map to obtain a twelfth feature map;
[0029] Use an activation function to process the twelfth feature map to obtain a thirteenth feature map;
[0030] Based on the thirteenth feature map, perform weighted processing on features of different frequencies through an adaptive fusion strategy to obtain a fourteenth feature map;
[0031] Based on the fourteenth feature map, the third feature map is obtained.
[0032] In a possible implementation manner, the method further includes:
[0033] Performing a second convolution process on the second feature map to obtain a fifteenth feature map;
[0034] Performing average pooling on the fifteenth feature map to obtain a sixteenth feature map;
[0035] The obtaining the third feature map based on the fourteenth feature map includes:
[0036] Based on the fourteenth feature map and the sixteenth feature map, the third feature map is obtained.
[0037] Based on the same inventive concept, an embodiment of the present application further provides a behavior detection device, including:
[0038] An acquisition module, configured to acquire preprocessed behavior data;
[0039] A convolution module, configured to perform at least one convolution process on the behavior data to obtain a first feature map;
[0040] A first residual module, configured to perform a convolution process on the first feature map by using a first convolutional residual block and update a mask to obtain a second feature map;
[0041] A second residual module, configured to perform a convolution process on the second feature map by using a second convolutional residual block and obtain a third feature map based on an adaptive fine-grained channel attention mechanism;
[0042] A detection module, configured to input the third feature map into a hybrid encoder to obtain a final detection result.
[0043] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the behavior detection method described in any one of the above is implemented.
[0044] Based on the same inventive concept, an embodiment of the present application further provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the behavior detection method described in any one of the above.
[0045] Based on the same inventive concept, an embodiment of the present application further provides a computer program product, which includes computer program instructions, and the computer instructions are used to cause the computer program product to execute the behavior detection method described in any one of the above.
[0046] As can be seen from the above, the behavior detection method and related devices provided in this application obtain preprocessed behavior data; perform at least one convolution process on the behavior data to obtain a first feature map; use a first convolutional residual block to perform a convolution process on the first feature map and update the mask to obtain a second feature map; use a second convolutional residual block to perform a convolution process on the second feature map, and based on the adaptive fine-grained channel attention mechanism, obtain a third feature map; input the third feature map into the hybrid encoder to obtain the final detection result. The embodiment of this application redesigned the feature extraction network of RT-DETR-R18, and innovatively introduced the partial convolution (PConv) technology, and specifically replaced part of the parameter-intensive 3×3 convolutional layers in the feature extraction network. This measure not only significantly reduces the number of parameters of the model and alleviates the computational burden, but also retains the powerful detection ability of the model. At the same time, the adaptive fine-grained channel attention mechanism (FCA) is fused after the PConv module, and refined weight adjustment is performed on the feature map output by the last PConv module. In this way, the model can more intelligently identify and strengthen key features, effectively improving the detection accuracy of non-safe behaviors in the mine environment, and finally the model can excellently complete the non-safe behavior detection task in the mine environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in this application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 It is a schematic flow chart of the behavior detection method according to the embodiment of this application;
[0049] Figure 2 It is a schematic flow chart of the detailed processing of the behavior data according to the embodiment of this application;
[0050] Figure 3 It is a schematic flow chart of the video segmentation algorithm according to the embodiment of this application;
[0051] Figure 4 It is a schematic structural diagram of the behavior detection device according to the embodiment of this application;
[0052] Figure 5 It is a schematic structural diagram of the electronic device according to the embodiment of this application. Detailed implementation manners
[0053] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0054] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the field to which the present application belongs. The terms "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0055] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.
[0056] For example, when receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.
[0057] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0058] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0059] As described in the background art section, the current internal environment of the mine is complex. Not only is the space narrow, messy, and there are a large number of floating dust particles, but also the internal lighting is uneven, with large shadow areas. The above objective environmental conditions will all lead to a significant decline in image quality, large noise interference, and difficulty in feature extraction. If the size of the detection target is small, the available feature information is limited, and problems such as feature loss are likely to occur during detection. Currently, due to the limitations of small memory and low computing power in hardware conditions, the detection model deployed on the mine non-safe behavior detection platform should be lightweight. However, there are a large number of redundant calculations in the feature extraction module of the real-time detection Transformer model, with a large amount of floating-point operations and high latency.
[0060] Considering the above comprehensively, the embodiment of the present application proposes a behavior detection method, which includes obtaining preprocessed behavior data; performing at least one convolution process on the behavior data to obtain a first feature map; using a first convolutional residual block to perform a convolution process on the first feature map and update the mask to obtain a second feature map; using a second convolutional residual block to perform a convolution process on the second feature map, and based on the adaptive fine-grained channel attention mechanism, obtain a third feature map; inputting the third feature map into a hybrid encoder to obtain a final detection result. The embodiment of the present application redesigned the feature extraction network of RT-DETR-R18, innovatively introduced the partial convolution technology, and specifically replaced part of the parameter-intensive 3×3 convolutional layers in the feature extraction network. This measure not only significantly reduces the number of parameters of the model and alleviates the computational burden, but also retains the powerful detection ability of the model. At the same time, the adaptive fine-grained channel attention mechanism is fused after the PConv module, and refined weight adjustment is implemented on the feature map output by the last PConv module. In this way, the model can more intelligently identify and strengthen key features, effectively improving the detection accuracy of non-safe behaviors in the mine environment. Finally, the model can excellently complete the non-safe behavior detection task in the mine environment.
[0061] Hereinafter, the technical solutions of the embodiments of the present application will be described in detail through specific embodiments.
[0062] Refer to Figure 1 , the behavior detection method of the embodiment of the present application includes the following steps:
[0063] Step S101, obtaining preprocessed behavior data;
[0064] Step S102, performing at least one convolution process on the behavior data to obtain a first feature map;
[0065] Step S103, using a first convolutional residual block to perform a convolution process on the first feature map and update the mask to obtain a second feature map;
[0066] Step S104: Convolve the second feature map using the second convolutional residual block to obtain a third feature map based on the adaptive fine-grained channel attention mechanism;
[0067] Step S105: Input the third feature map into the hybrid encoder to obtain the final detection result.
[0068] Reference Figure 2 , which is a schematic diagram of the detailed processing flow of the behavior data in the embodiment of the present application.
[0069] The specific meanings in the figure are shown in Table 1 below:
[0070] Table 1 Corresponding Meanings
[0071]
[0072]
[0073] The following is an illustration of the embodiments of the present application by combining Figure 1 and Figure 2 to illustrate the embodiments of the present application.
[0074] First, before detecting the behavior data, it is necessary to train the model for detecting the behavior data.
[0075] Specifically, in the dataset production stage: A dedicated non-safe behavior dataset MCSVD (Mine-Car Safety Violation Dataset) is rebuilt for the specific scenario of the mine monkey car passage. In the object detection algorithm based on deep learning, the quality of the dataset determines the training result of the model. Therefore, the self-built dataset MCSVD of this project is obtained by shooting videos on-site in the mine. The videos are segmented by frame, data is cleaned, data is labeled, and an offline data augmentation method RandomFixedMix is proposed to expand the dataset.
[0076] Model training stage: The mine non-safe behavior detection model is a deep learning neural network, which consists of three parts: a feature extraction network, a hybrid encoder, and a decoder. Among them, the improved ResNet-18 model proposed in this application is used for deep feature extraction in the feature extraction network. Specifically, a new residual module PCABlock (the first convolutional residual block) and PConvBlock are designed to replace the residual blocks in the network, and the outputs of the last three stages of the network are selected as the input data of the hybrid encoder. The hybrid encoder consists of an AIFI (Adaptive Feature Interaction Fusion) module and a CCFM (Cross-scale Context Fusion Module). The AIFI module is responsible for performing fine encoding on the highest-level S5 features, while the CCFM module effectively integrates multi-scale feature information by combining bottom-up and top-down feature fusion paths to generate high-quality image feature representations. In the decoding stage, RT-DETR uses the IoU-aware query mechanism to select a predefined number of image features from the feature sequence output by the encoder as the initial object queries, and then gradually refines the positions and confidence scores of the predicted bounding boxes through multiple iterative optimization processes, and finally outputs accurate object detection results.
[0077] Finally, in the detection stage: The trained model is used for detection in this stage, and it is judged whether there is a non-safe behavior according to the detection results.
[0078] For the dataset production stage, refer to Figure 3 , which is the schematic diagram of the video segmentation algorithm process for the embodiments of this application.
[0079] As Figure 3 shown, first specify the file path for storing videos, the storage path for segmented pictures, and the number of frames to be extracted. Then, it is judged whether there are still video files. If there are no video files, the processing is directly ended. If there are still video files, the next video is continued to be read. Then, it is segmented according to the specified number of frames, and the segmented results are stored in the specified folder. Then, it enters the judgment loop again until all videos are segmented.
[0080] Specifically, through cooperation with the mine side, monitoring videos from three perspectives of the head of the main shaft manriding, the corner of the main shaft manriding, and the middle part of the main shaft manriding, as well as videos from various angles taken by on-site mobile phones are obtained. The collected videos are frame-extracted every 10 frames using OpenCV. Finally, a total of JPEG format pictures are obtained. Among these pictures, there are a large number of pictures that cannot be used. To ensure the reliability of the dataset, data cleaning is finally carried out through manual screening, and pictures with extremely blurred images, pictures without the targets to be detected, and pictures with too high similarity are screened out.
[0081] Based on this, the visual data annotation tool labelimg is used for manual annotation to obtain a label file in YOLO format. A total of four target categories are annotated: helmet represents wearing a helmet normally, no_helmet represents not wearing a helmet, big_belongings represents taking the monkey car with large items, and no_riding represents the monkey car flight attendant.
[0082] In order to increase the sample quantity and diversity of the dataset MCSVD, this application proposes the RandomFixedMix algorithm to perform offline data augmentation on the dataset. The offline augmentation types include two categories: geometric transformation and color and clarity transformation. At the same time, a certain probability p is set for each augmentation operation to determine whether to execute.
[0083] Before data augmentation, the training set, validation set, and test set are randomly divided according to the ratio of 7:1:2, and the training set is augmented 6 times. At the same time, in order to ensure the ratio, the validation set is supplemented.
[0084] Furthermore, after obtaining the training data, the model is trained. The specific steps are as follows: Based on the RT DETR network, the partial convolution (PConv) is used to improve the BasicBlock in the feature extraction network. Specifically, the second standard convolution layer on the main branch of the original basic residual block BasicBlock is replaced, and thus a more concise and efficient PConvBlock module is constructed.
[0085] Based on the PConvBlock module, the adaptive fine-grained channel attention FCA is introduced. Specifically, the FCA attention mechanism is inserted after the second convolution layer of the main branch of the residual structure, and thus a new PCABlock module is constructed.
[0086] Based on the proposed PConvBlock module and PCABlock module, the feature extraction network of RT DETR is improved. Specifically, the PConvBlock module is used to replace the first two BasicBlocks of the original feature extraction network, and at the same time, the PCABlock module is used to replace the last two BasicBlocks.
[0087] The constructed initial detection model is trained using the training set and the validation set to obtain a trained detection model.
[0088] Furthermore, after the model is trained, the model is used to predict the behavior data.
[0089] Regarding step S101, in some embodiments, the behavior data is preprocessed through the following steps: obtaining a behavior video; extracting video frames from the behavior video to obtain behavior images; and converting the behavior images into deep learning tensors to obtain the behavior data.
[0090] In this embodiment, the trained detection model is used to detect the target to be detected, specifically, the original image or video frame obtained from the camera in the mine environment is converted into a tensor that can be accepted by the mine non-safe behavior detection model.
[0091] Further, regarding step S102, the behavior data is subjected to at least one convolution process to obtain a first feature map.
[0092] As Figure 2 shown, in this embodiment, the behavior data is subjected to three convolution processes. After that, the max pooling process is also performed on the data after the convolution process to reduce the spatial size of the feature map, and a first feature map is obtained.
[0093] In this embodiment, as Figure 2 shown, the input image first enters the feature extraction network, and through a 3x3 convolution layer with a stride of 2, preliminary feature extraction and downsampling are performed. Next, two 3x3 convolution layers are used for further feature extraction; the max pooling layer is used for downsampling to reduce the size of the feature map.
[0094] Further, regarding step S103, the first convolutional residual block is used to perform a convolution process on the first feature map and update the mask to obtain a second feature map.
[0095] In some embodiments, the using the first convolutional residual block to perform a convolution process on the first feature map and update the mask to obtain a second feature map includes: in the first convolutional residual block, performing a first convolution process on the first feature map to obtain a fourth feature map; performing a partial convolution process on the fourth feature map to obtain a fifth feature map; performing a batch normalization process on the fifth feature map to obtain a sixth feature map; using an activation function to process the sixth feature map to obtain a seventh feature map; and obtaining the second feature map based on the seventh feature map.
[0096] In some embodiments, the method further includes: performing a second convolution process on the first feature map to obtain an eighth feature map; performing an average pooling process on the eighth feature map to obtain a ninth feature map; and the obtaining the second feature map based on the seventh feature map includes: obtaining the second feature map based on the seventh feature map and the ninth feature map.
[0097] In this embodiment, as Figure 2As shown in the figure, the feature map is further processed by two partial convolutional residual blocks (PConvBlock), allowing the model to process data with missing or occluded regions. Specifically, a standard convolutional kernel is used to perform a convolutional operation on the feature map, while the mask is updated to ensure that the convolutional operation is only performed in the valid region. The weights of the convolutional kernel are adjusted according to the mask to ensure that the convolutional operation does not produce invalid results in the occluded region. Then, the ReLU activation function is used to perform a non-linear transformation on the convolved feature map. Finally, the input feature map is added to the convolved feature map to ensure the direct transmission of information and prevent the problems of gradient vanishing and explosion.
[0098] Further, for step S104, in some embodiments, the second feature map is convolved using the second convolutional residual block to obtain a third feature map based on the adaptive fine-grained channel attention mechanism, including: in the second convolutional residual block, the second feature map is subjected to a first convolutional process to obtain a tenth feature map; the tenth feature map is subjected to a partial convolutional process to obtain an eleventh feature map; the eleventh feature map is subjected to a batch normalization process to obtain a twelfth feature map; the activation function is used to process the twelfth feature map to obtain a thirteenth feature map; based on the thirteenth feature map, different-frequency features are weighted by an adaptive fusion strategy to obtain a fourteenth feature map; and based on the fourteenth feature map, the third feature map is obtained.
[0099] In some embodiments, the method further includes: performing a second convolutional process on the second feature map to obtain a fifteenth feature map; performing an average pooling process on the fifteenth feature map to obtain a sixteenth feature map; and obtaining the third feature map based on the fourteenth feature map includes: obtaining the third feature map based on the fourteenth feature map and the sixteenth feature map.
[0100] As Figure 2 shown in the figure, the feature map is further processed by two PCABlock residual modules. Different-frequency features are weighted by an adaptive fusion strategy to highlight important features and suppress unimportant features. Then, the ReLU activation function is used to perform a non-linear transformation on the convolved feature map. Finally, the input feature map is added to the convolved feature map to ensure the direct transmission of information and prevent the problems of gradient vanishing and explosion.
[0101] Further, the extracted features are fed into a hybrid encoder. The hybrid encoder contains multiple modules, such as CCFM (Contextual Channel Feature Modulation), for further feature fusion and enhancement. In the CCFM module, the feature map is processed through a series of operations (convolution, upsampling, Concat (connection), etc.) to enhance the expression ability of the features.
[0102] For step S105, the decoder outputs the final detection result, including the category and location information of the target. Based on the detection result output by the detection network, it is determined whether there is a non-safe behavior in the mine.
[0103] In this embodiment, the PConv module: allows the model to process data with missing or occluded regions, which is particularly beneficial for visual tasks in the mine environment because there may be various occlusions or situations where part of the field of view is covered due to equipment; in addition, it can reduce the computational complexity while ensuring a certain accuracy, and is suitable for deployment on resource-constrained edge devices.
[0104] The FCA module: can effectively capture global and local information at different scales, and perform feature weighting through an adaptive fusion strategy, which helps to highlight important features and suppress unimportant features. By finely adjusting the feature weights, the ability of the model to resist noise interference is enhanced, enabling the model to work stably under harsh environmental conditions.
[0105] As can be seen from the above embodiments, the behavior detection method described in the embodiments of the present application includes: obtaining preprocessed behavior data; performing at least one convolution process on the behavior data to obtain a first feature map; using a first convolutional residual block to perform a convolution process on the first feature map and update the mask to obtain a second feature map; using a second convolutional residual block to perform a convolution process on the second feature map, and based on an adaptive fine-grained channel attention mechanism, obtain a third feature map; inputting the third feature map into a hybrid encoder to obtain the final detection result. In the embodiments of the present application, by partially replacing the 3×3 convolutional layer in the original network with a partial convolutional layer (PConv), the number of model parameters is effectively reduced. This not only reduces the computational complexity of the model, but also reduces the memory resources required during the training process, making the model easier to deploy on resource-constrained devices. The design of PConv can more flexibly process the valid information in the feature map, avoiding the indiscriminate processing of invalid regions in traditional convolution operations. This mechanism helps to reduce the redundancy of the network and improve the efficiency of the model. The reduction of the number of parameters and the reduction of network redundancy work together to significantly improve the inference speed of the model. This is particularly important for application scenarios that require real-time detection, such as mine non-safe behavior detection tasks, where rapid response can provide timely warnings and improve safety. The present application introduces an adaptive fine-grained channel attention (FCA) mechanism, which can automatically adjust the importance of each feature in the channel dimension of the feature map. This means that the model can better focus on key features, improving the accuracy and robustness of detection, especially in a complex mine environment, this ability is particularly important. In addition, the traditional residual connection design is retained to ensure the direct transmission of information and prevent the problems of gradient disappearance and explosion.
[0106] It should be noted that the method of the embodiments of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present application, and these multiple devices will interact with each other to complete the described method.
[0107] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0108] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a behavior detection device.
[0109] Reference Figure 4 , the behavior detection device includes:
[0110] An acquisition module 41, configured to acquire preprocessed behavior data;
[0111] A convolution module 42, configured to perform at least one convolution process on the behavior data to obtain a first feature map;
[0112] A first residual module 43, configured to perform a convolution process on the first feature map using a first convolutional residual block and update the mask to obtain a second feature map;
[0113] A second residual module 44, configured to perform a convolution process on the second feature map using a second convolutional residual block and obtain a third feature map based on an adaptive fine-grained channel attention mechanism;
[0114] A detection module 45, configured to input the third feature map into a hybrid encoder to obtain a final detection result.
[0115] For the convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0116] The device of the above embodiment is used to implement the corresponding behavior detection method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0117] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the behavior detection method described in any of the above embodiments.
[0118] Figure 5 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0119] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0120] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0121] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0122] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0123] The bus 1050 includes a path for transmitting information among various components of the device, such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040.
[0124] It should be noted that although only the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050 are shown in the above device, in the specific implementation process, the device may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present specification, and do not have to include all the components shown in the figure.
[0125] The electronic device of the above embodiment is used to implement the corresponding behavior detection method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0126] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the behavior detection method described in any of the foregoing embodiments.
[0127] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0128] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the behavior detection method described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0129] Based on the same inventive concept, corresponding to the behavior detection method described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to execute the behavior detection method. Corresponding to the execution subjects corresponding to the steps in the respective embodiments of the behavior detection method, the processor that executes the corresponding steps can belong to the corresponding execution subject.
[0130] The computer program product of the above embodiment is used to cause the computer and / or the processor to execute the behavior detection method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0131] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; Under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.
[0132] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application will be implemented (that is, these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0133] Although the present application has been described in connection with specific embodiments of the present application, many substitutions, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.
[0134] Embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A behavior detection method, characterized in that: include: Obtain preprocessed behavioral data; Performing at least one convolution process on the behavior data to obtain a first feature map; Using a first convolution residual block to perform convolution processing on the first feature map, and updating the mask to obtain a second feature map; Using a second convolution residual block to perform convolution processing on the second feature map, based on an adaptive fine-grained channel attention mechanism, to obtain a third feature map; The third feature map is input into the hybrid encoder to obtain the final detection result.
2. The method according to claim 1, characterized in that The behavioral data is preprocessed through the following steps: Get behavioral videos; Based on the behavior video, extracting video frames therefrom to obtain a behavior image; The behavior image is converted into a deep learning tensor to obtain the behavior data.
3. The method according to claim 1, characterized in that The using the first convolution residual block to perform convolution processing on the first feature map and updating the mask to obtain the second feature map includes: In the first convolution residual block, performing a first convolution process on the first feature map to obtain a fourth feature map; Performing partial convolution processing on the fourth feature map to obtain a fifth feature map; Performing batch normalization processing on the fifth feature map to obtain a sixth feature map; Using an activation function, processing the sixth feature map to obtain a seventh feature map; Based on the seventh characteristic map, the second characteristic map is obtained.
4. The method according to claim 3, characterized in that The method further comprises: Performing a second convolution process on the first feature map to obtain an eighth feature map; Performing average pooling processing on the eight feature maps to obtain a ninth feature map; The obtaining the second feature map based on the seventh feature map includes: Based on the seventh characteristic map and the ninth characteristic map, the second characteristic map is obtained.
5. The method according to claim 1, characterized in that The method of performing convolution processing on the second feature map by using the second convolution residual block, and obtaining a third feature map based on an adaptive fine-grained channel attention mechanism, includes: In the second convolution residual block, performing a first convolution process on the second feature map to obtain a tenth feature map; Performing partial convolution processing on the tenth feature map to obtain an eleventh feature map; Performing batch normalization processing on the eleventh feature map to obtain a twelfth feature map; Processing the twelfth feature map using an activation function to obtain a thirteenth feature map; Based on the thirteenth feature map, weighted processing is performed on features of different frequencies through an adaptive fusion strategy to obtain a fourteenth feature map; Based on the fourteenth characteristic map, the third characteristic map is obtained.
6. The method according to claim 5, characterized in that The method further comprises: Performing a second convolution process on the second feature map to obtain a fifteenth feature map; Performing average pooling processing on the fifteenth feature map to obtain a sixteenth feature map; The step of obtaining the third characteristic map based on the fourteenth characteristic map comprises: Based on the fourteenth characteristic map and the sixteenth characteristic map, the third characteristic map is obtained.
7. A behavior detection device, characterized in that: include: an acquisition module, configured to acquire the preprocessed behavior data; A convolution module, configured to perform at least one convolution process on the behavior data to obtain a first feature map; A first residual module is configured to perform convolution processing on the first feature map using a first convolution residual block and update a mask to obtain a second feature map; A second residual module is configured to perform convolution processing on the second feature map using a second convolution residual block, and obtain a third feature map based on an adaptive fine-grained channel attention mechanism; The detection module is configured to input the third feature map into the hybrid encoder to obtain a final detection result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 6.