A violent target detection method based on VHPC-DETR
Through the improved VHPC-DETR model, the problems of existing violent target detection models in multi-scale feature fusion and robustness are solved, achieving more accurate and adaptable violence detection and improving social security.
Patent Information
- Application Number
- CN202510919296.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Existing violent target detection models have problems such as insufficient multi-scale feature fusion, poor robustness in complex scenarios, and computational redundancy, making it difficult to effectively detect and prevent violent behavior.
An improved model based on VHPC-DETR is adopted to optimize multi-scale feature fusion and computational efficiency through bidirectional hybrid feature pyramid network, dynamic hierarchical channel interactive convolution, and global scale semantic weaving and elastic feature calibration network, thereby improving the detection accuracy and robustness of the model.
The accuracy and robustness of violence detection have been improved, making it more suitable for violence detection in different scenarios and contents, thus enhancing social security.
Smart Images

Figure CN120411736B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection and social security technology, and specifically is a violent target detection method based on VHPC-DETR. Background Art
[0002] With the rapid development of smart cities and security monitoring systems, real-time detection of violence has become a core requirement in the public safety sector. Currently, violent incidents such as school violence and assaults are common, severely impacting social security. Addressing these issues requires first detecting and detecting violence. However, manual observation not only requires significant time and effort but is also prone to oversight. Violent target detection, by detecting violent behavior and tools, can prevent and promptly curb violence, saving manpower while improving social security.
[0003] Existing methods for detecting aggressive targets primarily rely on CNN-based models (such as the YOLO series) or Transformer architectures (such as the DETR series). For example, some previous methods used YOLOv5 to extract feature frames, followed by DeepSort tracking, and then used the SlowFast network to extract training features. However, these methods still suffer from numerous issues, including insufficient multi-scale feature fusion, poor robustness in complex scenarios, and computational redundancy. Summary of the Invention
[0004] To address the above issues, the present invention provides a violent target detection method based on VHPC-DETR, which is trained using open source datasets to solve existing problems in violent target detection and improve social security.
[0005] A method for detecting a violent target based on VHPC-DETR of the present invention comprises the following steps:
[0006] S1 obtains a violent image dataset and preprocesses the dataset;
[0007] S2 improves the target detection model RT-DETR to obtain the improved VHPC-DETR model;
[0008] S3 trains the constructed VHPC-DETR target detection model;
[0009] S4 uses the trained detection model to perform violent target detection.
[0010] Furthermore, step S1 specifically involves obtaining labeled violent image data, organizing the image data set, and then performing data augmentation and flipping operations. The data can be self-acquired labeled images containing violence, or a public violent image data set.
[0011] Furthermore, in step S2, the improvements of the improved VHPC-DETR model compared to the basic model RT-DETR include three parts: a bidirectional hybrid feature pyramid network, a dynamic hierarchical channel interaction convolution part, and a global scale semantic weaving and elastic feature calibration network.
[0012] Furthermore, the bidirectional hybrid feature pyramid network takes feature maps of different scales as input, which are then processed by a feature alignment module. First, the high-level feature maps are transposed and convolved to expand the output size to the P4 / 16 level, and then the low-level feature maps P3 / 8 features are downsampled to the P4 / 16 level through conventional convolution.
[0013] The aligned features are passed through the channel attention mechanism to generate dynamic weights. The calculation formula is:
[0014] W TD =σ(Conv2(ReLU(Conv1(F align )))) (1)
[0015] Where W TD is the channel attention weight matrix, σ is the Sigmoid function, Conv2 is a 1×1 convolution layer used to achieve channel recovery, ReLU is the activation function, Conv1 is a 1×1 convolution layer used to achieve channel compression, F align The aligned feature maps are then passed through the information injection module and the information fusion module to achieve multi-scale feature optimization.
[0016] Furthermore, the information injection module is mainly implemented through the channel attention module (ChannelAttention_HSFPN) and a series of convolution operations. The channel attention module first performs adaptive average pooling and maximum pooling on the input features in parallel to generate a 1×1 channel descriptor. Then, two layers of 1×1 convolution are shared for channel interaction. The first layer performs channel compression (ratio=4, such as 256→64) and uses ReLU activation. The second layer restores the channel (64→256). After adding the two-way output, the sigmoid function is used to generate a channel attention map. The aligned features are then compared with W. TD Multiply them together to generate attention-enhanced features, which are further optimized by 3×3 convolution (step 1, padding 1).
[0017] Furthermore, in the information fusion module, in the top-down path, the bottom-level features (P4 / 16-256C) are multiplied by the high-level down-transmitted features (P5 / 32→P4 / 16-256C) after channel attention processing, and then added to the original P5 / 32 features to output 256C features. In the bottom-up path, the middle-level features (P4 / 16-256C) are enhanced by global attention (GA) and multiplied by the lower-level features (P3 / 8→P4 / 16-256C) to output 256C features.
[0018] Furthermore, the dynamic hierarchical channel interaction convolution component (BasicPCock) is an improvement achieved by integrating the basic blocks of the last two layers (layers 6 and 7) of the original model's backbone network with the channel grouping convolution. After the image enters this part of the model, the BasicPCock network begins by receiving input features. The input features are processed along different paths. First, a branch of the input features points to the shortcut (parameter judgment) path. If the parameter is True, the input features are directly passed through this path to the subsequent merging step; if it is False, this path has no effect. Another branch of the input features enters. In the PCLPH (general processing path), the features first pass through a convolutional layer, which performs a convolution operation on the input features to extract features. The convolved features are then batch normalized. The normalized features are then activated by the ReLU function to introduce nonlinearity. The input features also have a branch that enters the SPRPH (special processing path). In this path, the core components are introduced to perform partial convolution operations using improved basic blocks. The deep features use BasicPCock to maintain standard convolution for some shallow features, maintain position sensitivity, and finally merge the path elements into an output.
[0019] Furthermore, the operation of the improved basic block includes the following steps: first, performing group convolution on the first 1 / 4 channels, and then fusing the output features with the residual connection through channel shuffle, as follows:
[0020] F out =Shuffle(Y)+X [1:cp] (2)
[0021] Among them F out is the output feature of the module, Y is the convolution result of the first 1 / 4 channel, X [1:cp] is the original feature of the remaining 3 / 4 channels.
[0022] Furthermore, the global-scale semantic weaving and elastic feature calibration network adjusts the CFC_CRB module and SFC_G2 module in the context and spatial feature calibration network according to the original model architecture, and integrates them with the feature fusion and feature enhancement parts in the neck network in the basic model to obtain the adjusted CFC_CRB module and the adjusted SFC_G2 module.
[0023] Furthermore, the adjusted CFC_CRB module first compresses the input features using 3×3 convolution (with a stride of 1 and padding of 1) and extracts multi-scale contextual features through cascaded pyramid pooling (PSP). The pooling grid size is [6, 3, 2, 1]. The query vector Q is constructed through 1×1 convolution, the similarity matrix is calculated and normalized, and the fusion value vector is used to reconstruct the contextual features. The channel-space dual attention mechanism is then used for processing. Finally, the residual connection retains the original features. The formula is:
[0024] F calibrated =F+Softmax(QK T )V (4)
[0025] Among them, F calibrated is the calibrated feature, F is the original feature input to the module, K T is the transpose of K, Softmax is the activation function, Q=Conv q (F),K,V=PSP(F), where Conv q (F) is a 1×1 convolution operation used to convert the input feature F into a query vector Q. PSP(F) represents a cascaded pyramid pooling operation on the feature F to obtain the key vector K and the value vector V.
[0026] Furthermore, in the adjusted SFC_G2 module, the shallow features are subjected to 3×3 convolution (with a stride of 1 and a padding of 1) to preserve details, and the deep features are subjected to convolution (with clear convolution parameters) and interpolation to jointly predict the network to generate offset injection information. The specific operation is as follows: first, the shallow features are combined with the predicted offset Δp, then a dynamic sampling grid is generated based on the offset, features are aligned through bilinear interpolation, and finally, Tanh activation is used to generate fusion weights α, β, satisfying α + β = 1, and features are fused according to the weights; wherein, the formula for predicting the offset Δp is:
[0027] Δp=Convoffset(Concat(Fcp,Fsp)) (5)
[0028] Conv offset is a convolutional layer used to predict the offset, Concat(F cp ,F sp ) is to transform the shallow feature Fcp and deep features F sp The splicing operation is performed on the channel dimension, and the formula for weighted fusion features is:
[0029] F fuse =αF sp +βF cp (6)
[0030] F fuse is the fused feature, F sp is the deep feature, F cp It is a shallow feature.
[0031] The beneficial effects of the present invention are:
[0032] The accuracy and robustness of violence detection are improved, and the method is more suitable for violence detection in different scenarios and with different contents. The present invention can detect violence more accurately and improve social security. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flow chart of the present invention.
[0034] Figure 2 This is the overall architecture diagram of the VHPC-DETR model adopted in the present invention.
[0035] Figure 3 It is the bidirectional hybrid feature pyramid network model of the present invention.
[0036] Figure 4 This is the partial convolution module used in the present invention.
[0037] Figure 5 This is the schematic diagram of the adjusted CFC_CRB module.
[0038] Figure 6 This is the schematic diagram of the adjusted SFC_G2 module.
[0039] Figure 7 This is a comparison diagram of the effects of the original model RT-DETR and the improved model VHPC-DETR in Example 1 of the present invention.
[0040] Figure 8 It is the MAP value of the experiment in Example 1 of the present invention.
[0041] In the figure: backbone network, ConvBN convolution batch normalization layer, MaxPool maximum pooling layer, BasicBlok basic residual block, BasicPCock dynamic stage channel interaction convolution part, channelAttention_HSFPN channel attention module, nn.Conv2d two-dimensional convolution layer, Add addition operation, RepC3 repeated C3 module, Multiply multiplication operation, nn.ConvTranspose2d two-dimensional transposed convolution layer, RTDETRDecoder basic model RT-DETR decoder, Conv convolution layer, AIFI adaptive feature interaction module, Top-Down path top-down path, LA local attention, GA global attention, another representation of FO RepC3, the following brackets indicate the operation, Bottom-Up Path bottom-up path, Input input, Shortcut parameter judgment, PCLPH general processing path, BatchNorm2d two-dimensional batch normalization layer, ReLU activation function, SPRPH special processing path, branch2b branch 2b, dim_untouched dimension remains unchanged, Output output, query_conv query operation convolution layer, calculation of the similarity between the query vector (Query) and the key vector (Key), Sim(Query*Key) two feature similarity calculation, Key PSPModule combines the structure of the key vector and the pyramid pooling module, Value PSPModule combines the structure of the value vector and the pyramid pooling module, Value*Sim value and similarity product, LocalAttenModule local attention module, Grid dynamic sampling grid, conv_offset convolution offset, Concat(cp+sp) splicing (conventional path features + special path features), Fusion fusion, l convolution offset obtained feature l, h convolution offset obtained feature h, Alt convolution offset obtained feature Alt, alt convolution offset obtained feature alt. DETAILED DESCRIPTION
[0042] It should be noted that in the description of the present invention, the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", "clockwise", "counterclockwise", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction.
[0043] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of a violent target detection system and method based on VHPC-DETR of the present invention in conjunction with the accompanying drawings.
[0044] The present invention uses the improved model VHPC-DETR to detect violent behaviors and tools, thereby preventing and stopping violence in a timely manner, reducing the occurrence of violence, and enhancing social security. Example
[0045] RT-DETR is a commonly used model in the existing technology, see https: / / arxiv.org / abs / 2304.08069.
[0046] VHPC-DETR is an improved model of the present invention. It is not an innovative part of the present invention and no references can be cited.
[0047] Figure 1 A flowchart of a method for detecting a violent target based on VHPC-DETR according to an embodiment of the present invention is provided. In the present invention, we take the Violence-Image-DatasetPu dataset as an example to elaborate on a method for detecting a violent target based on VHPC-DETR according to an embodiment of the present invention. Figure 1 The specific steps are as follows:
[0048] S1-1: Obtain a labeled Violence-Image-DatasetPu image dataset, where the Violence-Image-DatasetPu image dataset is a labeled Violence-Image-DatasetPu image dataset obtained from an open source dataset;
[0049] The specific content of the dataset is published in
[0050] https: / / github.com / ChinaZhangPeng / Violence-Image-Dataset
[0051] Bianculli, M.; Falcionelli, N.; Sernani, P.; Tomassini, S.; Contardo,P.; Lombardi, M.; Dragoni, AFJDib A dataset for automatic violence detection in videos. 2020, 33, 106587.
[0052] S1-2: Preprocess the acquired images, uniformly adjust the image size, and perform normalization to map the image pixel values to a specific numerical range to meet the input requirements of the model.
[0053] S2-1: Improve the original target detection model RT-DETR to generate an improved model VHPC-DETR. The overall architecture of VHPC-DETR used in this invention is shown in the figure below: Figure 2 shown
[0054] S2–2: The original model improves the dynamic hierarchical channel interaction convolution BasicPCock portion. After the image enters this portion of the model, the BasicPCock network begins by receiving input features, which are then processed along different pathways. First, a branch of the input features points to the shortcut path, though whether this path is enabled depends on its parameter. If the parameter is True, the input features are directly passed through this path to the subsequent merging step; if it is False, this path has no effect. Another branch of the input features enters. In PCLPH, features first pass through a convolutional layer, which convolves the input features to extract features. The convolved features are then batch normalized, which accelerates model convergence and improves stability. The normalized features are then activated using the ReLU function, introducing nonlinearity and enhancing the model's expressiveness. The input features also have a branch that enters the SPRPH. In this pathway, the core component uses improved basic blocks for partial convolution operations, while deep features use the BasicPCock portion. This partial convolution reduces computational effort by approximately 25% while preserving detailed features. The shallow features maintain standard convolution to maintain position sensitivity. Finally, the path elements are merged and output.
[0055] The operation of the improved basic block includes the following steps: first, group convolution is performed on the first 1 / 4 channels, and then the output features are fused with residual connections through channel shuffle. The formula is:
[0056] W TD =σ(Conv2(ReLU(Conv1(F align )))) (1)
[0057] Where W TD is the channel attention weight matrix, σ is the Sigmoid function, Conv2 is a 1×1 convolution layer used to achieve channel recovery, ReLU is the activation function, Conv1 is a 1×1 convolution layer used to achieve channel compression, F align The aligned feature maps are then passed through the information injection module and the information fusion module to achieve multi-scale feature optimization.
[0058] S2-3: Improve the original model with a bidirectional hybrid feature pyramid network. After the image enters this part of the model, the high-level feature map (P5 / 32) is first transposed convolution (kernel=3×3, stride=2, padding=1, output_padding=1), and the output size is expanded to the P4 / 16 level (the size changes from H / 32×W / 32 to H / 16×W / 16). Then, channel attention weighting is performed. The calculation formula is:
[0059] F out =Shuffle(Y)+X [1:cp] (2)
[0060] Among them F out is the output feature of the module, Y is the convolution result of the first 1 / 4 channel, X [1:cp] is the original feature of the remaining 3 / 4 channels, and then feature fusion is performed. The formula is:
[0061] F fuse =F low ×W CA +F align (3)
[0062] Then, the bottom-up operation is performed. First, the low-level features are downsampled, and the P3 / 8 features are downsampled to the P4 / 16 level through 3×3 convolution (Stride=2). Then, local attention enhancement is performed, and finally feature optimization is performed.
[0063] S2-4: S2-4: Construct a global-scale semantic weaving and elastic feature calibration network. After the image enters this part of the model, the adjusted CFC_CRB module first compresses the input features using 3×3 convolution (stride 1, padding 1) to compress the channels. Multi-scale contextual features are extracted through cascaded pyramid pooling (PSP) with a pooling grid size of [6, 3, 2, 1]. The query vector Q is constructed through 1×1 convolution, the similarity matrix is calculated and normalized, and the fused value vector is reconstructed into contextual features. Then, the channel-spatial dual attention mechanism is used for processing. Finally, the residual connection is used to retain the original features. The formula is:
[0064] F calibrated =F+Softmax(QK T )V (4)
[0065] Among them, Fcalibrated is the calibrated feature, F is the original feature input to the module, K T is the transpose of K, Softmax is the activation function, Q=Conv q (F),K,V=PSP(F), where Conv q(F) is a 1×1 convolution operation used to convert the input feature F into a query vector Q. PSP(F) represents a cascaded pyramid pooling operation on the feature F to obtain the key vector K and the value vector V.
[0066] The adjusted SFC_G2 module then performs a 3×3 convolution (with a stride of 1 and padding of 1) on shallow features to preserve details. Deep features are then convolved (with explicit convolution parameters) and interpolated before being fed into the prediction network to generate offset injection information. Specifically, the shallow features are combined with the predicted offset Δp. A dynamic sampling grid is then generated based on the offset. Features are aligned using bilinear interpolation. Finally, Tanh activation is used to generate fusion weights α and β, satisfying α + β = 1, and the features are fused accordingly. The formula for the predicted offset Δp is:
[0067] Δp=Conv offset (Concat(F cp ,F sp )) (5)
[0068] Conv offset is a convolutional layer used to predict the offset, Concat(F cp ,F sp ) is to transform the shallow feature F cp and deep features F sp The splicing operation is performed on the channel dimension. The formula for weighted feature fusion is:
[0069] F fuse =αF sp +βF cp (6)
[0070] F fuse is the fused feature, F sp is the deep feature, F cp It is a shallow feature.
[0071] S3-1: Divide the preprocessed Violence-Image-DatasetPu dataset into training set, validation set, and test set in a ratio of 7:1:2.
[0072] S3-2: Configure the training environment. The experiment uses Python 3.9 on a Windows system, equipped with an 18-vCPU AMD EPYC 9754 128-Core Processor, an RTX 3090 (24GB) GPU, 60GB of video memory, and CUDA 11.7.
[0073] S3-3 trains the processed dataset in the environment. During training, periodically evaluate the model using the validation set, calculating the loss and evaluation metrics on the validation set. Based on the evaluation results, adjust training parameters, such as the learning rate and early stopping, to prevent overfitting and ensure good performance on the validation set.
[0074] Model testing and application phase (S4-1 to S4-2):
[0075] S4-1: After training is complete, the model is tested using the test set. Images from the test set are fed into the trained model, and the model outputs predicted object categories and locations. Evaluation metrics such as mAP, recall, and precision are calculated on the test set. The calculation formula is:
[0076] (7)
[0077] (8)
[0078] (9)
[0079] mAP = (10)
[0080] S4-2: Compare the effects of the improved model VHPC-DETR and the basic model RT-DETR.
[0081] Table 1 Comparison of results between the new model and the original model
[0082] method Box(P R mAP50 RTDETR 0.822 0.802 0.847 VHPC-DETR 0.838 0.822 0.859
[0083] S4-3: Compare the improved model VHPC-DETR and the classic target detection model YOLO model.
[0084] Table 2 Comparison of the effects of other target detection models and the new model in the violent target task
[0085] method Violence-Image-DatasetPu Yolov5 0.802 Yolov6 0.812 Yolov8 0. 803 Yolov10 0.789 Yolov11 0.832 RTDETR 0.847 VHPC-DETR 0.859
[0086] In summary, the accuracy of this model is higher than that of the existing violent target detection model, and its effect is more applicable.
[0087] The foregoing description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by any person skilled in the art within the technical scope disclosed herein and within the spirit and principles of the present invention shall be covered by the scope of protection of the present invention. Furthermore, any matters not described in detail in this specification constitute prior art known to those skilled in the art.
[0088] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
Claims
1. A violent target detection method based on VHPC-DETR, characterized in that: The following steps are involved: S1 obtains a violent image dataset and preprocesses the dataset; S2 improves the target detection model RT-DETR to obtain the improved VHPC-DETR model; S3 trains the constructed VHPC-DETR target detection model; S4 uses the trained detection model to detect violent targets; In step S2, the improved VHPC-DETR model compared to the basic model RT-DETR includes three parts: a bidirectional hybrid feature pyramid network, a dynamic hierarchical channel interaction convolution part, and a global scale semantic weaving and elastic feature calibration network. The bidirectional hybrid feature pyramid network takes feature maps of different scales as input and processes them through the feature alignment module. First, the high-level feature map is transposed and convolved to expand the output size to the P4 / 16 level. Then, the low-level feature map P3 / 8 features are downsampled to the P4 / 16 level through conventional convolution. The aligned features are then used to generate dynamic weights through the channel attention mechanism. The calculation formula is: W TD =σ(Conv2(ReLU(Conv1(F align )))) (1) Where W TD is the channel attention weight matrix, σ is the Sigmoid function, Conv2 is a 1×1 convolution layer used to achieve channel recovery, ReLU is the activation function, Conv1 is a 1×1 convolution layer used to achieve channel compression, F align The aligned feature maps are then passed through the information injection module and the information fusion module to achieve multi-scale feature optimization; The dynamic hierarchical channel interactive convolution part is obtained by improving the basic blocks of the last two layers of the backbone network part of the original model and fusing the channel grouping part of the convolution. After the image enters this part of the model, the dynamic hierarchical channel interactive convolution part network starts from receiving the input features, and the input features will be processed along different paths. First, the input feature has a branch pointing to the parameter judgment path. If the parameter is True, the input feature will be directly passed to the subsequent merging step through this path; if the parameter is False, the path has no practical effect, and another branch of the input feature enters. In the general processing path, the feature first passes through a convolution layer, which performs a convolution operation on the input feature to extract the feature. The convolutional feature is then batch normalized. The normalized feature passes through the ReLU activation function to introduce nonlinear factors. The input feature also has a branch entering the special processing path. The core component is introduced to use the improved basic block for partial convolution operation. The deep feature uses the dynamic hierarchical channel interactive convolution part, and the shallow feature maintains the standard convolution to maintain position sensitivity. Finally, the path elements are merged and output; The global-scale semantic weaving and elastic feature calibration network adjusts the CFC_CRB module and SFC_G2 module in the context and spatial feature calibration network according to the original model architecture, and fuses them with the feature fusion and feature enhancement parts in the neck network in the basic model to obtain the adjusted CFC_CRB module and the adjusted SFC_G2 module.
2. A violent target detection method based on VHPC-DETR according to claim 1, characterized in that, Step S1 specifically involves obtaining labeled violent image data, organizing the image data set, and then performing operations such as data augmentation and flipping. The data can be labeled images containing violence obtained by oneself or a public violent image data set.
3. A violent target detection method based on VHPC-DETR according to claim 1, characterized in that, The information injection module is implemented through a channel attention module and a convolution operation. The channel attention module first performs adaptive average pooling and maximum pooling on the input features in parallel to generate a 1×1 channel descriptor. Then, two layers of 1×1 convolution are shared for channel interaction. The first layer performs channel compression and uses ReLU activation. The second layer performs channel restoration. After adding the two-way output, a channel attention map is generated by the Sigmoid function. The aligned features are then compared with W. TD Multiply them together to generate attention-enhanced features, which are further optimized by 3×3 convolution.
4. A method for detecting violent targets based on VHPC-DETR according to claim 1, characterized in that: In the information fusion module, in the top-down path, the bottom-level features are processed by channel attention and multiplied with the high-level downward features, and then added to the original features to output the features. In the bottom-up path, the middle-level features are enhanced by global attention and multiplied with the lower-level features to output the features.
5. A method for detecting violent targets based on VHPC-DETR according to claim 1, characterized in that: The operation of the improved basic block includes the following steps: first, group convolution is performed on the first 1 / 4 channels, and then the output features are fused with residual connections through channel shuffle. The formula is: F out =Shuffle(Y)+X [1:cp] (2) Among them F out is the output feature of the module, Y is the convolution result of the first 1 / 4 channel, X [1:cp] is the original feature of the remaining 3 / 4 channels.
6. A method for detecting violent targets based on VHPC-DETR according to claim 1, characterized in that: The adjusted CFC_CRB module first compresses the input features using 3×3 convolution channels, extracts multi-scale contextual features through cascaded pyramid pooling, and the pooling grid size is [6, 3, 2, 1]. The query vector Q is constructed through 1×1 convolution, the similarity matrix is calculated and normalized, and the fusion value vector is used to reconstruct the contextual features. The channel-space dual attention mechanism is then used for processing, and the residual connection is used to retain the original features. The formula is: F calibrated =F+Softmax(QK T )V (4) Among them, F calibrated is the calibrated feature, F is the original feature input to the module, K T is the transpose of K, Softmax is the activation function, Q = Conv q (F),K,V=PSP(F), where Conv q (F) is a 1×1 convolution operation used to convert the input feature F into a query vector Q. PSP(F) represents a cascaded pyramid pooling operation on the feature F to obtain a key vector K and a value vector V. In the adjusted SFC_G2 module, shallow features are subjected to 3×3 convolution to preserve details, and deep features are subjected to convolution and interpolation to jointly predict the network to generate offset injection information. First, the shallow features are combined with the predicted offset Δp, and then a dynamic sampling grid is generated based on the offset. The features are aligned through bilinear interpolation, and finally Tanh activation is used to generate fusion weights α and β, satisfying α + β = 1, and the features are fused according to the weights. The formula for predicting the offset Δp is: Δp=Conv offset (Concat(F cp ,F sp )) (5) Conv offset is a convolutional layer used to predict the offset, Concat(F cp ,F sp ) is to transform the shallow feature F cp and deep features F sp The splicing operation is performed on the channel dimension, and the formula for weighted fusion features is: F fuse =αF sp +βF cp (6) F fuse is the fused feature, F sp is the deep feature, F cp It is a shallow feature.
Citation Information
Patent Citations
Improved target detection method in automatic driving scene based on RT-DETR
CN118644824A
Rail operator safety wearing detection method, device and equipment and storage medium
CN119992463A