Goods dumping detection method based on large convolution reverse residual network

The multi-scale cargo tipping detection method using a large convolutional inverse residual network combines the inverse residual attention module and the spatial pyramid perception convolution module with the Focal-SIoU loss function to solve the problem of low accuracy in cargo tipping detection under imbalanced samples, achieving higher detection accuracy and reliability.

CN120876835APending Publication Date: 2025-10-31CHONGQING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511042470.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In intelligent detection of cargo tipping, existing methods struggle to effectively improve the detection accuracy of a small number of normal samples under unbalanced sample conditions, resulting in low detection accuracy and a high risk of misjudgment.

Method used

A multi-scale cargo tipping detection method based on a large convolutional inverse residual network is adopted. By combining the backbone network layer, the neck network layer and the detection head, the inverse residual attention module and the spatial pyramid perception convolution module are used to enhance feature extraction and fusion. The model training is optimized by combining the Focal-SIoU loss function to improve the accuracy and reliability of cargo tipping detection.

Benefits of technology

It improves the accuracy of cargo tipping detection under unbalanced samples, reduces the false detection rate, enhances the sensitivity and adaptability to cargo tipping conditions, and ensures the reliability and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876835A_ABST
    Figure CN120876835A_ABST
Patent Text Reader

Abstract

The invention discloses a cargo dumping detection method based on a large convolution reverse residual network, and the method comprises the steps: employing an optimized and constructed multi-scale cargo dumping detection network model, enabling a backbone network layer to pass through an initial convolution layer and a C2f module, enabling the backbone network layer to pass through a reverse residual attention iRMB module, enhancing the feature extraction capability of a cargo dumping region, and enabling the feature extraction capability of the cargo dumping region to be improved; global information is efficiently captured through a spatial pyramid perception convolution SPPF-UR module, so that multi-scale cargo dumping information is obtained. A feature map enters a neck network layer, and multi-scale features are effectively integrated through up-sampling and feature splicing. And finally, at a detection head, through the concern of a Focal Loss and SIOU enhanced Focal-SIOU loss function on the dumping positioning precision, the overall detection performance under an unbalanced sample is improved, finally, information such as the target category, the position and the confidence coefficient of the detected goods is output, and the accuracy and the reliability of goods dumping detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of visual image processing technology and neural network model technology, specifically to a cargo tipping detection method based on a large convolutional inverse residual network. Background Technology

[0002] In the intelligent detection of cargo tipping, some scenarios present an imbalanced sample problem where the number of tipped samples far exceeds the number of normal samples, easily leading to low detection accuracy for the few normal samples. Therefore, introducing an intelligent cargo tipping recognition model based on a large convolutional inverse residual network under imbalanced sample conditions during the cargo tipping process improves the detection accuracy for the few normal samples, quickly and accurately identifies tipping scenarios, reduces human error, and accurately determines the cargo status, which is of great significance for ensuring the safe and reliable operation of cargo tipping detection.

[0003] In recent years, some scholars have proposed imbalanced sample detection methods based on deep learning and have made some progress. Existing methods are mainly divided into two categories: data-level and model-level. Data-level methods commonly employ strategies such as data augmentation and data sampling. However, data-level methods fail to fully extract the deep semantic information of the data and struggle to effectively control the types of generated samples, resulting in low data augmentation efficiency. Therefore, scholars have begun to optimize the imbalance problem from a model perspective. Among these, single-stage target detection algorithms are widely used due to their fast detection speed and ability to better meet the detection needs of practical engineering. However, existing methods have weak capabilities in extracting global features and tend to favor majority class samples under imbalanced sample conditions. There is still room for improvement in the accuracy of existing cargo dumping target detection methods. Summary of the Invention

[0004] To address the shortcomings of the existing technology, this invention provides a cargo tipping detection method based on a large convolutional inverse residual network, thereby improving the accuracy and reliability of cargo tipping detection.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0006] A cargo tipping detection method based on a large convolutional inverse residual network is disclosed. The method acquires an image of the cargo to be detected and inputs it into a pre-trained multi-scale cargo tipping detection network model to obtain the cargo tipping state detection result of the cargo image. The multi-scale cargo tipping detection network model includes a backbone layer, a neck layer, and a detection head. The backbone layer extracts multi-scale features from the cargo image input to the multi-scale cargo tipping detection network model to obtain a multi-scale cargo tipping feature map. The neck layer performs feature fusion processing on the multi-scale cargo tipping feature map to obtain a fused feature map. The detection head obtains the cargo tipping state detection result of the cargo image based on the fused feature map, which serves as the output of the multi-scale cargo tipping detection network model.

[0007] As a preferred embodiment, in the multi-scale cargo dumping detection network model:

[0008] The backbone network layer includes a series of convolutional layers, four convolutional feature extraction units, one inverted residual attention module, and one spatial pyramid perceptual convolutional module; each convolutional feature extraction unit includes a cascaded convolutional layer and a C2f module.

[0009] The neck network layer includes two feature fusion upsampling units and two convolutional connection feature fusion units connected in sequence. The feature fusion upsampling unit includes an upsampling module, a connection layer, and a C2f module connected in sequence; the convolutional connection feature fusion unit includes a convolutional layer, a connection layer, and a C2f module connected in sequence.

[0010] The detection head is used to obtain the cargo tipping state detection result of the cargo image based on the output of the neck network layer, and as the output of the multi-scale cargo tipping detection network model;

[0011] In this system, the input of the multi-scale cargo dumping detection network model serves as the input of the backbone network layer. Within the backbone network layer, the outputs of the second and third convolutional feature extraction units also serve as the inputs to the connection layers of the second and first feature fusion upsampling units in the neck network layer, respectively. In addition to serving as the input of the spatial pyramid perception convolutional module in the backbone network layer to the neck network layer, the output of the spatial pyramid perception convolutional module also serves as the input to the connection layer of the second convolutional connection feature fusion unit in the neck network layer. Furthermore, within the neck network layer, the output of the first feature fusion upsampling unit also serves as the input to the connection layer of the first convolutional connection feature fusion unit, and the outputs of the second feature fusion upsampling unit and the two convolutional connection feature fusion units all serve as the inputs to the detection head.

[0012] As a preferred embodiment, the inverse residual attention module includes a moving block branch with an integrated attention mechanism, and an inverse residual branch consisting of a cascaded first 1×1 convolutional layer, a 3×3 depthwise separable convolutional layer, and a second 1×1 convolutional layer. The input feature map of the inverse residual attention module serves as the input to the moving block branch and the inverse residual branch, respectively. The moving block branch is used to extract the attention matrix from the input feature map and output it to the inverse residual branch. In the inverse residual branch, the output of the first 1×1 convolutional layer is multiplied by the attention matrix output from the moving block branch to obtain an attention convolutional feature map, which serves as the input to the 3×3 depthwise separable convolutional layer. The output of the 3×3 depthwise separable convolutional layer is then added to the attention convolutional feature map and used as the input to the second 1×1 convolutional layer. The output of the second 1×1 convolutional layer is then added to the input feature map of the inverse residual attention module and used as the output of the inverse residual attention module.

[0013] As a preferred embodiment, the spatial pyramid perceptual convolutional module includes a cascaded convolutional layer, three two-dimensional max-pooling modules, a connection layer, and a perceptual large-kernel convolutional module. The outputs of the convolutional layer and the three two-dimensional max-pooling modules, in addition to being cascaded, also serve as inputs to the connection layer. After being concatenated by the connection layer, they are output to the perceptual large-kernel convolutional module, whose output serves as the output of the spatial pyramid perceptual convolutional module. The input feature map of the spatial pyramid perceptual convolutional module undergoes feature extraction processing through the convolutional layer, and then sequentially passes through the three two-dimensional max-pooling modules for max-pooling operations before being passed to the connection layer. The outputs of the convolutional layer and the three two-dimensional max-pooling modules also serve as inputs to the connection layer. After being concatenated by the connection layer, they are output to the perceptual large-kernel convolutional module for further global feature extraction, resulting in the output of the spatial pyramid perceptual convolutional module.

[0014] As a preferred embodiment, the perceptual large kernel convolutional module includes four feature extraction stages: feature extraction stage 1, feature extraction stage 2, feature extraction stage 3, and feature extraction stage 4. Adjacent feature extraction stages are connected by a transition segment, which is a 3×3 convolutional layer with a stride of 2. Each feature extraction stage includes several cascaded kernel units, each of which includes one cascaded large kernel block (LarK Block) and two small kernel blocks (SmaKBlock). The input feature map of the perceptual large kernel convolutional module is first downsampled and channel-expanded through a 3×3 convolutional layer with a stride of 4, and then output to feature extraction stage 1. It is then processed sequentially through several of the four feature extraction stages and 3×3 convolutional layers of length 2 between adjacent feature extraction stages. The feature map obtained from feature extraction stage 4 is then used as the output of the perceptual large kernel convolutional module.

[0015] As a preferred embodiment, the SmaK Block comprises a cascaded 3×3 depthwise separable convolutional layer, a batch normalization layer, a compression activation layer, a feedforward network layer, and another batch normalization layer. The input feature map of the SmaK Block is processed sequentially through the 3×3 depthwise separable convolutional layer, a batch normalization layer, a compression activation layer, a feedforward network layer, and another batch normalization layer to obtain a feature map, which is then added to the input feature map of the SmaK Block to obtain the output feature map of the SmaK Block. The LarK Block comprises a cascaded dilated reparameterization layer, a batch normalization layer, a compression activation layer, a feedforward network layer, and another batch normalization layer. The input feature map of the LarK Block is processed sequentially through the dilated reparameterization layer, a batch normalization layer, a compression activation layer, a feedforward network layer, and another batch normalization layer to obtain a feature map, which is then added to the input feature map of the LarK Block to obtain the output feature map of the LarK Block.

[0016] As a preferred embodiment, the dilated reparameterization layer includes five dilated convolution modules connected in parallel, including one dilated convolution module consisting of a cascaded 9×9 dilated convolution block and a batch normalization module, one dilated convolution module consisting of a cascaded 5×5 dilated convolution block and a batch normalization module, and three dilated convolution modules consisting of a cascaded 6×6 dilated convolution block and a batch normalization module. The input feature map of the dilated reparameterization layer is processed by the five dilated convolution modules to obtain dilated convolution maps, which are then converted into large-kernel convolution feature maps through structural reparameterization and used as the output of the dilated reparameterization layer.

[0017] As a preferred embodiment, the multi-scale cargo tipping detection network model is trained in the following manner:

[0018] S101: Obtain the cargo image sample dataset, label both the cargo tipping image samples and the cargo normal image samples in the cargo image sample dataset, and divide them into training set and test set according to a preset ratio;

[0019] S102: Input the training set into the multi-scale cargo dumping detection network model for training, and optimize the parameters of the multi-scale cargo dumping detection network model with the goal of minimizing the loss function until the multi-scale cargo dumping detection network model converges, thus obtaining the trained multi-scale cargo dumping detection network model.

[0020] S103: Test the multi-scale cargo dumping detection network model using the test set to confirm the recognition performance of the trained multi-scale cargo dumping detection network model; if the recognition performance meets the requirements, end the training of the multi-scale cargo dumping detection network model; otherwise, return to step S102.

[0021] As a preferred option, the Focal-SIoU loss function is selected for localization regression loss, and the BCE loss function is selected for classification loss.

[0022] As a preferred embodiment, the expression for the Focal-SIoU loss function is:

[0023] L Focal-SIoU =IoU γ L SIoU ;

[0024] Among them, L Focal-SIoU Indicates Focal-SIoU loss; IoU, L SIoU Let IoU loss and SIoU loss be represented respectively; γ is the adjustment factor; and we have:

[0025]

[0026] γ = 2 - Λ;

[0027] Among them, B, B gt Let Area(·) represent the predicted cargo dumping area location frame and the actual cargo dumping area location frame, respectively; Λ represents the angle cost between the predicted and actual cargo dumping area location frames; Δ represents the center distance cost between the predicted and actual cargo dumping area location frames; Ω represents the cost of the overlapping area between the predicted and actual cargo dumping area location frames; and we have:

[0028]

[0029] Where e is the natural constant; θ is a preset natural number parameter; and parameter ρ x ρ y ω w ω h They are as follows:

[0030]

[0031] Among them, c x (b,b gt ), c y (b,b gt ) represent the center point b of the predicted cargo dumping area location frame and the center point b of the actual cargo dumping area location frame, respectively.gt The x-coordinate distance and y-coordinate distance between them; ρ(b,b) gt () represents the center point b of the predicted cargo dumping area positioning frame and the center point b of the actual cargo dumping area positioning frame. gt The Euclidean distance between them; w and h represent the width and height of the predicted cargo tipping area positioning frame, respectively; w gt h gt These represent the width and height of the positioning frame for the actual cargo dumping area, respectively; C w C h These represent the width and height of the smallest bounding rectangle covering the predicted cargo dumping area location frame and the actual cargo dumping area location frame, respectively.

[0032] Compared with the prior art, the present invention has the following technical effects:

[0033] 1. This invention presents a cargo tipping detection method based on a large convolutional inverse residual network. Utilizing an optimized multi-scale cargo tipping detection network model, the backbone layer (Backbone) passes through an initial convolutional layer and a C2f module, followed by an inverse residual attention (iRMB) module to enhance feature extraction of the cargo tipping area. Then, a spatial pyramid perceptual convolutional (SPPF-UR) module efficiently captures global information to obtain multi-scale cargo tipping information. The feature map enters the neck layer (Neck), where upsampling and feature concatenation effectively integrate multi-scale features. Finally, in the detection head (Head), a Focal Loss combined with a SIoU-enhanced Focal-SIoU loss function focuses on the accuracy of the tipping location, thereby improving the overall detection performance under imbalanced samples. The final output includes the detected cargo target category, location, and confidence level, enhancing the accuracy and reliability of cargo tipping detection.

[0034] 2. In the multi-scale cargo tipping detection network model of this invention, the Spatial Pyramid Perceptual Convolutional Module (SPPF-UR) enhances the model's ability to recognize a small number of samples by designing a multi-layered perceptual large-kernel convolutional module (UniRepLKNet). UniRepLKNet integrates subtle location and contextual information through a larger receptive field, improving sensitivity to tipping states and effectively identifying subtle abnormal features. Through better feature representation and multi-scale information fusion, the network can more accurately distinguish between tipped and normal samples, reducing false detection rates and improving overall detection accuracy. Due to differences in perspective, cargo images vary significantly in size; some features are obvious and large, while others are subtle and difficult to detect. For larger features, spatial pyramid pooling captures overall features through large-size pooling kernels, while for subtle features, small-size pooling kernels capture local details. Furthermore, the SPPF-UR module's design can process features at different scales while enhancing adaptability to various cargo tipping patterns, ensuring the model can effectively learn important information from different scales.

[0035] 3. In the multi-scale cargo dumping detection network model of the present invention, the iRMB module, through its residual structure design, simultaneously captures multi-scale and multi-level features of the input data, adapting to the shape and features of different types of images. This enables the model to extract features more flexibly from different angles when processing complex cargo dumping, thereby enhancing the recognition ability of cargo dumping samples and improving the localization and detection performance of different types of cargo areas. Furthermore, it enhances the model's learning ability on dumping samples, enabling the model to still obtain effective information when facing a minority of dumping samples, avoiding overtraining on most normal samples.

[0036] 4. In the multi-scale cargo tipping detection network model of this invention, a Focal-SIoU loss function is designed in the detection head as the localization regression loss. The introduction of SIoU ensures high accuracy of the detection box regression. By considering the overlapping area between the predicted box and the real box and their relative position, it does not rely solely on the area overlap. The application of FocalLoss can specifically adjust the weight ratio of positive and negative samples in the overall loss, thereby effectively alleviating the dominance of a large number of normal samples on the learning direction of model parameters, enhancing the learning strength of a few samples, thereby reducing overfitting and underfitting problems caused by sample imbalance, effectively correcting the position and size of the predicted box, and ensuring the reliability of cargo tipping detection under imbalanced samples. Attached Figure Description

[0037] To make the objectives, technical solutions, and advantages of the invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:

[0038] Figure 1This is a schematic diagram of the architecture of the multi-scale cargo tipping detection network model used in the method of the present invention;

[0039] Figure 2 This is a schematic diagram of the SPPF-UR module structure;

[0040] Figure 3 This is a schematic diagram of the UniRepLKNet structure in the SPPF-UR module;

[0041] Figure 4 SmaK Block and LarK Block structure diagram

[0042] Figure 5 This is a schematic diagram of the structure of the Dilated Reparam Block in UniRepLKNet;

[0043] Figure 6 This is a schematic diagram of the iRMB module structure;

[0044] Figure 7 This is a comparison chart of the detection accuracy of different types of samples at a data ratio of 2.29:1 in the example.

[0045] Figure 8 This is a comparison chart of the detection performance of different models in the embodiments. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but only to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0047] The present invention will now be described in further detail with reference to the accompanying drawings.

[0048] This invention discloses a cargo tipping detection method based on a large convolutional inverse residual network. After acquiring an image of the cargo to be detected, the method inputs it into a pre-trained multi-scale cargo tipping detection network model to obtain the cargo tipping state detection result. The multi-scale cargo tipping detection network model used in this invention employs the YOLOv8 network as its basic model architecture and is further optimized and improved as follows: First, an inverse residual attention module (iRMB) is designed in the backbone layer to enhance the feature extraction and localization capability of the cargo tipping area. Simultaneously, a spatial pyramid perceptual convolutional module (SPPF-UR) combining UniRepLKNet and SPPF is designed, introducing large convolutional kernels to enhance the model's attention to global feature information, thereby better capturing multi-scale cargo tipping feature information and improving the accuracy of cargo tipping recognition. Furthermore, a Focal-SIoU loss function is designed in the detection head to dynamically adjust the weight allocation of sample categories, thereby improving the model's ability to detect tipping samples from minority classes.

[0049] Specifically, such as Figure 1 As shown, the multi-scale cargo tipping detection network model used in the method of this invention includes a backbone network layer, a neck network layer, and a detection head. The backbone network layer is used to extract multi-scale features from the cargo image input to the multi-scale cargo tipping detection network model to obtain a multi-scale cargo tipping feature map. The neck network layer is used to perform feature fusion processing on the multi-scale cargo tipping feature map to obtain a fused feature map. The detection head is used to obtain the cargo tipping state detection result of the cargo image to be detected based on the fused feature map, which is used as the output of the multi-scale cargo tipping detection network model.

[0050] The following section provides a more detailed description of the cargo tipping detection method based on a large convolutional inverse residual network and the multi-scale cargo tipping detection network model used in this invention.

[0051] 1. Multi-scale cargo tipping detection network model

[0052] The main improvements of the multi-scale cargo dumping detection network model used in this invention are reflected in the SPPF-UR module, the iRMB module, and the Focal-SIoU function. For example... Figure 1 As shown, in this multi-scale cargo dumping detection network model:

[0053] The backbone network layer consists of sequentially connected convolutional layers, four convolutional feature extraction units, one inverted residual attention module, and one spatial pyramid perceptual convolutional module; each convolutional feature extraction unit includes cascaded convolutional layers and a C2f module;

[0054] The neck network layer includes two feature fusion upsampling units and two convolutional connected feature fusion units connected in sequence. The feature fusion upsampling unit includes an upsampling module, a connection layer, and a C2f module connected in sequence; the convolutional connected feature fusion unit includes a convolutional layer, a connection layer, and a C2f module connected in sequence.

[0055] The detection head is used to obtain the cargo tipping state detection result of the cargo image based on the output of the neck network layer, and serves as the output of the multi-scale cargo tipping detection network model;

[0056] In this system, the input of the multi-scale cargo dumping detection network model serves as the input of the backbone network layer. Within the backbone network layer, the outputs of the second and third convolutional feature extraction units also serve as the inputs to the connection layers of the second and first feature fusion upsampling units in the neck network layer, respectively. In addition to serving as the input of the spatial pyramid perception convolutional module in the backbone network layer to the neck network layer, the output of the spatial pyramid perception convolutional module also serves as the input to the connection layer of the second convolutional connection feature fusion unit in the neck network layer. Furthermore, within the neck network layer, the output of the first feature fusion upsampling unit also serves as the input to the connection layer of the first convolutional connection feature fusion unit, and the outputs of the second feature fusion upsampling unit and the two convolutional connection feature fusion units all serve as the inputs to the detection head.

[0057] This multi-scale cargo tipping detection network model, in its processing, goes through an initial convolutional layer and a C2f module in the backbone layer, followed by an inverse residual attention (iRMB) module to enhance feature extraction of the cargo tipping area. Then, a spatial pyramid perceptual convolutional (SPPF-UR) module efficiently captures global information to obtain multi-scale cargo tipping data. The feature maps enter the neck layer, where upsampling and feature concatenation effectively integrate multi-scale features. Finally, in the head layer, a Focal-SIoU loss function enhanced by Focal Loss focuses on the accuracy of the tipping location, thereby improving overall detection performance under imbalanced samples. The final output includes the detected cargo target category, location, and confidence level, enhancing the accuracy and reliability of cargo tipping detection.

[0058] Next, we will introduce in detail several key optimized and improved module units in the multi-scale cargo dumping detection network model.

[0059] 1.1 Spatial Pyramid Perceptual Convolutional Module (SPPF-UR Module)

[0060] To capture multi-scale cargo dumping information, SPPF was combined with the large kernel convolutional network UniRepLKNet to construct a spatial pyramid perception convolutional module, which can be abbreviated as SPPF-UR module.

[0061] Specifically, in the detection of unbalanced cargo tipping samples, the number of normal samples is usually small and their features are limited. To improve detection accuracy, by combining local and global features, the network can capture detailed features while preserving macroscopic image context information. Therefore, the SPPF module is combined with UniRepLKNet to construct the SPPF-UR module, further enhancing the network's ability to process input image information.

[0062] The architecture of the Spatial Pyramid Perception Convolutional Module SPPF-UR is as follows: Figure 2 As shown, the system includes a cascaded convolutional layer, three 2D max-pooling modules, a connection layer, and a UniRepLKNet perceptual kernel convolutional module. The outputs of the convolutional layer and the three 2D max-pooling modules, in addition to being passed through the cascaded layers, are also used as inputs to the connection layer. After being concatenated by the connection layer, the outputs are sent to the UniRepLKNet perceptual kernel convolutional module. The output of the UniRepLKNet perceptual kernel convolutional module is used as the output of the spatial pyramid perceptual convolutional module.

[0063] In the specific processing, the input feature map of the Spatial Pyramid Perception Convolutional Module (SPPF-UR) is processed by feature extraction through the convolutional layer, and then passed through three two-dimensional max pooling modules for max pooling before being passed to the connection layer. The outputs of the convolutional layer and the three two-dimensional max pooling modules are also used as inputs to the connection layer. After being concatenated by the connection layer, the output is sent to the Perception Large Kernel Convolutional Module (UniRepLKNet) for further global feature extraction, resulting in the output of the Spatial Pyramid Perception Convolutional Module (SPPF-UR).

[0064] The UniRepLKNet perceptual large-kernel convolutional module expands the receptive field by using large-sized convolutional kernels, capturing more global feature information while maintaining network width and depth, thus enhancing its ability to capture complex features. The overall structure of the UniRepLKNet perceptual large-kernel convolutional module is as follows: Figure 3As shown, it includes four feature extraction stages: feature extraction stage 1 (Stage 1), feature extraction stage 2 (Stage 2), feature extraction stage 3 (Stage 3), and feature extraction stage 4 (Stage 4). Adjacent feature extraction stages are connected by a transition segment, which is a 3×3 convolutional layer with a stride of 2. Each feature extraction stage includes several cascaded kernel units of different sizes, and each kernel unit includes one cascaded large kernel block (LarK Block) and two small kernel blocks (SmaK Block).

[0065] In the specific processing, the input feature map of the perceptual large kernel convolution module UniRepLKNet is first downsampled and expanded by a 3×3 convolutional layer with a stride of 4, and then output to the feature extraction stage 1. Then, it is processed by several 4-feature extraction stages and 3×3 convolutional layers of length 2 between adjacent feature extraction stages. The feature map obtained by the feature extraction stage 4 is used as the output of the perceptual large kernel convolution module UniRepLKNet.

[0066] The structures of the large kernel block (LarK Block) and the small kernel block (SmaK Block) are as follows: Figure 4 As shown.

[0067] The structure of the SmaK Block is as follows: Figure 4 As shown in (a), the kernel includes a cascaded 3×3 depthwise separable convolutional layer (DWconv), a batch normalization layer (BN), a compression activation layer (SE block), a feedforward network layer (FFN), and another batch normalization layer (BN). The input feature map of the small kernel block (SmaK Block) is processed sequentially through the 3×3 depthwise separable convolutional layer (DW conv), a batch normalization layer (BN), a compression activation layer (SE block), a feedforward network layer (FFN), and another batch normalization layer (BN) to obtain the feature map. This feature map is then added to the input feature map of the small kernel block (SmaK Block) to obtain the output feature map of the small kernel block (SmaKBlock).

[0068] The structure of a large kernel block is as follows: Figure 4As shown in (b), the large kernel block (LarK Block) consists of a cascaded dilated reparameter block, a batch normalization (BN) layer, a squeezed excitation (SE) block, a feedforward network (FFN) layer, and another batch normalization (BN) layer. The input feature map of the large kernel block is processed sequentially through the dilated reparameter block, a batch normalization (BN) layer, a squeezed excitation (SE) layer, a feedforward network (FFN) layer, and another batch normalization (BN) layer to obtain the feature map. This feature map is then added to the input feature map of the large kernel block (LarK Block) to obtain the output feature map of the large kernel block (LarK Block).

[0069] The main architectures of the SmaK Block and the LarK Block are basically the same, both consisting of depthwise convolutional layers, a batch normalization (BN) layer, a squeeze activation layer (SE block), a feedforward network layer (FFN), and another batch normalization (BN) layer. The difference lies in that the SmaK Block uses 3×3 depthwise separable convolutions (DWconv) for its depthwise convolutional layers, while the LarK Block uses dilated reparameter blocks. The architectures of the depthwise separable convolutional layer (DWconv), the batch normalization layer (BN), the squeeze activation layer (SE block), and the feedforward network layer (FFN) are commonly used modules in existing technologies and will not be elaborated further. The dilated reparameter block will be introduced below.

[0070] The Dilated Reparam Block (DRB) module, in addition to using large kernel convolutions, also employs parallel dilated convolutions and transforms the entire block into an equivalent large kernel convolution through structural reparameterization, thereby enhancing the model's ability to capture features at different scales. Specifically, such as... Figure 5 As shown, the Dilated Reparameterization Layer (DRB) comprises five parallel dilated convolutional modules: one consisting of a cascaded 9×9 dilated convolutional block and a batch normalization module; one consisting of a cascaded 5×5 dilated convolutional block and a batch normalization module; and three consisting of a cascaded 6×6 dilated convolutional block and a batch normalization module. During processing, the input feature map of the DRB is processed by each of the five dilated convolutional modules to obtain dilated convolutional maps, which are then transformed into a large-kernel convolutional feature map through structural reparameterization, serving as the output of the DRB.

[0071] After using dilated convolution, the dilated reparameterized layer (DRB) can not only identify local details of cargo dumping but also detect global features. In imbalanced datasets, minority class dumping samples are usually few in number and unevenly distributed. The DRB module can more effectively learn the features of these sparse samples. Through the reparameterized training mechanism, the DRB module can better utilize minority samples for learning, thereby improving the generalization ability of minority class dumping samples. Reparameterization ensures that the model maintains high efficiency without losing the fine-grained information captured during training. This mechanism significantly improves the model's detection accuracy and reduces the false negative rate.

[0072] Overall, the Spatial Pyramid Perceptual Convolutional Module (SPPF-UR) enhances the model's ability to recognize a small number of samples by designing the multi-layered perceptual large-kernel convolutional module UniRepLKNet. UniRepLKNet integrates subtle location and contextual information through a larger receptive field, improving sensitivity to tipping states and effectively identifying subtle anomalies. Through better feature representation and multi-scale information fusion, the network can more accurately distinguish between tipped and normal samples, reducing false detection rates and improving overall detection accuracy. Due to differences in perspective, cargo images vary significantly in size; some features are obvious and large, while others are subtle and difficult to detect. For larger features, spatial pyramid pooling captures the overall features using large-size pooling kernels, while for subtle features, small-size pooling kernels capture local details. Furthermore, the SPPF-UR module's design can process features at different scales while enhancing adaptability to various cargo tipping patterns, ensuring the model can effectively learn important information from different scales.

[0073] 1.2iRMB Module

[0074] To fully extract cargo location features from a small number of normal data samples, an inverted residual attention module (iRMB) was designed. This module combines residual structure and attention mechanism to enhance the model's ability to extract location features from complex cargo dumping. Specifically, as... Figure 6 As shown, the inverse residual attention module iRMB includes a moving block branch with an integrated attention mechanism, and an inverse residual branch consisting of a cascaded first 1×1 convolutional layer, a 3×3 depth-separable convolutional layer, and a second 1×1 convolutional layer.

[0075] In the specific processing, the input feature map of the inverse residual attention module iRMB is used as the input to the moving block branch and the inverse residual branch, respectively. The moving block branch is used to extract the attention matrix of the input feature map and output it to the inverse residual branch. In the inverse residual branch, the output of the first 1×1 convolutional layer is multiplied by the attention matrix output by the moving block branch to obtain the attention convolutional feature map, which is used as the input of the 3×3 depth separable convolutional layer. The output of the 3×3 depth separable convolutional layer is then added to the attention convolutional feature map and used as the input of the second 1×1 convolutional layer. The output of the second 1×1 convolutional layer is then added to the input feature map of the inverse residual attention module iRMB and used as the output of the inverse residual attention module iRMB.

[0076] Specifically, for the input feature map X of the iRMB module, in the moving block branch integrating the attention mechanism, the query vector Q and key vector K corresponding to the input feature map X are first processed to obtain the attention matrix AttnMat:

[0077] Q = XW Q K = XW K ;

[0078]

[0079] These are the linear transformation matrices corresponding to the query and key in the self-attention mechanism; d k The dimension of the attention head; Softmax(·) represents the normalized exponential function operation;

[0080] Then, the attention matrix AttnMat is passed to the inverse residual branch of the iRMB module, and after processing, the output feature map F of the iRMB module is obtained, which can be expressed as:

[0081] F=X+Conv(X′+DWconv(X′));

[0082]

[0083] Where Conv(·) and DWconv(·) represent convolution and depthwise separable convolution operations, respectively; T is the transpose symbol.

[0084] When normal samples are scarce, the iRMB module, through its residual structure design, simultaneously captures multi-scale and multi-level features of the input data, adapting to the morphology and features of different image types. This allows the model to extract features more flexibly from different angles when handling complex cargo dumping, thereby enhancing its ability to identify cargo dumping samples and improving the localization and detection performance of different types of cargo areas. The iRMB module introduces residual connections, effectively mitigating the gradient vanishing problem and enabling the model to maintain good learning capabilities in deep networks. Furthermore, the multi-branch structure reduces dependence on single features through information fusion from different paths, improving the model's robustness. The iRMB module introduces a channel attention mechanism, optimizing the channel weight allocation of feature maps, focusing on important feature regions, assigning higher weights to important channels, and suppressing redundant and irrelevant background features. The attention mechanism in the iRMB module is adaptive, dynamically adjusting the region of interest based on different input feature maps, allowing the network to adaptively adjust the attention direction according to the characteristics of the data, optimizing feature extraction in key regions. Through adaptive focusing, the model can more accurately identify complex cargo dumping, thereby improving detection accuracy.

[0085] The iRMB module enhances the model's feature extraction capability for normal states by combining an inverse residual structure and uses an attention mechanism to focus on key areas, thereby enhancing the model's learning ability for dumped samples. This allows the model to still acquire effective information when faced with a minority of dumped samples, avoiding overtraining on most normal samples and thus improving the accuracy of cargo dumping area location detection.

[0086] 2. Training of a multi-scale cargo tipping detection network model

[0087] The multi-scale cargo tipping detection network model is trained in the following way:

[0088] S101: Obtain the cargo image sample dataset, label both the cargo tipping image samples and the cargo normal image samples in the cargo image sample dataset, and divide them into training set and test set according to a preset ratio;

[0089] S102: Input the training set into the multi-scale cargo dumping detection network model for training, and optimize the parameters of the multi-scale cargo dumping detection network model with the goal of minimizing the loss function until the multi-scale cargo dumping detection network model converges, thus obtaining the trained multi-scale cargo dumping detection network model.

[0090] S103: Test the multi-scale cargo dumping detection network model using the test set to confirm the recognition performance of the trained multi-scale cargo dumping detection network model; if the recognition performance meets the requirements, end the training of the multi-scale cargo dumping detection network model; otherwise, return to step S102.

[0091] During the training of the multi-scale cargo dumping detection network model, the BCE (Binary Cross-Entropy) loss function was chosen for classification loss. This is also the classification loss function used in the YOLOv8 network, which helps to balance the classification convergence and regression speed and classification accuracy during model training. For the localization regression loss, to improve the detection accuracy of minority samples, this invention designs a Focal-SIoU loss function to assign higher weights to minority samples, helping the model focus on learning the features of these minority samples, thereby improving the accuracy of sample detection.

[0092] Specifically, the expression for the Focal-SIoU loss function is:

[0093] L Focal-SIoU =IoU γ L SIoU ;

[0094] Among them, L Focal-SIoU Indicates Focal-SIoU loss; IoU, L SIoU Let IoU loss and SIoU loss be represented respectively; γ is the adjustment factor; and we have:

[0095]

[0096] γ = 2 - Λ;

[0097] Among them, B, B gt Let Area(·) represent the predicted cargo dumping area location frame and the actual cargo dumping area location frame, respectively; Λ represents the angle cost between the predicted and actual cargo dumping area location frames; Δ represents the center distance cost between the predicted and actual cargo dumping area location frames; Ω represents the cost of the overlapping area between the predicted and actual cargo dumping area location frames; and we have:

[0098]

[0099] Where e is the natural constant; θ is a preset natural number parameter; and parameter ρ x ρ y ω w ω h They are as follows:

[0100]

[0101] Among them, c x (b,b gt ), c y (b,b gt) represent the center point b of the predicted cargo dumping area location frame and the center point b of the actual cargo dumping area location frame, respectively. gt The x-coordinate distance and y-coordinate distance between them; ρ(b,b) gt () represents the center point b of the predicted cargo dumping area positioning frame and the center point b of the actual cargo dumping area positioning frame. gt The Euclidean distance between them; w and h represent the width and height of the predicted cargo tipping area positioning frame, respectively; w gt h gt These represent the width and height of the positioning frame for the actual cargo dumping area, respectively; C w C h These represent the width and height of the smallest bounding rectangle covering the predicted cargo dumping area location frame and the actual cargo dumping area location frame, respectively.

[0102] In the Focal-SIoU loss function, the introduction of SIoU ensures high accuracy in bounding box regression by considering the overlap area between the predicted and ground truth boxes and their relative positions, rather than solely relying on area overlap. The application of FocalLoss allows for targeted adjustment of the weight ratio of positive and negative samples in the overall loss, effectively mitigating the dominance of a large number of normal samples on the model's parameter learning direction and enhancing the learning power for a minority of samples. Therefore, the Focal-SIoU loss function, combining FocalLoss and SIoU, enables the model to effectively learn from a minority of samples while reducing overfitting and underfitting problems caused by sample imbalance, effectively correcting the position and size of the predicted bounding boxes, and ensuring the reliability of cargo tipping detection under imbalanced samples.

[0103] 3 Examples

[0104] In this embodiment, a dataset is used to test and verify the method of the present invention and other methods to better demonstrate the technical advantages of the method of the present invention. This will be described in detail below.

[0105] 3.1 Dataset and Experimental Parameters

[0106] This experiment used cargo images collected at the company's unloading site as the dataset, covering different angles and lighting conditions. The Labelimg library was used to label the dataset with bounding boxes, containing two categories: normal and abnormal. Data allocation and proportions are shown in Table 1.

[0107] Table 1 Data Allocation Table

[0108]

[0109] 3.2 Comparison of Experimental Results

[0110] 3.2.1 Ablation Experiment

[0111] In this embodiment, for ease of description, the multi-scale cargo dumping detection network model used in this invention is labeled as the LCIRNet model.

[0112] To verify the effectiveness of the LCIRNet model proposed in this invention in the cargo state detection task under imbalanced samples, an ablation experiment was designed based on the improved method. The SPPF-UR module, iRMB module and Focal-SIoU function were designed on the basis of YOLOv8. The modules were accumulated in sequence, and the experimental results under the first set of data ratios are shown in Table 2.

[0113] Table 2 shows the ablation experiment results at a data ratio of 2.29:1.

[0114]

[0115] As shown in Table 2, under the same experimental conditions, the proposed LCIRNet model has improved precision, recall and mAP to varying degrees. The precision, recall and mAP of the proposed LCIRNet model are 98%, 99.3% and 99.4%, respectively. In Experiment 1, the SPPF-UR module was designed, which increased the recall rate for identifying tipped goods by 2.6%, but decreased the precision. In Experiment 2, the iRMB module was designed, which increased the accuracy for identifying tipped goods by 0.8%, the precision rate by 0.3%, and the recall rate remained the same. In Experiment 3, the iRMB module was designed based on Experiment 1, and the model's precision, recall rate, and detection precision increased compared to Experiment 1, but the model's precision did not change significantly. The last group of experiments was based on the proposed LCIRNet model, and Focal-SIoU was designed based on Experiment 3. Compared to the YOLOv8 model, the proposed LCIRNet model improved precision, recall rate, and mAP@0.5 by 0.8%, 1.9%, and 0.3%, respectively, further demonstrating the effectiveness of the method of this invention.

[0116] Experiment 1, by designing the SPPF-UR module, captured cargo dumping features at different scales, improving the model's recall and detection accuracy for minority samples. However, the introduction of global features may increase the response of background regions, leading to a slight decrease in precision. Experiment 2 shows that the model, through residual structure and attention mechanism, becomes more sensitive to key regions, reducing false negatives and improving the model's precision, recall, and detection accuracy. Experiment 3 demonstrates that the SPPF-UR module provides global contextual information, enhancing the perception of minority samples, while the iRMB module further optimizes the model's detection performance for minority samples through reinforcement learning of local key regions. The combination of these two modules achieves a balance between global and local features to a certain extent, thereby improving the model's adaptability to minority samples. Therefore, the proposed LCIRNet model significantly improves global information fusion and local feature extraction capabilities through SPPF-UR, iRMB, and Focal-SIoU modules, adjusts loss weights, and enhances the model's learning effect on minority samples, showing significant improvements in precision, recall, and detection accuracy, meeting the needs of practical engineering applications.

[0117] 3.2.2 Comparative Experiment

[0118] To verify the cargo tipping detection performance of the LCIRNet model proposed in this invention under unbalanced sample conditions, it was compared with four different models. The experimental results are shown in Table 3.

[0119] Table 3. Data Comparison and Experimental Results

[0120]

[0121] Table 3 shows that the proposed LCIRNet model achieves the best detection performance under imbalanced sample conditions. In the first dataset, the proposed LCIRNet model achieves a detection accuracy of 99.4% for cargo states. The accuracy of the proposed LCIRNet model is 0.3% higher than that of YOLOv8. The proposed LCIRNet model captures global contextual information through the SPPF-UR module, which helps the model capture multi-scale cargo tipping state information. The iRMB module utilizes residual structure and attention mechanisms to highlight features in a few normal sample regions, improving the model's detection capability for a few samples. Focal-SIoU assigns higher weights to a few samples, thereby improving the model's detection accuracy, further demonstrating the effectiveness of the proposed method.

[0122] To further verify the detection performance of the proposed LCIRNet model, based on the data ratio of 2.29:1 in Table 3 above, a comparative experiment was conducted between the proposed LCIRNet model and some existing object detection algorithms, including YOLOv3, YOLOv5, YOLOv8, and YOLO11. The results of the comparative experiment are shown in Table 4.

[0123] Table 4 shows the experimental results of the model algorithm comparison at a data ratio of 2.29:1.

[0124]

[0125] As shown in Table 4, the LCIRNet model method proposed in this invention improves precision, recall, and mAP to varying degrees. The precision, recall, and mAP of the proposed LCIRNet model are 98%, 99.3%, and 99.4%, respectively. Regarding mAP, it increases by 0.8%, 0.7%, and 0.3% compared to YOLOv3, YOLOv5, and YOLOv8, respectively, while the mAP of YOLO11 is on par with the method of this invention. In terms of recall, the method of this invention improves by 1.4%, 1.9%, and 1.9% compared to the corresponding models, respectively, and is on par with YOLO11. Regarding precision, the method of this invention improves by 0.8%, 1.8%, and 0.2% compared to YOLOv3, YOLOv8, and YOLOv11, respectively.

[0126] To more intuitively demonstrate the detection accuracy of existing models for different categories of samples, a comparative experiment was conducted using different models. The results are shown in Table 5. Figure 7 As shown.

[0127] Table 5 compares the detection accuracy of different categories of samples at a data ratio of 2.96:1.

[0128]

[0129] From Table 5 and Figure 7 It is evident that the method of this invention achieves the best detection performance for a small number of samples. Specifically, the method of this invention achieves a detection accuracy of 99.3% for samples of goods in a normal state and 99.5% for samples of goods that have been tipped over. The detection accuracy for normal samples is significantly higher than that of YOLOv3, YOLOv5, and YOLOv8. Compared to YOLOv11, the detection accuracy of the method of this invention is comparable for both tipped and normal samples. The LCIRNet model proposed in this invention expands the receptive field range through the SPPF-UR module using large convolutional kernels, thereby capturing global features more comprehensively. The iRMB module highlights the key features of a small number of normal samples. Finally, the Focal-SIoU module dynamically allocates weights, increasing the attention given to normal samples and improving detection accuracy, further demonstrating the effectiveness of the method of this invention.

[0130] 3.2.3 Visualization Verification Experiment

[0131] To verify the effectiveness of the proposed LCIRNet model in real-world object detection scenarios, we selected some images from a self-made dumping dataset from a certain company for detection. The detection results were compared to... Figure 8 As shown. Figure 8 (a), (b), (c), and (b) in the figure represent the monitoring results of four different cargo images.

[0132] Depend on Figure 8 It can be seen that the proposed LCIRNet model achieves the best detection performance compared to YOLOv3, YOLOv5, YOLOv8, and YOLO11. Figure 8 (a) and Figure 8 (b) It can be seen that the detection accuracy of YOLOv3, YOLOv5, and YOLOv8 for the tipping state of goods is lower than that of the proposed LCIRNet model; Figure 8 (c) It can be seen that the proposed LCIRNet model can not only identify cargo overturned samples, but also improve the detection accuracy of normal cargo samples. The detection accuracy of the comparison model is lower than that of the proposed LCIRNet model and there are false negatives. The method of this invention uses SPPF-UR, iRMB and Focal-SIoU modules to enable the proposed LCIRNet model to effectively capture global and local features by expanding the receptive field and integrating multi-scale information. By combining residual structure and attention mechanism, the feature representation of minority sample regions is strengthened, while the interference of normal samples on detection is suppressed, and the model’s attention to minority samples is improved. By dynamically adjusting the weight distribution of easy and difficult samples, the model’s optimization ability for minority samples is enhanced, and the impact of class imbalance on the loss function is mitigated. Therefore, the LCIRNet model proposed in this invention can not only detect cargo overturned state more accurately under imbalanced samples and reduce the false negative rate of overturning, but also improve the detection accuracy of normal samples, meeting the intelligent detection needs of cargo state under imbalanced samples in actual construction sites.

[0133] 4 Overview

[0134] To address the issue of low detection accuracy for minority samples due to the imbalance between tilted and normal cargo samples in intelligent cargo tilt detection, this invention proposes a cargo tilt detection method based on the LCIRNet model under imbalanced sampling conditions. Experimental results show that the proposed method achieves precision, recall, and mAP of 98%, 99.3%, and 99.4%, respectively, for cargo tilt detection. Compared with the original YOLOv8 model, precision, recall, and mAP are improved by 0.8%, 1.9%, and 0.3%, respectively. Compared with YOLOv3 and YOLOv5, the proposed LCIRNet model significantly reduces the false negative and false positive rates while improving detection accuracy, meeting the requirements for practical engineering applications. Furthermore, the applicability of the proposed LCIRNet model for minority sample detection is demonstrated through test graph visualization results, providing a feasible approach for minority sample detection under imbalanced sampling conditions.

[0135] In summary, compared with the prior art, the present invention has the following technical advantages:

[0136] 1. This invention presents a cargo tipping detection method based on a large convolutional inverse residual network. Utilizing an optimized multi-scale cargo tipping detection network model, the backbone layer (Backbone) passes through an initial convolutional layer and a C2f module, followed by an inverse residual attention (iRMB) module to enhance feature extraction of the cargo tipping area. Then, a spatial pyramid perceptual convolutional (SPPF-UR) module efficiently captures global information to obtain multi-scale cargo tipping information. The feature map enters the neck layer (Neck), where upsampling and feature concatenation effectively integrate multi-scale features. Finally, in the detection head (Head), a Focal Loss combined with a SIoU-enhanced Focal-SIoU loss function focuses on the accuracy of the tipping location, thereby improving the overall detection performance under imbalanced samples. The final output includes the detected cargo target category, location, and confidence level, enhancing the accuracy and reliability of cargo tipping detection.

[0137] 2. In the multi-scale cargo tipping detection network model of this invention, the Spatial Pyramid Perceptual Convolutional Module (SPPF-UR) enhances the model's ability to recognize a small number of samples by designing a multi-layered perceptual large-kernel convolutional module (UniRepLKNet). UniRepLKNet integrates subtle location and contextual information through a larger receptive field, improving sensitivity to tipping states and effectively identifying subtle abnormal features. Through better feature representation and multi-scale information fusion, the network can more accurately distinguish between tipped and normal samples, reducing false detection rates and improving overall detection accuracy. Due to differences in perspective, cargo images vary significantly in size; some features are obvious and large, while others are subtle and difficult to detect. For larger features, spatial pyramid pooling captures overall features through large-size pooling kernels, while for subtle features, small-size pooling kernels capture local details. Furthermore, the SPPF-UR module's design can process features at different scales while enhancing adaptability to various cargo tipping patterns, ensuring the model can effectively learn important information from different scales.

[0138] 3. In the multi-scale cargo dumping detection network model of the present invention, the iRMB module, through its residual structure design, simultaneously captures multi-scale and multi-level features of the input data, adapting to the shape and features of different types of images. This enables the model to extract features more flexibly from different angles when processing complex cargo dumping, thereby enhancing the recognition ability of cargo dumping samples and improving the localization and detection performance of different types of cargo areas. Furthermore, it enhances the model's learning ability on dumping samples, enabling the model to still obtain effective information when facing a minority of dumping samples, avoiding overtraining on most normal samples.

[0139] 4. In the multi-scale cargo tipping detection network model of this invention, a Focal-SIoU loss function is designed in the detection head as the localization regression loss. The introduction of SIoU ensures high accuracy of the detection box regression. By considering the overlapping area between the predicted box and the real box and their relative position, it does not rely solely on the area overlap. The application of FocalLoss can specifically adjust the weight ratio of positive and negative samples in the overall loss, thereby effectively alleviating the dominance of a large number of normal samples on the learning direction of model parameters, enhancing the learning strength of a few samples, thereby reducing overfitting and underfitting problems caused by sample imbalance, effectively correcting the position and size of the predicted box, and ensuring the reliability of cargo tipping detection under imbalanced samples.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described with reference to preferred embodiments, those skilled in the art should understand that various changes in form and detail can be made without departing from the spirit and scope of the invention as defined in the appended claims.

Claims

1. A method for detecting cargo tipping based on a large convolutional inverse residual network, characterized in that, The image of the cargo to be detected is obtained and input into a pre-trained multi-scale cargo tipping detection network model to obtain the cargo tipping state detection result of the cargo image to be detected. The multi-scale cargo tipping detection network model includes a backbone network layer, a neck network layer, and a detection head. The backbone network layer is used to extract multi-scale features from the cargo image input to the multi-scale cargo tipping detection network model to obtain a multi-scale cargo tipping feature map. The neck network layer is used to perform feature fusion processing on the multi-scale cargo tipping feature map to obtain a fused feature map. The detection head is used to obtain the cargo tipping state detection result of the cargo image to be detected based on the fused feature map, which is used as the output of the multi-scale cargo tipping detection network model.

2. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 1, characterized in that, In the multi-scale cargo dumping detection network model: The backbone network layer includes a series of convolutional layers, four convolutional feature extraction units, one inverted residual attention module, and one spatial pyramid perceptual convolutional module; each convolutional feature extraction unit includes a cascaded convolutional layer and a C2f module. The neck network layer includes two feature fusion upsampling units and two convolutional connection feature fusion units connected in sequence. The feature fusion upsampling unit includes an upsampling module, a connection layer, and a C2f module connected in sequence; the convolutional connection feature fusion unit includes a convolutional layer, a connection layer, and a C2f module connected in sequence. The detection head is used to obtain the cargo tipping state detection result of the cargo image based on the output of the neck network layer, and as the output of the multi-scale cargo tipping detection network model; In this system, the input of the multi-scale cargo dumping detection network model serves as the input of the backbone network layer. Within the backbone network layer, the outputs of the second and third convolutional feature extraction units also serve as the inputs to the connection layers of the second and first feature fusion upsampling units in the neck network layer, respectively. In addition to serving as the input of the spatial pyramid perception convolutional module in the backbone network layer to the neck network layer, the output of the spatial pyramid perception convolutional module also serves as the input to the connection layer of the second convolutional connection feature fusion unit in the neck network layer. Furthermore, within the neck network layer, the output of the first feature fusion upsampling unit also serves as the input to the connection layer of the first convolutional connection feature fusion unit, and the outputs of the second feature fusion upsampling unit and the two convolutional connection feature fusion units all serve as the inputs to the detection head.

3. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 2, characterized in that, The inverse residual attention module includes a moving block branch with an integrated attention mechanism, and an inverse residual branch consisting of a cascaded first 1×1 convolutional layer, a 3×3 depth-separable convolutional layer, and a second 1×1 convolutional layer. The input feature map of the inverse residual attention module is used as the input to the moving block branch and the inverse residual branch, respectively; the moving block branch is used to extract the attention matrix of the input feature map and output it to the inverse residual branch; In the inverse residual branch, the output of the first 1×1 convolutional layer is multiplied by the attention matrix output of the moving block branch to obtain an attention convolutional feature map, which is used as the input of the 3×3 depth separable convolutional layer. The output of the 3×3 depth separable convolutional layer is then added to the attention convolutional feature map and used as the input of the second 1×1 convolutional layer. The output of the second 1×1 convolutional layer is then added to the input feature map of the inverse residual attention module and used as the output of the inverse residual attention module.

4. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 2, characterized in that, The spatial pyramid perceptual convolutional module includes a cascaded convolutional layer, three two-dimensional max pooling modules, a connection layer, and a perceptual large kernel convolutional module. In addition to being cascaded and passed, the outputs of the convolutional layer and the three two-dimensional max pooling modules are also used as inputs to the connection layer. After being concatenated by the connection layer, they are output to the perceptual large kernel convolutional module. The output of the perceptual large kernel convolutional module is used as the output of the spatial pyramid perceptual convolutional module. The input feature map of the spatial pyramid perceptual convolutional module is processed by feature extraction through the convolutional layer, and then passed through three two-dimensional max pooling modules for max pooling before being passed to the connection layer. The outputs of the convolutional layer and the three two-dimensional max pooling modules are also used as inputs to the connection layer. After being concatenated by the connection layer, the output is sent to the perceptual large kernel convolutional module for further global feature extraction, resulting in the output of the spatial pyramid perceptual convolutional module.

5. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 4, characterized in that, The perceptual large kernel convolution module includes four feature extraction stages: feature extraction stage 1, feature extraction stage 2, feature extraction stage 3, and feature extraction stage 4. Adjacent feature extraction stages are connected by a transition segment, which is a 3×3 convolutional layer with a stride of 2. Each feature extraction stage includes several cascaded large and small kernel units, and each large and small kernel unit includes one cascaded large kernel block and two small kernel blocks. The input feature map of the perceptual large kernel convolution module is first downsampled and expanded by a 3×3 convolutional layer with a stride of 4, and then output to the feature extraction stage 1. After being processed by several 4-feature extraction stages and 3×3 convolutional layers of length 2 between adjacent feature extraction stages, the feature map obtained by the feature extraction stage 4 is used as the output of the perceptual large kernel convolution module.

6. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 5, characterized in that, The SmaK Block comprises a cascaded 3×3 depthwise separable convolutional layer, a batch normalization layer, a compressed activation layer, a feedforward network layer, and another batch normalization layer. The input feature map of the SmaK Block is processed sequentially through the 3×3 depthwise separable convolutional layer, a batch normalization layer, a compressed activation layer, a feedforward network layer, and another batch normalization layer to obtain the feature map. This result is then added to the input feature map of the SmaK Block to obtain the output feature map of the SmaK Block. The large kernel block (LarK Block) comprises a cascaded extended reparameterization layer, a batch normalization layer, a compressed excitation layer, a feedforward network layer, and another batch normalization layer. The input feature map of the large kernel block is processed sequentially through the extended reparameterization layer, a batch normalization layer, a compressed excitation layer, a feedforward network layer, and another batch normalization layer to obtain the feature map. This result is then added to the input feature map of the large kernel block (LarK Block) to obtain the output feature map of the large kernel block (LarK Block).

7. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 6, characterized in that, The dilated reparameterization layer includes five dilated convolution modules in parallel, including one dilated convolution module consisting of a cascaded 9×9 dilated convolution block and a batch normalization module, one dilated convolution module consisting of a cascaded 5×5 dilated convolution block and a batch normalization module, and three dilated convolution modules consisting of a cascaded 6×6 dilated convolution block and a batch normalization module. The input feature map of the dilated reparameterization layer is processed by five dilated convolution modules to obtain dilated convolution maps, which are then converted into large kernel convolution feature maps by structural reparameterization and used as the output of the dilated reparameterization layer.

8. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 1, characterized in that, The multi-scale cargo dumping detection network model is trained in the following manner: S101: Obtain the cargo image sample dataset, label both the cargo tipping image samples and the cargo normal image samples in the cargo image sample dataset, and divide them into training set and test set according to a preset ratio; S102: Input the training set into the multi-scale cargo dumping detection network model for training, and optimize the parameters of the multi-scale cargo dumping detection network model with the goal of minimizing the loss function until the multi-scale cargo dumping detection network model converges, thus obtaining the trained multi-scale cargo dumping detection network model. S103: Test the multi-scale cargo dumping detection network model using the test set to confirm the recognition performance of the trained multi-scale cargo dumping detection network model; if the recognition performance meets the requirements, then end the training of the multi-scale cargo dumping detection network model. Otherwise, return to step S102.

9. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 8, characterized in that, Among the loss functions, the Focal-SIoU loss function is selected for localization regression loss, and the BCE loss function is selected for classification loss.

10. The cargo tipping detection method based on a large convolutional inverse residual network according to claim 9, characterized in that, The expression for the Focal-SIoU loss function is: L Focal-SIoU =IoU γ L SIoU ; Among them, L Focal-SIoU Indicates Focal-SIoU loss; IoU, L SIoU Let IoU loss and SIoU loss be represented respectively; γ is the adjustment factor; and we have: γ = 2 - Λ; Among them, B, B gt Let Area(·) represent the predicted cargo dumping area location frame and the actual cargo dumping area location frame, respectively; Λ represents the angle cost between the predicted and actual cargo dumping area location frames; Δ represents the center distance cost between the predicted and actual cargo dumping area location frames; Ω represents the cost of the overlapping area between the predicted and actual cargo dumping area location frames; and we have: Where e is the natural constant; θ is a preset natural number parameter; and parameter ρ x ρ y ω w ω h They are as follows: Among them, c x (b,b gt ), c y (b,b gt ) represent the center point b of the predicted cargo dumping area location frame and the center point b of the actual cargo dumping area location frame, respectively. gt The x-coordinate distance and y-coordinate distance between them; ρ(b,b) gt () represents the center point b of the predicted cargo dumping area positioning frame and the center point b of the actual cargo dumping area positioning frame. gt The Euclidean distance between them; w and h represent the width and height of the predicted cargo tipping area positioning frame, respectively; w gt h gt These represent the width and height of the positioning frame for the actual cargo dumping area, respectively; C w C h These represent the width and height of the smallest bounding rectangle covering the predicted cargo dumping area location frame and the actual cargo dumping area location frame, respectively.

Citation Information

Cited By

  • Urban road ponding rapid monitoring method and system based on structure re-parameterization

    CN121937880A