A manhole cover hidden danger detection method based on object detection

By improving the dynamic upsampling and inverting residual attention downsampling module of the YOLOv9 model, the problem of low detection accuracy caused by the complex manhole cover environment is solved, and efficient, accurate and real-time detection of manhole cover hidden dangers is achieved.

CN119992073BActive Publication Date: 2025-07-04JIANGNAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510450797.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-04
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

In the prior art, the manhole cover environment background is complex and the target is small, resulting in low model detection accuracy and it is difficult to meet the needs of efficient and accurate manhole cover hidden danger detection.

Method used

The dynamic upsampling module and the inverted residual attention downsampling module are used to improve the YOLOv9 model, and the sampling offset is generated through the dynamic upsampling module to adjust the sampling network. Combined with the inverted residual attention downsampling module, the attention mechanism and depth separation convolution are introduced, which enhances feature extraction capabilities and improves the detection accuracy of small objects and details in manhole cover images.

Benefits of technology

It improves the accuracy and robustness of manhole cover hidden danger detection, adapts to complex environments, realizes efficient detection of manhole cover hidden dangers, meets real-time monitoring needs, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992073B_ABST
    Figure CN119992073B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical fields of deep learning and object detection, and discloses a manhole cover hidden danger detection method based on object detection. The manhole cover image is input into the backbone network of the object detection model. Through multiple sequentially connected extraction units, multiple primary feature maps of different scales are extracted and input into the neck network based on the FPN structure. After the fusion feature maps corresponding to each scale are output, they are input into the head network, and the predicted labels, predicted bounding boxes and confidence levels corresponding to each scale of the manhole cover image are output. The dynamic upsampling module generates sampling offsets to adjust the sampling network, better capture small target information, and improve the detection and segmentation accuracy of details in the manhole cover image; the inverted residual attention downsampling module introduces an attention mechanism, depthwise separable convolution and a residual mechanism, effectively compresses data, reduces information loss, enhances the attention to key features of the manhole cover image, improves the feature expression of the manhole cover hidden danger by the model, and further improves the detection accuracy of the manhole cover hidden danger.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and object detection, and particularly to a manhole cover hidden danger detection method based on object detection. Background Art

[0002] With the acceleration of the urbanization process, the safety and reliability of urban infrastructure have been increasingly emphasized. As an important part of urban roads and pipeline systems, manhole covers play an important role in ensuring the safe passage of pedestrians and vehicles in urban infrastructure management. The problem of hidden dangers directly relates to public safety and urban image. However, for a long time, due to various factors, including natural wear, vehicle pressure, and improper use, manhole covers may have various potential safety hazards, including manhole cover breakage, uncovered manhole covers, lost manhole covers, and manhole ring problems.

[0003] Traditional methods for detecting manhole cover hidden dangers mainly rely on manual inspections, which are inefficient and costly, and difficult to meet the detection requirements in large-scale and complex environments. Therefore, using intelligent means to improve the efficiency and accuracy of manhole cover hidden danger detection has become an urgent problem to be solved in urban management.

[0004] In recent years, deep learning technology has made remarkable progress in the field of image recognition and processing, especially in object detection tasks. Deep learning models, such as convolutional neural networks (CNNs), have achieved efficient recognition and classification of objects in images by automatically learning image features. Object detection algorithms based on deep learning, such as Faster R-CNN, YOLO, and SSD, have been widely used in many fields due to their high detection speed and excellent accuracy.

[0005] However, in the field of intelligent detection of manhole cover hidden dangers, there are still some challenges and problems in the existing technology. First, the diversity of the shape, material, and environmental background of manhole covers, as well as the complexity of the urban environment, such as light changes and weather conditions, pose high requirements for the adaptability and accuracy of intelligent detection systems. Second, existing object detection algorithms still have problems with low detection accuracy when dealing with small objects, occluded objects, and objects in complex backgrounds. In addition, the need to process a large amount of image data in real time poses a challenge to the computational efficiency of the algorithm. Summary of the Invention

[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problem of low model detection accuracy in the prior art due to the complex environmental background and small targets of manhole covers.

[0007] To solve the above technical problem, the present invention provides a manhole cover hidden danger detection method based on object detection, including:

[0008] Input the real-time collected manhole cover image into the backbone network of the target detection model. Through multiple sequentially cascaded extraction units, extract multiple primary feature maps of different scales; each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module cascaded in sequence;

[0009] Input the multiple primary feature maps of different scales into the neck network of the target detection model based on the FPN structure, and output the fused feature map corresponding to each scale; the top-down path in the neck network includes a dynamic upsampling module; the bottom-up path in the neck network includes an inverted residual attention downsampling module;

[0010] Input the fused feature map corresponding to each scale into the corresponding detection head in the head network of the target detection model, and output the prediction labels, prediction bounding boxes, and confidence levels of the fused feature maps of each scale corresponding to the real-time collected manhole cover image; the prediction labels are manhole cover hidden danger categories, including intact manhole cover, damaged manhole cover, uncovered manhole cover, manhole ring problem, and lost manhole cover;

[0011] Among them, for the inverted residual attention downsampling module, after performing average pooling on the input feature map, it is evenly divided into two paths according to the number of channels. One path passes through convolution, and the other path passes through max pooling and convolution, and then the two paths are concatenated by channel to obtain a concatenated feature map; extract the attention feature map of the concatenated feature map, and after performing depthwise separable convolution, fuse it with the attention feature map and the concatenated feature map to output a downsampled feature map;

[0012] For the dynamic upsampling module, divide the input feature map into two paths. One path passes through convolution, and the other path passes through convolution and activation, and then multiply the two paths element-wise to obtain a combined feature map; perform pixel shuffle on the combined feature map to obtain a sampling offset; add the sampling offset to the original sampling grid to obtain a new sampling grid, and sample the input feature map to output an upsampled feature map.

[0013] Preferably, the inverted residual attention downsampling module is specifically used for:

[0014] After the input feature map of the inverted residual attention downsampling module passes through average pooling, it is evenly divided into two-channel feature maps according to the number of channels;

[0015] One-channel feature map passes through a 3×3 convolution and is sent to the concatenation unit; the other-channel feature map passes through a max pooling layer and a 1×1 convolution and is then sent to the concatenation unit;

[0016] The concatenation unit outputs a concatenated feature map, which is divided into two paths. One path of the feature map is directly sent to the fusion unit; the other path of the feature map is based on the self-attention mechanism. After obtaining the attention map, multiply it by itself to obtain an attention feature map;

[0017] Let the attention feature map pass through a 3×3 depthwise separable convolution and then fuse with the attention feature map, followed by a 1×1 convolution, and then be fed into the fusion unit;

[0018] Fuse the feature maps in the fusion unit and output them as the downsampled feature map.

[0019] Preferably, the backbone network includes a convolutional unit, an efficient aggregation module, and multiple extraction units connected in series in sequence.

[0020] Preferably, input multiple primary feature maps of different scales into the neck network based on the FPN structure of the object detection model, and output the fused feature map corresponding to each scale, including:

[0021] Pass the primary feature map of the first scale through the SP pooling module to obtain the pooled feature map of the first scale; after processing the pooled feature map of the first scale through the dynamic upsampling module, splice it with the primary feature map of the second scale and output the upsampled and spliced feature map of the second scale;

[0022] Pass the upsampled and spliced feature map of the k-th scale through the SP pooling module to obtain the pooled feature map of the k-th scale; after processing the pooled feature map of the k-th scale through the dynamic upsampling module, splice it with the primary feature map of the k+1-th scale and output the upsampled and spliced feature map of the k-th scale; k = 1, 2,..., K - 1;

[0023] Pass the upsampled and spliced feature map of the K-th scale through the efficient aggregation module and output the fused feature map of the K-th scale;

[0024] After processing the fused feature map of the K-th scale through the inverted residual attention downsampling module, splice it with the pooled feature map of the K - 1-th layer to obtain the downsampled and spliced feature map of the K - 1-th layer, and then process it through the efficient aggregation module to output the fused feature map of the K - 1-th scale; successively obtain the fused feature maps of the K - 2-th scale to the first scale.

[0025] Preferably, the efficient aggregation module is a generalized efficient layer aggregation network, which is specifically used for:

[0026] After performing convolution on the input feature map of the efficient aggregation module, split it into two paths on average according to the number of channels, and directly send both paths of feature maps into the splicing unit of the efficient aggregation module;

[0027] Arbitrarily select one path of the feature map, pass it through RepNCSP and convolution in sequence, then divide it into two paths, directly send one path of the feature map into the splicing unit of the efficient aggregation module, and send the other path of the feature map into the splicing unit of the efficient aggregation module after passing through RepNCSP and convolution;

[0028] After splicing the four-way feature maps in the splicing unit of the efficient aggregation module and then performing convolution, the output of the efficient aggregation module is obtained;

[0029] Among them, the RepNCSP is used to: divide the input feature map of the RepNCSP into two paths, one path is sent to the splicing unit in the RepNCSP after convolution; the other path is divided into two paths after convolution, and after performing convolutions of different scales and adding them together, and then performing convolution to obtain a convolutional feature map, which is sent to the splicing unit in the RepNCSP; after splicing the two-way feature maps in the splicing unit of the RepNCSP and then performing convolution output, the output of the RepNCSP is obtained.

[0030] Preferably, the object detection model is obtained by replacing the upsampling module and the downsampling module in the YOLOv9 model with the dynamic upsampling module and the inverted residual attention downsampling module respectively.

[0031] Preferably, the training process of the object detection model includes:

[0032] Obtain manhole cover images corresponding to different manhole cover hazard categories as training images, and obtain the true labels and true prediction boxes corresponding to the training images;

[0033] Input the training images into the programmable gradient information module and the backbone network of the object detection model;

[0034] Based on the multiple feature maps of different scales output by the backbone network and the feature maps passing through the efficient aggregation module and the inverted residual attention downsampling module in the programmable gradient information module, perform cross-branch fusion, and output the predicted labels, predicted bounding boxes and confidence levels corresponding to the feature maps of each scale;

[0035] Based on the predicted labels and the true labels, calculate the cross-entropy classification loss of the training images;

[0036] Use the IoU loss to calculate the regression loss of the training images based on the overlapping degree between the true bounding boxes and the predicted bounding boxes;

[0037] Based on the weighted sum of the cross-entropy classification loss and the regression loss, construct a total loss function;

[0038] Train the object detection model, perform gradient backpropagation based on the value of the total loss function, and let the optimizer update the model parameters until the preset number of iterations is reached, and obtain the pre-trained object detection model.

[0039] Preferably, the total loss function is expressed as:

[0040] ;

[0041] Wherein, denotes the total loss function; denotes the classification loss, and the expression is , denotes the true label, denotes the predicted probability of the class ; denotes the regression loss, and the expression is , denotes the predicted bounding box, denotes the true bounding box, denotes the intersection of the predicted bounding box and the true bounding box and the union of which the ratio, and the expression is ; and are loss weight hyperparameters.

[0042] Preferably, the programmable gradient information module is specifically used for:

[0043] Input the multiple primary feature maps of different scales output by the backbone network into the cross-branch fusion modules of the corresponding scales in the programmable gradient information module respectively;

[0044] Let the training image pass through the convolution, efficient aggregation module and inverted residual attention downsampling module, and then input it into the cross-branch fusion module of the Kth scale;

[0045] Initialize k = K, update k = k - 1 until k = 1; Output the output of the cross-branch fusion module of the kth scale through the efficient aggregation module to obtain the aggregated feature map of the kth scale; After passing the aggregated feature map of the kth scale through the inverted residual attention downsampling module, input it into the cross-branch fusion module of the k - 1th scale; Output the aggregated feature map of the kth scale through the detection head to obtain the class label, bounding box and confidence of the kth scale;

[0046] Obtain the class label, bounding box and confidence of each scale as the output of the programmable gradient information module.

[0047] The above technical solutions of the present invention have the following beneficial effects compared with the prior art:

[0048] In the manhole cover hidden danger detection method based on object detection described in the present invention, a dynamic upsampling module and an inverted residual attention downsampling module are proposed.

[0049] The dynamic upsampling module defines upsampling as a point sampling problem. By generating sampling offsets to adjust the sampling network, it can process the features corresponding to the manhole cover image more flexibly, enabling the model to better capture the information of small targets or fine structures in the manhole cover image, improving the detection and segmentation accuracy of the details in the manhole cover hidden dangers, and thus enhancing the detection accuracy of manhole covers with different shapes and materials in different environments. Moreover, under the condition of obtaining stronger performance, the dynamic upsampling module still has fewer parameters, floating-point operation counts, GPU memory, and latency, has high timeliness, and is more suitable for real-time monitoring of road manhole covers.

[0050] After depthwise separable convolution in the inverted residual attention downsampling module, an attention mechanism, depthwise separable convolution, and residual mechanism are introduced, expanding the convolution-based inverted residual block to an attention-based model. This enables the model to combine the static short-range modeling ability of CNN and the dynamic long-range feature interaction ability of Transformer. It can not only effectively compress data but also enhance the attention to key features while reducing information loss, improving the model's feature expression and detection ability for targets. Consequently, it can accurately extract the features corresponding to the targets in the manhole cover image, enhancing the detection accuracy of hidden dangers in the manhole cover image.

[0051] Based on the YOLOv9 model, this invention introduces a dynamic upsampling module and an inverted residual attention downsampling module, enhancing the feature extraction ability and environmental adaptability of the YOLOv9 model, improving the detection accuracy and robustness of the model for the hidden danger state of manhole covers in complex environments, and still maintaining high accuracy when dealing with small targets, occluded targets, and targets in complex backgrounds. At the same time, it realizes the efficient processing of a large amount of image data, meeting the requirements of real-time monitoring. Moreover, when training the object detection model based on the YOLOv9 model, a programmable gradient information module is utilized to generate reliable gradients through an auxiliary reversible branch, programming the gradient information propagation at different semantic levels, enabling the deep features to still maintain the key characteristics for performing the target task, avoiding the semantic loss that may be caused by the traditional deep supervision process, improving the extraction accuracy of the image features corresponding to the hidden dangers in the manhole cover, enhancing the model training accuracy, and thus improving the detection accuracy of manhole cover hidden dangers. Brief Description of the Drawings

[0052] To make the content of this invention easier to understand clearly, the following further elaborates on this invention according to the specific embodiments of this invention in combination with the attached drawings, where:

[0053] Figure 1 is the step flow chart of the manhole cover hidden danger detection method provided by this invention;

[0054] Figure 2 is a schematic diagram of manhole cover hidden dangers;

[0055] Figure 3 Schematic diagram of the inverted residual attention downsampling module;

[0056] Figure 4 Schematic diagram of the dynamic upsampling module;

[0057] Figure 5 Schematic diagram of the efficient aggregation module;

[0058] Figure 6 Schematic diagram of the SP pooling module;

[0059] Figure 7 Schematic diagram of the object detection model based on the improved YOLOv9 model;

[0060] Figure 8 Schematic diagram of the cross-branch fusion module;

[0061] Figure 9 Schematic diagram of the user interface;

[0062] Figure 10 Schematic diagram of the detection result of the user interface;

[0063] Figure 11 Schematic diagram of the input selection of the user interface;

[0064] Figure 12 Schematic diagram of the detection result of the intact manhole cover;

[0065] Figure 13 Schematic diagram of the detection result of the uncovered manhole cover;

[0066] Figure 14 Schematic diagram of the detection result of the damaged manhole cover;

[0067] Figure 15 Schematic diagram of the detection result of the manhole ring problem;

[0068] Figure 16 Schematic diagram of the detection result of the lost manhole cover. Specific implementation manner

[0069] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited are not intended to limit the present invention.

[0070] Referring to Figure 1 As shown, the step flow chart of the manhole cover hidden danger detection method based on object detection provided by the present invention specifically includes:

[0071] S101: Input the real-time collected manhole cover image into the backbone network of the target detection model. Through multiple sequentially cascaded extraction units, extract multiple primary feature maps of different scales; each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module cascaded in sequence.

[0072] S102: Input the multiple primary feature maps of different scales into the neck network of the target detection model based on the FPN structure, and output the fused feature map corresponding to each scale; the top-down path in the neck network includes a dynamic upsampling module; the bottom-up path in the neck network includes an inverted residual attention downsampling module.

[0073] Among them, the Feature Pyramid Network (FPN) structure includes a top-down upsampling path and a bottom-up downsampling path, with horizontal connections between the upsampling path and the downsampling path; the top-down path uses the dynamic upsampling module to perform upsampling operations on the feature map, and the bottom-up path uses the inverted residual attention downsampling module to perform downsampling operations on the feature map.

[0074] S103: Input the fused feature map corresponding to each scale into the corresponding detection head in the head network of the target detection model, and output the prediction label, prediction bounding box, and confidence of the fused feature map of each scale corresponding to the real-time collected manhole cover image; the prediction label is the manhole cover hidden danger category, including intact manhole cover, damaged manhole cover, uncovered manhole cover, manhole ring problem, and lost manhole cover.

[0075] Among them, for the inverted residual attention downsampling module, after performing average pooling on the input feature map, it is evenly divided into two paths according to the number of channels. One path goes through convolution, and the other path goes through max pooling and convolution, and then the two paths are concatenated by channel to obtain a concatenated feature map; extract the attention feature map of the concatenated feature map, and after performing depthwise separable convolution, fuse it with the attention feature map and the concatenated feature map to output the downsampled feature map.

[0076] For the dynamic upsampling module, divide the input feature map into two paths. One path goes through convolution, and the other path goes through convolution and activation, and then multiply the two paths pointwise to obtain a combined feature map; perform pixel shuffling on the combined feature map to get the sampling offset; add the sampling offset to the original sampling grid to obtain a new sampling grid, and sample the input feature map to output the upsampled feature map.

[0077] Refer to Figure 2 As shown in the figure, it is a schematic diagram of manhole cover hidden dangers; specifically, the prediction label is the manhole cover hidden danger category, including intact good, damaged broke, uncovered uncovered, manhole ring problem circle, and lost lose.

[0078] Reference Figure 3 As shown in Figure 3 , it is a schematic diagram of an inverted residual attention downsampling module; the inverted residual attention downsampling module is specifically used for:

[0079] After the input feature map of the inverted residual attention downsampling module passes through average pooling, it is evenly divided into two-channel feature maps according to the number of channels;

[0080] One-channel feature map is sent to the splicing unit after passing through a 3×3 convolution; the other-channel feature map is sent to the splicing unit after passing through a max pooling layer and a 1×1 convolution;

[0081] The splicing unit outputs a spliced feature map, which is divided into two paths. One path of the feature map is directly sent to the fusion unit; the other path of the feature map is based on the self-attention mechanism. After obtaining the attention map, it is multiplied by itself to obtain the attention feature map;

[0082] After the attention feature map passes through a 3×3 depthwise separable convolution, it is fused with the attention feature map, and then passes through a 1×1 convolution and is sent into the fusion unit;

[0083] The feature maps in the fusion unit are fused and output as the downsampled feature map.

[0084] The inverted residual attention downsampling module (Attention-based Inverted Residual Downsampling, AIRD) connects an inverted residual mobile module behind a split-channel convolution module, extending the convolution-based inverted residual block to an attention-based model, enabling the model to combine the static short-range modeling ability of CNN and the dynamic long-range feature interaction ability of Transformer. Among them, the 3*3 convolution is a conventional convolution with a stride of 2 and the output channels remaining unchanged. Therefore, after the feature map passes through it, the size will be halved and the number of channels remains unchanged, which can be understood as downsampling the feature map; the 1*1 convolution on the right has a stride of 1 and the output channels remain unchanged. Therefore, after the feature map passes through it, both the number of channels and the size remain unchanged, and only a 1*1 convolution operation is performed; the first 1*1 convolution after splicing the features does not change the size but adjusts the number of channels to generate the V value in the attention operation. Generally, the QKV values in the attention operation are obtained through a linear layer. Here, the 1*1 convolution acts as a linear layer; the 3*3 depthwise separable convolution is a special convolution kernel. Different from ordinary convolutions, each depthwise separable convolution only performs convolution calculation operations on a single channel (normal convolutions calculate all channels of the feature map), which is relatively lightweight and has less computational complexity compared to ordinary convolutions. The 3*3 depthwise separable convolution here does not change the number of channels and dimensions; the role of the bottom 1*1 convolution is to adjust the number of channels so that it can be the same as the number of channels of the spliced features for addition operations to achieve residual connection.

[0085] Reference Figure 4 As shown in Figure 4 , it is a schematic diagram of the dynamic upsampling module; different from traditional dynamic convolution, dynamic upsampling (Dynamic Upsample, DyUp) adopts a more resource-efficient upsampling method from the perspective of point sampling. It bypasses dynamic convolution and instead redefines the upsampling problem as a point sampling problem, strengthening the behavior of traditional upsampling. Dynamic upsampling enables the model to have fewer parameters, floating-point operation counts, GPU memory, and latency while achieving stronger performance. Dynamic upsampling not only has superior performance in the field of object detection but also has strong performance in semantic segmentation and instance segmentation.

[0086] Specifically, the backbone network includes a convolutional unit, an efficient aggregation module, and multiple extraction units connected in series in sequence. The backbone network is used to obtain multiple primary feature maps of different scales, including: inputting the real-time collected manhole cover image into the backbone network, passing through the convolutional unit, the efficient aggregation module, and multiple extraction units in sequence; each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module connected in series; the output of the efficient aggregation module in each extraction unit is used as the primary feature map of the corresponding scale.

[0087] Specifically, inputting multiple primary feature maps of different scales into the neck network based on the FPN structure of the object detection model to output the fused feature map corresponding to each scale, including:

[0088] Passing the primary feature map of the first scale through the SP pooling module to obtain the pooled feature map of the first scale; after processing the pooled feature map of the first scale through the dynamic upsampling module, splicing it with the primary feature map of the second scale to output the upsampled and spliced feature map of the second scale;

[0089] Passing the upsampled and spliced feature map of the k-th scale through the SP pooling module to obtain the pooled feature map of the k-th scale; after processing the pooled feature map of the k-th scale through the dynamic upsampling module, splicing it with the primary feature map of the k + 1-th scale to output the upsampled and spliced feature map of the k-th scale; k = 1, 2,..., K - 1;

[0090] Passing the upsampled and spliced feature map of the K-th scale through the efficient aggregation module to output the fused feature map of the K-th scale;

[0091] After processing the fused feature map of the K-th scale through the inverted residual attention downsampling module, splicing it with the pooled feature map of the K - 1-th layer to obtain the downsampled and spliced feature map of the K - 1-th layer, and then passing it through the efficient aggregation module for processing to output the fused feature map of the K - 1-th scale; successively obtain the fused feature maps of the K - 2-th scale to the first scale.

[0092] Reference Figure 5As shown in the figure, it is a schematic diagram of an efficient aggregation module; the efficient aggregation module is a Generalized Efficient Layer Aggregation Network, which is specifically used for:

[0093] After convolving the input feature map of the efficient aggregation module, it is evenly split into two paths according to the number of channels, and both paths of feature maps are directly sent into the splicing unit of the efficient aggregation module;

[0094] Arbitrarily select one path of the feature map, which is successively passed through RepNCSP and convolution, and then split into two paths. One path of the feature map is directly sent into the splicing unit of the efficient aggregation module, and the other path of the feature map is passed through RepNCSP and convolution and then sent into the splicing unit of the efficient aggregation module;

[0095] After splicing the four paths of feature maps in the splicing unit of the efficient aggregation module, through convolution, the output of the efficient aggregation module is obtained;

[0096] Among them, the RepNCSP combines the reparameterization technology with the CSP (Cross Stage Partial) structure, which not only retains the efficient feature fusion ability of CSP, but also improves the inference speed through the reparameterization technology. It is used for: splitting the input feature map of RepNCSP into two paths, one path is sent into the splicing unit in RepNCSP after convolution; the other path is convolved and then split into two paths again, and after convolutions of different scales and addition, and then convolution, the convolution feature map is obtained and sent into the splicing unit in RepNCSP; after splicing the two paths of feature maps in the splicing unit of RepNCSP, through convolution output, the output of RepNCSP is obtained.

[0097] Refer to Figure 6 As shown in the figure, it is a schematic diagram of an SP pooling module; the SP pooling module is based on Spatial Pyramid Pooling (SPP); the SP pooling module is a Spatial Pyramid Pooling Enhanced with ELAN, which is used for:

[0098] Let the input features of the SP pooling module pass through the convolution unit, the first max-pooling unit, the second max-pooling unit and the third max-pooling unit in sequence;

[0099] After splicing the outputs of each unit and passing through the convolution unit for output, the output of the SP pooling is obtained.

[0100] In the manhole cover hidden danger detection method based on object detection proposed in the present invention, a dynamic upsampling module and an inverted residual attention downsampling module are proposed. The dynamic upsampling module defines upsampling as a point sampling problem and adjusts the sampling network by generating sampling offsets, which can process the features corresponding to the manhole cover image more flexibly, enabling the model to better capture the information of small targets or fine structures in the manhole cover image, improving the detection and segmentation accuracy of the details in the manhole cover hidden danger, and thus improving the hidden danger detection accuracy of manhole covers with different shapes and materials in different environments; moreover, the dynamic upsampling module still has fewer parameters, floating-point operation counts, GPU memory, and latency while obtaining stronger performance, has high timeliness, and is more suitable for real-time monitoring of road manhole covers. After channel-wise convolution, the inverted residual attention downsampling module introduces an attention mechanism, depthwise separable convolution, and a residual mechanism, extending the convolution-based inverted residual block to an attention-based model, enabling the model to combine the static short-range modeling ability of CNN and the dynamic long-range feature interaction ability of Transformer, not only effectively compressing data but also enhancing the attention to key features while reducing information loss, improving the model's feature expression and detection ability for targets, and thus being able to accurately extract the features corresponding to the targets in the manhole cover image and improving the detection accuracy of hidden dangers in the manhole cover image.

[0101] Based on the above embodiments, in the embodiments of the present invention, the upsampling module and the downsampling module in the YOLOv9 model are respectively replaced with the dynamic upsampling module and the inverted residual attention downsampling module to obtain an object detection model for detecting manhole cover hidden dangers; among them, the training process of the object detection model based on the improved YOLOv9 model includes:

[0102] S201: Obtain manhole cover images corresponding to different manhole cover hidden danger categories as training images, and obtain the true labels and true prediction boxes corresponding to the training images;

[0103] S202: Input the training images into the programmable gradient information module and the backbone network of the object detection model;

[0104] S203: Based on the feature maps of multiple different scales output by the backbone network and the feature maps in the programmable gradient information module that have passed through the efficient aggregation module and the inverted residual attention downsampling module, perform cross-branch fusion (Cross-BranchFuse) to output the predicted labels, predicted bounding boxes, and confidence levels corresponding to the feature maps of each scale;

[0105] S204: Calculate the cross-entropy classification loss of the training images based on the predicted labels and the true labels;

[0106] S205: Calculate the regression loss of the training image using the IoU loss based on the overlap degree between the ground truth bounding box and the predicted bounding box.

[0107] S206: Construct the total loss function based on the weighted sum of the cross-entropy classification loss and the regression loss, expressed as: ;

[0108] where, represents the total loss function; represents the classification loss, and the expression is , represents the ground truth label, represents the predicted probability of the class ; represents the regression loss, and the expression is , represents the predicted bounding box, represents the ground truth bounding box, represents the intersection of the predicted bounding box and the ground truth bounding box and the union ratio, and the expression is ; and

[0109] are loss weight hyperparameters.

[0110] Refer to Figure 7 shown in the figure, which is a schematic diagram of the object detection model based on the improved YOLOv9 model; the model consists of four parts: a programmable gradient information module, a backbone network, a neck network, and a head network.

[0111] The programmable gradient information module is specifically used for:

[0112] Input the multiple primary feature maps with different scales output by the backbone network into the cross-branch fusion modules with corresponding scales in the programmable gradient information module respectively;

[0113] Let the training image pass through a convolution, an efficient aggregation module, and an inverted residual attention downsampling module, and then input it into the cross-branch fusion module of the Kth scale;

[0114] Initialize \(k = K\), update \(k=k - 1\) until \(k = 1\); pass the output of the cross-branch fusion module at the \(k\)-th scale through the efficient aggregation module to output the aggregated feature map at the \(k\)-th scale; after passing the aggregated feature map at the \(k\)-th scale through the inverted residual attention downsampling module, input it into the cross-branch fusion module at the \((k - 1)\)-th scale; pass the aggregated feature map at the \(k\)-th scale through the detection head to obtain the class labels, bounding boxes, and confidences at the \(k\)-th scale.

[0115] Obtain the class labels, bounding boxes, and confidences at each scale as the output of the programmable gradient information module.

[0116] The Programmable Gradient Information (PGI) module generates reliable gradients through an auxiliary reversible branch, enabling deep features to still maintain the key characteristics for performing the target task. The design of the auxiliary reversible branch can avoid the semantic loss that may occur in the traditional deep supervision process, which integrates multi-path features. In other words, PGI achieves optimal training results by programming gradient information propagation at different semantic levels. The reversible architecture of PGI is built on the auxiliary branch, so there is no additional cost. Since PGI can freely choose the loss function suitable for the target task, it is more general than the deep supervision mechanism, which is only applicable to very deep neural networks. During training, the programmable gradient information module passes information to the main module, and when inferring, there is no need to pass through the programmable gradient information module anymore.

[0117] Refer to Figure 8 As shown, it is a schematic diagram of the cross-branch fusion module; the cross-branch fusion module includes calculating the nearest neighbor difference of multiple input features and adding them to obtain the output of the cross-branch fusion module.

[0118] In this embodiment, a Dynamic Upsample (DyUp) module, a Programmable Gradient Information (PGI) module, and an Attention-based Inverted Residual Downsampling (AIRD) module are used in combination with YOLOv9 to enhance the feature extraction ability and environmental adaptability of the model, not only improving the accuracy of intelligent detection of manhole cover hazards but also enhancing the robustness of the model under different lighting and weather conditions.

[0119] Specifically, in this embodiment, the training data with annotations is used to adjust the model parameters so that it can learn appropriate feature representations from the data, thereby learning the knowledge and rules for solving the hidden danger problems of manhole covers. For the N images from different categories in a batch B input to the object detection model, the main process of model training includes:

[0120] ① Image basic feature extraction;

[0121] Through the innovative backbone network of YOLOv9, namely the Generalized Efficient Layer Aggregation Network, the basic features of each image are obtained. Through a series of upsampling and downsampling operations, feature maps of multiple scales are generated. The operations usually include pooling layers, convolutional layers, and interpolation layers;

[0122] In this embodiment, two innovative upsampling function and downsampling function modules are used, namely the dynamic upsampling module and the inverted residual self-attention module, which can extract semantic information more effectively. Each scale of feature map corresponds to a different resolution of the input image. Fusing feature maps of different scales can enhance the model's detection ability for objects of different sizes. The fusion method can be weighted summation, concatenation, or a more complex attention mechanism. In this way, the low-level feature maps with high resolution but less semantic information and the high-level feature maps with low resolution but rich semantic information can complement each other to improve the detection accuracy. Finally, three feature maps of different sizes are obtained, arranged from large to small in scale, which are , , .

[0123] ② According to the extracted basic features, perform classification and regression;

[0124] The feature maps are input into the fully connected layer for classification, and the cross-entropy classification loss is expressed as:

[0125] ;

[0126] Among them, is the predicted probability of the model for class , is the true label.

[0127] For the regression task, this embodiment uses the IoU (Intersection over Union) loss to measure the overlap degree between the predicted bounding box and the true bounding box. IoU is the ratio of the intersection to the union of two bounding boxes. Assume is the predicted bounding box, is the true bounding box, and the IoU loss is expressed as:

[0128] ;

[0129] Therefore, the regression loss, denoted as: ;

[0130] In actual training, the total loss is the weighted sum of the classification loss and the regression loss, denoted as:

[0131] ;

[0132] This total loss function only considers the case of single-scale feature maps. In actual calculation, considering that the model will calculate feature maps of different scales, therefore, the classification loss and the regression loss corresponding to the feature maps of each scale are calculated separately to form the total loss function, denoted as:

[0133] ;

[0134] Among them, and are respectively the classification loss and the regression loss of the feature map of the th scale.

[0135] ③ According to the total loss function, use the optimizer to optimize the model parameters so that the model can accurately detect various hidden problems of manhole covers.

[0136] Based on the above description, the model training process includes:

[0137] Input: Dataset , is an image, is the corresponding label, and the label contains the bounding box and the category ; the total number of training epochs required ;

[0138] Model structure: backbone network (feature extractor ), neck network (feature fusion ), head network (model prediction head ), auxiliary branch programmable gradient information module;

[0139] Training process:

[0140] ① Initialize the model model, initialize the optimizer optimizer, initialize the total loss function L, and define the data loader train_loader;

[0141] ② Training loop:

[0142] for the current training epoch t = 1 to MAX_T (preset total training epochs) do:

[0143] Set the model to the training mode;

[0144] For the current iteration number i of traversing the dataset from 1 to MAX_I (the number of iterations required to iterate through the dataset) do:

[0145] Clear the gradients retained by the optimizer;

[0146] Sample a sample , and at the same time obtain the class label of the sample and the ground truth bounding box ;

[0147] Obtain the multi-scale features of the sample through the backbone of the model ;

[0148] Obtain the fused features through the neck of the model ;

[0149] Obtain the final prediction results through the head of the model, including: predicted object bounding box , confidence and class probability distribution ;

[0150] Calculate the classification loss, regression loss, and total loss, perform gradient backpropagation according to the total loss, and let the optimizer update the model parameters;

[0151] End for;

[0152] End for;

[0153] ③ Save the trained model model.

[0154] The embodiment of the present invention uses the manhole cover hidden danger detection method provided by the present invention to perform detection based on the 2024 Service Innovation Competition A03 manhole cover dataset; the 2024 Service Innovation Competition A03 manhole cover dataset is a dataset specifically for intelligent detection of manhole cover hidden dangers, containing more than 1300 manhole cover pictures in various different states, such as intact, damaged, missing, uncovered, and manhole ring problems, etc.; this dataset aims to achieve automatic detection of manhole cover hidden dangers through object detection technology to improve urban management efficiency and public safety. In the training stage, the stochastic gradient descent method is used, with a decay rate of , and a momentum of 0.9; the model is trained for 100 epochs, and the batch size of each batch is 16. To enhance the training data, the following data augmentation methods are used in this embodiment: randomly select an image, directly resize it to 256×256, and then crop it to 224×224; horizontal flip.

[0155] As shown in Table 1, in this embodiment, based on YOLOv9, the effectiveness of the method proposed by the present invention was verified; the prediction results were evaluated on the A03 manhole cover data test set of the 2024 Service Innovation Competition.

[0156] Table 1 Comparison of prediction results of YOLOv9 models with different functional modules deployed

[0157]

[0158] Referring to Table 1, it can be seen that the mAP value of the manhole cover hidden danger detection method based on object detection provided by the present invention can reach 0.78, which can effectively detect the manhole cover hidden danger problem.

[0159] The YOLO model has been rapidly iterating. In this embodiment, a comparison was made between the YOLOv9-based model and the object detection model improved based on the YOLOv5 model: when both use AR downsampling and Dy upsampling, using YOLOv9-c as the baseline model, the map50 is 0.779, and when using YOLOv5-l, it is 0.753. The present invention is based on the YOLOv9 model, introducing a dynamic upsampling module and an inverted residual attention downsampling module, enhancing the feature extraction ability and environmental adaptability of the YOLOv9 model, improving the detection accuracy and robustness of the model for the hidden danger state of manhole covers in complex environments, and at the same time achieving efficient processing of a large amount of image data, meeting the requirements of real-time monitoring. Moreover, when the object detection model based on the YOLOv9 model is trained, a programmable gradient information module is utilized to generate reliable gradients through an auxiliary reversible branch, programming the gradient information propagation at different semantic levels, so that the deep features still maintain the key characteristics for performing the target task, avoiding the semantic loss that may be caused by the traditional deep supervision process, improving the model training accuracy, enabling the network to better adapt to the difficult data distribution of manhole cover problems, and at the same time, being able to extract richer and more accurate local features, effectively reducing false detections.

[0160] Based on the above embodiments, in the embodiments of the present invention, when deploying after the model training is completed, this embodiment uses pytq5 to develop the application interface. During inference, first, the input image is preprocessed in the same way as during training, then the pre-trained model is loaded, the image is input into the loaded model and the prediction results are obtained, and then the prediction results are fed back to the user side. Referring to Figure 9 As shown, it is a schematic diagram of the user interface; the user interface includes various data input formats. The parts boxed in the figure from top to bottom are: file selection for detection from a file, start camera detection, detection from a network video stream, and click the button to select a file. Referring to Figure 10As shown, it is a schematic diagram of the user interface detection result; when a file is selected for detection, the text detection result will be displayed on the left after the detection is completed, and the image detection result will be displayed on the right side of the interface. Refer to Figure 11 As shown, it is a schematic diagram of the user interface input selection; the user interface provides parameter settings for detection, and the supported input parameters include IoU (Intersection over Union), confidence level, and frame delay.

[0161] Based on the manhole cover hidden danger detection model provided by the present invention, images of manhole covers with various different hidden danger categories are detected, and the detection results are as Figures 12 to 16 shown. Refer to Figure 12 As shown, it is a schematic diagram of the detection result of a sound manhole cover; refer to Figure 13 As shown, it is a schematic diagram of the detection result of an uncovered manhole cover; refer to Figure 14 As shown, it is a schematic diagram of the detection result of a damaged manhole cover; refer to Figure 15 As shown, it is a schematic diagram of the detection result of manhole ring problems; refer to Figure 16 As shown, it is a schematic diagram of the detection result of a missing manhole cover.

[0162] The manhole cover hidden danger detection method based on object detection provided by the present invention can be deployed as a corresponding detection system. The main functions of the system include but are not limited to real-time monitoring and detection, abnormal detection and warning, and user interface and interaction. Through the deployment and application of this system, the efficiency and level of urban manhole cover safety management can be improved, accidents caused by manhole cover hidden dangers can be reduced, and a safer living environment can be provided for urban residents.

[0163] In the manhole cover hidden danger detection method based on object detection proposed in the present invention, a dynamic upsampling module and an inverted residual attention downsampling module are proposed. The dynamic upsampling module defines upsampling as a point sampling problem. By generating sampling offsets to adjust the sampling network, it can process the features corresponding to the manhole cover image more flexibly, enabling the model to better capture the information of small targets or fine structures in the manhole cover image, improving the detection and segmentation accuracy of the details in the manhole cover hidden danger, and thus improving the hidden danger detection accuracy of manhole covers with different shapes and materials in different environments. Moreover, when the dynamic upsampling module obtains stronger performance, it still has fewer parameters, floating-point operation counts, GPU memory, and latency, has high timeliness, and is more suitable for real-time monitoring of road manhole covers. After channel-wise convolution, the inverted residual attention downsampling module introduces an attention mechanism, depthwise separable convolution, and a residual mechanism, extending the convolution-based inverted residual block to an attention-based model, enabling the model to combine the static short-range modeling ability of CNN and the dynamic long-range feature interaction ability of Transformer. It can not only effectively compress data but also enhance the attention to key features while reducing information loss, improving the model's feature expression and detection ability for targets, and thus being able to accurately extract the features corresponding to the targets in the manhole cover image, improving the detection accuracy of hidden dangers in the manhole cover image. Based on the YOLOv9 model, the present invention introduces a dynamic upsampling module and an inverted residual attention downsampling module, enhancing the feature extraction ability and environmental adaptability of the YOLOv9 model, improving the detection accuracy and robustness of the model for the hidden danger state of manhole covers in complex environments, and still having high accuracy when dealing with small targets, occluded targets, and targets in complex backgrounds. At the same time, it realizes the efficient processing of a large amount of image data, meeting the requirements of real-time monitoring. Moreover, when the object detection model based on the YOLOv9 model is trained, a programmable gradient information module is used to generate reliable gradients through an auxiliary reversible branch, programming the gradient information propagation at different semantic levels, so that the deep features still maintain the key characteristics of performing the target task, avoiding the semantic loss that may be caused by the traditional deep supervision process, improving the extraction accuracy of the image features corresponding to the hidden dangers in the manhole cover, improving the model training accuracy, and thus improving the manhole cover hidden danger detection accuracy.

[0164] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a means for implementing the functions specified in one block or multiple blocks.

[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a means for implementing the functions specified in one block or multiple blocks.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a means for implementing the functions specified in one block or multiple blocks.

[0168] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A manhole cover hidden danger detection method based on object detection, characterized in that Including: Input the real-time collected manhole cover image into the backbone network of the object detection model. Through multiple sequentially connected extraction units, extract multiple primary feature maps of different scales; each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module connected in series in sequence; Input the multiple primary feature maps of different scales into the neck network of the object detection model based on the FPN structure, and output the fused feature map corresponding to each scale; the top-down path in the neck network includes a dynamic upsampling module; the bottom-up path in the neck network includes an inverted residual attention downsampling module; Input the fused feature map corresponding to each scale into the corresponding detection head in the head network of the object detection model, and output the prediction label, prediction bounding box and confidence of the fused feature map of each scale corresponding to the real-time collected manhole cover image; the prediction label is the manhole cover hidden danger category, including intact manhole cover, damaged manhole cover, uncovered manhole cover, manhole ring problem and lost manhole cover; Among them, for the inverted residual attention downsampling module, after performing average pooling on the input feature map, it is evenly divided into two paths according to the number of channels. One path passes through convolution, and the other path passes through max pooling and convolution, and then the two paths are concatenated according to the channels to obtain a concatenated feature map; extract the attention feature map of the concatenated feature map, and after performing depthwise separable convolution, fuse it with the attention feature map and the concatenated feature map, and output the downsampled feature map; For the dynamic upsampling module, divide the input feature map into two paths. One path passes through convolution, and the other path passes through convolution and activation, and then the two paths are multiplied pointwise to obtain a combined feature map; perform pixel shuffling on the combined feature map to obtain a sampling offset; Add the sampling offset to the original sampling grid, sample the input feature map, and output the upsampled feature map.

2. The manhole cover hidden danger detection method based on object detection according to claim 1, characterized in that The inverted residual attention downsampling module is specifically used for: After the input feature map of the inverted residual attention downsampling module passes through average pooling, it is evenly divided into two-channel feature maps according to the number of channels; One-channel feature map passes through a 3×3 convolution and is sent to the concatenation unit; the other-channel feature map passes through a max pooling layer and a 1×1 convolution and is sent to the concatenation unit; The concatenation unit outputs a concatenated feature map, which is divided into two paths. One path of the feature map is directly sent to the fusion unit; the other path of the feature map is based on the self-attention mechanism. After obtaining the attention map, it is multiplied pointwise with itself to obtain the attention feature map; After the attention feature map passes through a 3×3 depthwise separable convolution, it is fused with the attention feature map, and then passes through a 1×1 convolution and is sent into the fusion unit; Fuse the feature maps in the fusion unit and output them as the downsampled feature map.

3. The manhole cover hidden danger detection method based on object detection according to claim 1, characterized in that, The backbone network includes a convolution unit, an efficient aggregation module and multiple extraction units connected in series in sequence.

4. The manhole cover hidden danger detection method based on object detection according to claim 1, characterized in that, Input the multiple primary feature maps of different scales into the neck network of the object detection model based on the FPN structure, and output the fused feature map corresponding to each scale, including: The primary feature map of the first scale is passed through the SP pooling module to obtain the pooled feature map of the first scale; after the pooled feature map of the first scale is processed by the dynamic upsampling module, it is concatenated with the primary feature map of the second scale to output the upsampled concatenated feature map of the second scale; The upsampled concatenated feature map of the k-th scale is passed through the SP pooling module to obtain the pooled feature map of the k-th scale; after the pooled feature map of the k-th scale is processed by the dynamic upsampling module, it is concatenated with the primary feature map of the (k + 1)-th scale to output the upsampled concatenated feature map of the k-th scale; k = 1, 2, …, K - 1; The upsampled concatenated feature map of the K-th scale is passed through the efficient aggregation module to output the fused feature map of the K-th scale; The fused feature map of the K-th scale is processed by the inverted residual attention downsampling module and then concatenated with the pooled feature map of the (K - 1)-th layer to obtain the downsampled concatenated feature map of the (K - 1)-th layer, which is then processed by the efficient aggregation module to output the fused feature map of the (K - 1)-th scale; the fused feature maps of the (K - 2)-th scale to the first scale are obtained in sequence.

5. The manhole cover hidden danger detection method based on object detection according to claim 1, characterized in that, The efficient aggregation module is a generalized efficient layer aggregation network, specifically used for: After convolving the input feature map of the efficient aggregation module, it is evenly split into two paths according to the number of channels, and both paths of feature maps are directly sent into the concatenation unit of the efficient aggregation module; Arbitrarily take one path of the feature map, which is successively passed through RepNCSP and convolution, and then split into two paths. One path of the feature map is directly sent into the concatenation unit of the efficient aggregation module, and the other path of the feature map is passed through RepNCSP and convolution and then sent into the concatenation unit of the efficient aggregation module; After the four paths of feature maps in the concatenation unit of the efficient aggregation module are concatenated, convolution is performed to obtain the output of the efficient aggregation module; Among them, the RepNCSP is used for: splitting the input feature map of RepNCSP into two paths, one path is passed through convolution and then sent into the concatenation unit in RepNCSP; the other path is passed through convolution and then split into two paths again, convolved at different scales and added, and then convolved to obtain the convolutional feature map, which is sent into the concatenation unit in RepNCSP; after the two paths of feature maps in the concatenation unit in RepNCSP are concatenated, convolution output is performed to obtain the output of RepNCSP.

6. The manhole cover hidden danger detection method based on object detection according to claim 1, wherein, The object detection model is obtained by replacing the upsampling module and the downsampling module in the YOLOv9 model with the dynamic upsampling module and the inverted residual attention downsampling module respectively.

7. The manhole cover hidden danger detection method based on object detection according to claim 6, characterized in that The training process of the object detection model includes: Obtaining manhole cover images corresponding to different manhole cover hazard categories as training images, and obtaining the true labels and true prediction boxes corresponding to the training images; Inputting the training images into the programmable gradient information module and the backbone network of the object detection model; Based on the feature maps of multiple different scales output by the backbone network and the feature maps passing through the efficient aggregation module and the inverted residual attention downsampling module in the programmable gradient information module, cross-branch fusion is performed to output the predicted labels, predicted bounding boxes, and confidence levels corresponding to the feature maps of each scale; Based on the predicted labels and the true labels, calculating the cross-entropy classification loss of the training images; Using the IoU loss, calculate the regression loss of the training image based on the overlap degree between the ground truth bounding box and the predicted bounding box; Construct the total loss function based on the weighted sum of the cross-entropy classification loss and the regression loss; Train the object detection model, perform gradient backpropagation based on the value of the total loss function, and let the optimizer update the model parameters until the preset number of iterations is reached to obtain the pre-trained object detection model.

8. The manhole cover hidden danger detection method based on object detection according to claim 7, characterized in that The total loss function is expressed as: ; Among them, represents the total loss function; represents the classification loss, and the expression is , represents the true label, represents the predicted probability of the class ; represents the regression loss, and the expression is , represents the predicted bounding box, represents the true bounding box, represents the intersection of the predicted bounding box and the true bounding box and the union ratio, and the expression is ; and are loss weight hyperparameters.

9. The manhole cover hidden danger detection method based on object detection according to claim 8, characterized in that, The programmable gradient information module is specifically used for: Input the multiple primary feature maps of different scales output by the backbone network into the cross-branch fusion modules of corresponding scales in the programmable gradient information module respectively; Let the training image pass through the convolution, efficient aggregation module and inverted residual attention downsampling module, and then input it into the cross-branch fusion module of the Kth scale; Initialize k = K, update k = k - 1 until k = 1; output the aggregated feature map of the kth scale after passing the output of the cross-branch fusion module of the kth scale through the efficient aggregation module; input the aggregated feature map of the kth scale into the cross-branch fusion module of the (k - 1)th scale after passing through the inverted residual attention downsampling module; output the class label, bounding box and confidence of the kth scale after passing the aggregated feature map of the kth scale through the detection head; Obtain the class label, bounding box and confidence of each scale as the output of the programmable gradient information module.

Citation Information

Patent Citations

  • Yolov5 target detection method based on cross-stage routing attention module and residual information fusion module

    CN116721398A

  • Improved YOLOv4-based power grid infrastructure target detection method

    CN117649514A