Well lid hidden danger detection method based on target detection
By introducing dynamic upsampling module and inverting residual attention downsampling module in manhole cover hazard detection, the problem of low detection accuracy caused by complex environmental background and small targets is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510450797.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In the detection of hidden dangers of manhole covers, the prior art causes low detection accuracy due to the complex environmental background and small targets.
The method based on object detection is adopted, and the feature extraction ability and environmental adaptability of the model are enhanced through the dynamic upsampling module and the inverted residual attention downsampling module.
It improves the accuracy and robustness of manhole cover hidden danger detection, and can better capture features in small targets and complex backgrounds to meet the needs of real-time monitoring.
Smart Images

Figure CN119992073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and target detection, and in particular to a method for detecting hidden dangers of manhole covers based on target detection. Background Art
[0002] With the acceleration of urbanization, the safety and reliability of urban infrastructure are increasingly valued. As an important part of urban roads and pipeline systems, manhole covers bear an important responsibility in urban infrastructure management to ensure the safe passage of pedestrians and vehicles. Their potential safety hazards are directly related to public safety and the image of the city. However, for a long time, due to various factors, including natural wear and tear, vehicle pressure and improper use, manhole covers may have various potential safety hazards, including manhole cover damage, uncovered manhole covers, lost manhole covers and manhole ring problems.
[0003] Traditional methods of detecting hidden dangers in manhole covers mainly rely on manual inspections, which are inefficient and costly, and are difficult to adapt to the detection needs in large-scale and complex environments. Therefore, using intelligent means to improve the efficiency and accuracy of manhole cover hidden danger detection has become an urgent problem to be solved in urban management.
[0004] In recent years, deep learning technology has made remarkable progress in the field of image recognition and processing, especially in object detection tasks. Deep learning models, such as convolutional neural networks (CNN), achieve efficient recognition and classification of objects in images by automatically learning image features. Object detection algorithms based on deep learning, such as Faster R-CNN, YOLO, and SSD, have been widely used in many fields due to their efficient detection speed and excellent accuracy.
[0005] However, in the field of intelligent detection of hidden dangers in manhole covers, existing technologies still face some challenges and problems. First, the diversity of shapes, materials and environmental backgrounds of manhole covers, as well as the complexity of urban environments, such as changes in lighting and weather conditions, place high demands on the adaptability and accuracy of intelligent detection systems. Secondly, existing target detection algorithms still have the problem of low detection accuracy when dealing with small targets, occluded targets, and targets in complex backgrounds. In addition, the need to process large amounts of image data in real time poses a challenge to the algorithm's computational efficiency. Summary of the invention
[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problem of low model detection accuracy in the prior art due to the complex background of the environment in which the manhole cover is located and the small target.
[0007] In order to solve the above technical problems, the present invention provides a method for detecting hidden dangers of manhole covers based on target detection, comprising: The manhole cover image collected in real time is input into the backbone network of the target detection model, and multiple primary feature maps of different scales are extracted through multiple extraction units connected in series. Each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module connected in series. Input multiple primary feature maps of different scales into the neck network based on the FPN structure of the target detection model, and output the fused feature map corresponding to each scale; the top-down path in the neck network includes a dynamic upsampling module; the bottom-up path in the neck network includes an inverted residual attention downsampling module; The fused feature map corresponding to each scale is input into the corresponding detection head in the head network of the target detection model, and the predicted label, predicted bounding box and confidence of the fused feature map of each scale corresponding to the real-time collected manhole cover image are output; the predicted label is the category of manhole cover hidden danger, including intact manhole cover, damaged manhole cover, uncovered manhole cover, manhole ring problem and lost manhole cover; The inverted residual attention downsampling module performs average pooling on the input feature map and then divides it into two paths according to the number of channels. One path is convolved, and the other path is max pooled and convolved. The two paths are spliced by channel to obtain a spliced feature map. The attention feature map of the spliced feature map is extracted, and after depth-separable convolution, it is fused with the attention feature map and the spliced feature map to output a downsampled feature map. The dynamic upsampling module divides the input feature map into two paths, one path undergoes convolution, and the other path undergoes convolution and activation, and then multiplies the two paths to obtain a combined feature map; shuffles pixels of the combined feature map to obtain a sampling offset; adds the sampling offset to the original sampling grid to obtain a new sampling grid, samples the input feature map, and outputs an upsampled feature map.
[0008] Preferably, the inverted residual attention downsampling module is specifically used for: After average pooling, the input feature map of the inverted residual attention downsampling module is evenly divided into two channel feature maps according to the number of channels; One channel feature map is sent to the splicing unit after 3×3 convolution; the other channel feature map is sent to the splicing unit after the maximum pooling layer and 1×1 convolution; The concatenation unit outputs a concatenated feature map, which is divided into two paths. One path is directly sent to the fusion unit; the other path is based on the self-attention mechanism. After obtaining the attention map, it is dot-multiplied with itself to obtain the attention feature map. The attention feature map is subjected to a 3×3 depth-separable convolution, fused with the attention feature map, and then subjected to a 1×1 convolution before being sent to the fusion unit. The feature maps in the fusion unit are fused and output as the downsampled feature map.
[0009] Preferably, the backbone network comprises a convolution unit, an efficient aggregation module and a plurality of extraction units connected in series.
[0010] Preferably, multiple primary feature maps of different scales are input into the neck network based on the FPN structure of the target detection model, and a fusion feature map corresponding to each scale is output, including: The primary feature map of the first scale is passed through the SP pooling module to obtain the pooled feature map of the first scale; the pooled feature map of the first scale is processed by the dynamic upsampling module and then spliced with the primary feature map of the second scale to output the upsampled spliced feature map of the second scale; The upsampled spliced feature map of the kth scale is passed through the SP pooling module to obtain the pooled feature map of the kth scale; the pooled feature map of the kth scale is processed by the dynamic upsampling module and then spliced with the primary feature map of the k+1th scale to output the upsampled spliced feature map of the kth scale; k=1, 2, ..., K-1; The upsampled spliced feature map of the Kth scale is passed through an efficient aggregation module to output a fused feature map of the Kth scale; The fused feature map of the Kth scale is processed by the inverted residual attention downsampling module and then spliced with the K-1th layer pooling feature map. After obtaining the K-1th layer downsampling splicing feature map, it is processed by the efficient aggregation module to output the K-1th scale fused feature map; the fused feature maps of the K-2th scale to the 1st scale are obtained in sequence.
[0011] Preferably, the efficient aggregation module is a generalized efficient layer aggregation network, specifically used for: After convolution of the input feature map of the efficient aggregation module, it is evenly split into two paths according to the number of channels, and both feature maps are directly sent to the splicing unit of the efficient aggregation module; After randomly selecting one of the feature maps and passing through RepNCSP and convolution in sequence, it is divided into two paths. One of the feature maps is directly sent to the splicing unit of the efficient aggregation module, and the other feature map is sent to the splicing unit of the efficient aggregation module after passing through RepNCSP and convolution. After splicing the four feature maps in the splicing unit of the efficient aggregation module, the output of the efficient aggregation module is obtained through convolution. The RepNCSP is used to: divide the input feature map of the RepNCSP into two paths, one of which is sent to the splicing unit in the RepNCSP after convolution; the other path is divided into two paths again after convolution, and convolutions of different scales are performed and added, and then convolution is performed to obtain a convolution feature map, which is sent to the splicing unit in the RepNCSP; after splicing the two feature maps in the splicing unit in the RepNCSP, the output is convolved to obtain the output of the RepNCSP.
[0012] Preferably, the target detection model is obtained by replacing the upsampling module and the downsampling module in the YOLOv9 model with the dynamic upsampling module and the inverted residual attention downsampling module respectively.
[0013] Preferably, the training process of the target detection model includes: Obtain manhole cover images corresponding to different manhole cover hazard categories as training images, and obtain the true labels and true prediction boxes corresponding to the training images; Input the training image into the programmable gradient information module and backbone network of the target detection model; Based on the feature maps of multiple different scales output by the backbone network, cross-branch fusion is performed with the feature maps of the programmable gradient information module that have passed through the efficient aggregation module and the inverted residual attention downsampling module, and the predicted label, predicted bounding box and confidence corresponding to the feature map of each scale are output; Calculate the cross entropy classification loss of the training image based on the predicted label and the true label; Using IoU loss, the regression loss of the training image is calculated based on the overlap between the true bounding box and the predicted bounding box. Construct a total loss function based on the weighted sum of the cross entropy classification loss and the regression loss; The target detection model is trained, and the gradient is returned based on the value of the total loss function. The optimizer updates the model parameters until the preset number of iterations is reached to obtain the pre-trained target detection model.
[0014] Preferably, the total loss function is expressed as: ; in, represents the total loss function; Represents the classification loss, expressed as , represents the true label, Indicates the category The predicted probability of Represents the regression loss, expressed as , represents the predicted bounding box, represents the ground-truth bounding box, Represents the intersection of the predicted bounding box and the true bounding box and union The ratio of is expressed as ; and is the loss weight hyperparameter.
[0015] Preferably, the programmable gradient information module is specifically used for: The primary feature maps of multiple different scales output by the backbone network are respectively input into the cross-branch fusion modules of corresponding scales in the programmable gradient information module; After the training image passes through the convolution, efficient aggregation module and inverted residual attention downsampling module, it is input into the cross-branch fusion module of the Kth scale; Initialize k=K, update k=k-1 until k=1; pass the output of the cross-branch fusion module of the kth scale through the efficient aggregation module to output the aggregated feature map of the kth scale; pass the aggregated feature map of the kth scale through the inverted residual attention downsampling module and input it into the cross-branch fusion module of the k-1th scale; pass the aggregated feature map of the kth scale through the detection head output to obtain the category label, bounding box and confidence of the kth scale; Get the class labels, bounding boxes, and confidences at each scale as the output of the programmable gradient information module.
[0016] The above technical solution of the present invention has the following beneficial effects compared with the prior art: The method for detecting hidden dangers of manhole covers based on target detection described in the present invention proposes a dynamic upsampling module and an inverted residual attention downsampling module.
[0017] The dynamic upsampling module defines upsampling as a point sampling problem. By generating a sampling offset to adjust the sampling network, it can process the features corresponding to the manhole cover image more flexibly, allowing the model to better capture the information of small targets or fine structures in the manhole cover image, and improve the detection and segmentation accuracy of the details of the manhole cover hidden dangers, thereby improving the accuracy of hidden danger detection for manhole covers of different shapes and materials in different environments. In addition, the dynamic upsampling module still has fewer parameters, floating-point operations, GPU memory and latency while achieving stronger performance, has higher timeliness, and is more suitable for real-time monitoring of road manhole covers.
[0018] After the channel-wise convolution, the inverted residual attention downsampling module introduces the attention mechanism, depthwise separable convolution and residual mechanism, and extends the convolution-based inverted residual block to the attention-based model. The model combines the static short-range modeling capability of CNN and the dynamic long-range feature interaction capability of Transformer. It can not only effectively compress data, but also enhance the focus on key features while reducing information loss, thereby improving the model's feature expression and detection capabilities for the target, and then accurately extracting the features corresponding to the target in the manhole cover image, thereby improving the accuracy of hidden danger detection in the manhole cover image.
[0019] The present invention is based on the YOLOv9 model, introduces a dynamic upsampling module and an inverted residual attention downsampling module, enhances the feature extraction capability and environmental adaptability of the YOLOv9 model, improves the accuracy and robustness of the model in detecting the hidden danger state of manhole covers in complex environments, and still has a high accuracy when processing small targets, occluded targets, and targets under complex backgrounds; at the same time, it realizes efficient processing of a large amount of image data, meeting the needs of real-time monitoring. In addition, when training the target detection model based on the YOLOv9 model, a programmable gradient information module is used to generate reliable gradients through auxiliary reversible branches, and gradient information propagation is programmed at different semantic levels, so that deep features still maintain the key characteristics of executing target tasks, avoiding semantic loss that may be caused by traditional deep supervision processes, improving the extraction accuracy of image features corresponding to hidden dangers in manhole covers, improving the model training accuracy, and thus improving the detection accuracy of hidden dangers in manhole covers. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein: Figure 1 It is a flowchart of the steps of the method for detecting hidden dangers of manhole covers based on target detection provided by the present invention; Figure 2 This is a schematic diagram of hidden dangers of manhole covers; Figure 3 Schematic diagram of the inverted residual attention downsampling module; Figure 4 It is a schematic diagram of the dynamic upsampling module; Figure 5 It is a schematic diagram of the efficient aggregation module; Figure 6 It is a schematic diagram of the SP pooling module; Figure 7 Schematic diagram of the target detection model based on the improved YOLOv9 model; Figure 8 It is a schematic diagram of the cross-branch fusion module; Fig. 9 is a user interface schematic diagram; Fig.10 This is a schematic diagram of the user interface detection results; Fig.11 Select a schematic for the user interface input; Fig.12 This is a schematic diagram of the intactness test results of the manhole cover; Fig.13 This is a schematic diagram of the detection result of the manhole cover not being covered; Fig.14 This is a schematic diagram of the manhole cover damage detection results; Fig.15This is a schematic diagram of the well circle problem detection results; Fig.16 Schematic diagram of the manhole cover missing detection results. DETAILED DESCRIPTION
[0021] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0022] Reference Figure 1 As shown, the step flow chart of the manhole cover hidden danger detection method based on target detection provided by the present invention specifically includes: S101: inputting the manhole cover image collected in real time into the backbone network of the target detection model, and extracting multiple primary feature maps of different scales through multiple extraction units connected in series; each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module connected in series; S102: inputting a plurality of primary feature maps of different scales into a neck network based on an FPN structure of a target detection model, and outputting a fused feature map corresponding to each scale; a top-down path in the neck network includes a dynamic upsampling module; a bottom-up path in the neck network includes an inverted residual attention downsampling module; Among them, the Feature Pyramid Network (FPN) structure includes a top-down upsampling path and a bottom-up downsampling path, and the upsampling path and the downsampling path are horizontally connected; the top-down path uses a dynamic upsampling module to upsample the feature map, and the bottom-up path uses an inverted residual attention downsampling module to downsample the feature map.
[0023] S103: Input the fused feature map corresponding to each scale into the corresponding detection head in the head network of the target detection model, and output the predicted label, predicted bounding box and confidence of the fused feature map of each scale corresponding to the manhole cover image collected in real time; the predicted label is the category of hidden dangers of the manhole cover, including intact manhole cover, damaged manhole cover, uncovered manhole cover, manhole ring problem and lost manhole cover; The inverted residual attention downsampling module performs average pooling on the input feature map and then divides it into two paths according to the number of channels. One path is convolved, and the other path is max pooled and convolved. The two paths are spliced by channel to obtain a spliced feature map. The attention feature map of the spliced feature map is extracted, and after depth-separable convolution, it is fused with the attention feature map and the spliced feature map to output a downsampled feature map. The dynamic upsampling module divides the input feature map into two paths, one path undergoes convolution, and the other path undergoes convolution and activation, and then multiplies the two paths to obtain a combined feature map; shuffles pixels of the combined feature map to obtain a sampling offset; adds the sampling offset to the original sampling grid to obtain a new sampling grid, samples the input feature map, and outputs an upsampled feature map.
[0024] Reference Figure 2 As shown in FIG. 1 , a schematic diagram of hidden dangers of manhole covers is shown; specifically, the predicted labels are categories of hidden dangers of manhole covers, including good manhole covers, broken manhole covers, uncovered manhole covers, circle manhole covers, and lost manhole covers.
[0025] Reference Figure 3 As shown, it is a schematic diagram of the inverted residual attention downsampling module; the inverted residual attention downsampling module is specifically used for: After average pooling, the input feature map of the inverted residual attention downsampling module is evenly divided into two channel feature maps according to the number of channels; One channel feature map is sent to the splicing unit after 3×3 convolution; the other channel feature map is sent to the splicing unit after the maximum pooling layer and 1×1 convolution; The concatenation unit outputs a concatenated feature map, which is divided into two paths. One path is directly sent to the fusion unit; the other path is based on the self-attention mechanism. After obtaining the attention map, it is dot-multiplied with itself to obtain the attention feature map. The attention feature map is subjected to a 3×3 depth-separable convolution, fused with the attention feature map, and then subjected to a 1×1 convolution before being sent to the fusion unit. The feature maps in the fusion unit are fused and output as the downsampled feature map.
[0026] The Attention-based Inverted Residual Downsampling (AIRD) module connects an inverted residual moving module to a channel-wise convolution module, extending the convolution-based inverted residual block to an attention-based model, so that the model combines the static short-range modeling capability of CNN and the dynamic long-range feature interaction capability of Transformer. Among them, the 3*3 convolution is a regular convolution with a step size of 2 and the output channel remains unchanged. Therefore, the size of the feature map will be halved after passing through it, and the number of channels remains unchanged. It can be understood as downsampling the feature map; the 1*1 convolution on the right has a step size of 1 and the output channel remains unchanged. Therefore, the number of channels and size of the feature map remain unchanged after passing through it. A simple 1*1 convolution operation is performed; the first 1*1 convolution after the splicing feature does not change the size, but adjusts the number of channels to generate the V value in the attention operation. The general attention operation generates QKV values. It is obtained through the linear layer, where 1*1 convolution is used as the linear layer; 3*3 depth-wise separable convolution is a special convolution kernel. Different from ordinary convolution, each depth-wise separable convolution only performs convolution calculation on a single channel (normal convolution will calculate all channels of the feature map), which is relatively lightweight and requires less calculation than ordinary convolution. The 3*3 depth-wise separable convolution here does not change the number of channels and dimensions; the function of the bottom 1*1 convolution is to adjust the number of channels so that it can be the same as the number of splicing feature channels, so as to perform the addition operation and realize the residual connection.
[0027] Reference Figure 4 The figure shows a schematic diagram of the dynamic upsampling module. Dynamic upsampling (DyUp) is different from traditional dynamic convolution. Dynamic upsampling adopts a more resource-efficient upsampling method from the perspective of point sampling. It bypasses dynamic convolution and redefines the upsampling problem as a point sampling problem, strengthening the behavior of traditional upsampling. Dynamic upsampling enables the model to have fewer parameters, floating-point operations, GPU memory, and latency while achieving stronger performance. Dynamic upsampling not only has superior performance in the field of object detection, but also has strong performance in semantic segmentation and instance segmentation.
[0028] Specifically, the backbone network includes a convolution unit, an efficient aggregation module, and a plurality of extraction units connected in series. The backbone network is used to obtain a plurality of primary feature maps of different scales, including: inputting the manhole cover image collected in real time into the backbone network, and passing through the convolution unit, the efficient aggregation module, and a plurality of extraction units in sequence; each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module connected in series; and the output of the efficient aggregation module in each extraction unit is used as the primary feature map of the corresponding scale.
[0029] Specifically, multiple primary feature maps of different scales are input into the neck network based on the FPN structure of the target detection model, and the fused feature map corresponding to each scale is output, including: The primary feature map of the first scale is passed through the SP pooling module to obtain the pooled feature map of the first scale; the pooled feature map of the first scale is processed by the dynamic upsampling module and then spliced with the primary feature map of the second scale to output the upsampled spliced feature map of the second scale; The upsampled spliced feature map of the kth scale is passed through the SP pooling module to obtain the pooled feature map of the kth scale; the pooled feature map of the kth scale is processed by the dynamic upsampling module and then spliced with the primary feature map of the k+1th scale to output the upsampled spliced feature map of the kth scale; k=1, 2, ..., K-1; The upsampled spliced feature map of the Kth scale is passed through an efficient aggregation module to output a fused feature map of the Kth scale; The fused feature map of the Kth scale is processed by the inverted residual attention downsampling module and then spliced with the K-1th layer pooling feature map. After obtaining the K-1th layer downsampling splicing feature map, it is processed by the efficient aggregation module to output the K-1th scale fused feature map; the fused feature maps of the K-2th scale to the 1st scale are obtained in sequence.
[0030] Reference Figure 5 The figure shows a schematic diagram of an efficient aggregation module; the efficient aggregation module is a generalized efficient layer aggregation network (Generalized Efficient Layer Aggregation Network), which is specifically used for: After convolution of the input feature map of the efficient aggregation module, it is evenly split into two paths according to the number of channels, and both feature maps are directly sent to the splicing unit of the efficient aggregation module; After randomly selecting one of the feature maps and passing through RepNCSP and convolution in sequence, it is divided into two paths. One of the feature maps is directly sent to the splicing unit of the efficient aggregation module, and the other feature map is sent to the splicing unit of the efficient aggregation module after passing through RepNCSP and convolution. After splicing the four feature maps in the splicing unit of the efficient aggregation module, the output of the efficient aggregation module is obtained through convolution. Among them, the RepNCSP combines the reparameterization technology (Reparameterization) with the CSP (Cross Stage Partial) structure, which not only retains the efficient feature fusion capability of CSP, but also improves the inference speed through the reparameterization technology, and is used to: divide the input feature map of RepNCSP into two paths, one of which is sent to the splicing unit in RepNCSP after convolution; the other path is divided into two paths again after convolution, and convolutions of different scales are performed and then added, and convolution is performed to obtain a convolution feature map, which is sent to the splicing unit in RepNCSP; after splicing the two feature maps in the splicing unit in RepNCSP, the output is convolved to obtain the output of RepNCSP.
[0031] Reference Figure 6 , which is a schematic diagram of the SP pooling module; the SP pooling module is based on Spatial Pyramid Pooling (SPP); the SP pooling module is Spatial Pyramid Pooling Enhanced with ELAN, which is used for: The input features of the SP pooling module are sequentially passed through the convolution unit, the first maximum pooling unit, the second maximum pooling unit, and the third maximum pooling unit; After concatenating the output of each unit, it passes through the convolution unit output to obtain the SP pooled output.
[0032] The method for detecting hidden dangers of manhole covers based on target detection described in the present invention proposes a dynamic upsampling module and an inverted residual attention downsampling module. The dynamic upsampling module defines upsampling as a point sampling problem, and adjusts the sampling network by generating a sampling offset, which can more flexibly process the features corresponding to the manhole cover image, allowing the model to better capture the information of small targets or fine structures in the manhole cover image, improve the detection and segmentation accuracy of the details in the hidden dangers of manhole covers, and thus improve the accuracy of hidden danger detection for manhole covers of different shapes and materials under different environments; and the dynamic upsampling module still has fewer parameters, floating-point operations, GPU memory and latency while achieving stronger performance, has higher timeliness, and is more suitable for real-time monitoring of road manhole covers. After the channel-wise convolution, the inverted residual attention downsampling module introduces the attention mechanism, depthwise separable convolution and residual mechanism, and extends the convolution-based inverted residual block to the attention-based model. The model combines the static short-range modeling capability of CNN and the dynamic long-range feature interaction capability of Transformer. It can not only effectively compress data, but also enhance the focus on key features while reducing information loss, thereby improving the model's feature expression and detection capabilities for the target, and then accurately extracting the features corresponding to the target in the manhole cover image, thereby improving the accuracy of hidden danger detection in the manhole cover image.
[0033] Based on the above embodiment, in an embodiment of the present invention, the upsampling module and the downsampling module in the YOLOv9 model are respectively replaced with the dynamic upsampling module and the inverted residual attention downsampling module to obtain a target detection model to detect hidden dangers of manhole covers; wherein the training process of the target detection model based on the improved YOLOv9 model includes: S201: Obtain manhole cover images corresponding to different manhole cover hazard categories as training images, and obtain true labels and true prediction boxes corresponding to the training images; S202: Inputting the training image into the programmable gradient information module and the backbone network of the target detection model; S203: Based on the feature maps of multiple different scales output by the backbone network, cross-branch fusion is performed with the feature maps of the programmable gradient information module that have passed through the efficient aggregation module and the inverted residual attention downsampling module, and the predicted label, predicted bounding box and confidence corresponding to the feature map of each scale are output; S204: Calculate the cross entropy classification loss of the training image based on the predicted label and the true label; S205: Using the IoU loss, the regression loss of the training image is calculated based on the overlap between the real bounding box and the predicted bounding box; S206: Based on the weighted sum of the cross entropy classification loss and the regression loss, a total loss function is constructed, which is expressed as: ; in, represents the total loss function; Represents the classification loss, expressed as , represents the true label, Indicates the category The predicted probability of Represents the regression loss, expressed as , represents the predicted bounding box, represents the ground-truth bounding box, Represents the intersection of the predicted bounding box and the true bounding box and union The ratio of is expressed as ; and is the loss weight hyperparameter.
[0034] S207: Train the target detection model, perform gradient feedback based on the value of the total loss function, and let the optimizer update the model parameters until a preset number of iterations is reached to obtain a pre-trained target detection model.
[0035] Reference Figure 7 As shown in the figure, it is a schematic diagram of the target detection model based on the improved YOLOv9 model; the model consists of four parts: programmable gradient information module, backbone network, neck network and head network.
[0036] The programmable gradient information module is specifically used for: The primary feature maps of multiple different scales output by the backbone network are respectively input into the cross-branch fusion module of the corresponding scale in the programmable gradient information module; After the training image passes through the convolution, efficient aggregation module and inverted residual attention downsampling module, it is input into the cross-branch fusion module of the Kth scale; Initialize k=K, update k=k-1 until k=1; pass the output of the cross-branch fusion module of the kth scale through the efficient aggregation module to output the aggregated feature map of the kth scale; pass the aggregated feature map of the kth scale through the inverted residual attention downsampling module and input it into the cross-branch fusion module of the k-1th scale; pass the aggregated feature map of the kth scale through the detection head output to obtain the category label, bounding box and confidence of the kth scale; Get the class labels, bounding boxes, and confidences at each scale as the output of the programmable gradient information module.
[0037] The Programmable Gradient Information (PGI) module generates reliable gradients through auxiliary reversible branches, so that deep features still maintain the key characteristics of performing the target task. The design of the auxiliary reversible branch can avoid the semantic loss that may occur in the traditional deep supervision process, which integrates multi-path features. In other words, PGI achieves the best training results by programming gradient information propagation at different semantic levels. PGI's reversible architecture is built on auxiliary branches, so there is no additional cost. Since PGI can freely choose the loss function suitable for the target task, it is more general than the deep supervision mechanism, which is only applicable to very deep neural networks. During training, the programmable gradient information module will pass information to the main module, and when inference, there is no need to go through the programmable gradient information module.
[0038] Reference Figure 8 As shown, it is a schematic diagram of a cross-branch fusion module; the cross-branch fusion module includes: calculating the nearest neighbor difference values of multiple input features, and adding them to obtain the output of the cross-branch fusion module.
[0039] This embodiment adopts a dynamic upsample module (DyUp), a programmable gradient information module (PGI) and an attention-based inverted residual downsampling module (AIRD), combined with YOLOv9, to enhance the feature extraction capability and environmental adaptability of the model, which not only improves the accuracy of intelligent detection of hidden dangers in manhole covers, but also improves the robustness of the model under different lighting and weather conditions.
[0040] Specifically, this embodiment uses labeled training data to adjust model parameters so that it can learn appropriate feature representations from the data, thereby learning knowledge and rules for solving the hidden dangers of manhole covers; for N images from different categories in a batch B of any input target detection model, the main process of model training includes: ① Extraction of basic image features; The basic features of each image are obtained through the innovative backbone network of YOLOv9, namely the Generalized Efficient Layer Aggregation Network. Feature maps of multiple scales are generated through a series of upsampling and downsampling operations. The operations usually include pooling layers, convolution layers, and interpolation layers. This embodiment uses two innovative upsampling and downsampling function modules, namely the dynamic upsampling module and the inverted residual self-attention module, which can extract semantic information more effectively. The feature map of each scale corresponds to different resolutions of the input image. The fusion of feature maps of different scales can enhance the model's detection ability for objects of different sizes; the fusion method can be weighted summation, splicing, or a more complex attention mechanism. In this way, low-level feature maps with high resolution but less semantic information and high-level feature maps with low resolution but rich semantic information can complement each other and improve detection accuracy. In the end, three feature maps of different sizes are obtained, and the scales are arranged from large to small, as follows: , , .
[0041] ② Classification and regression are performed based on the extracted basic features; The feature map is input into the fully connected layer for classification, and the cross entropy classification loss is expressed as: ; in, is the model for the category The predicted probability of is the true label.
[0042] For regression tasks, this embodiment uses IoU (Intersection over Union) loss to measure the overlap between the predicted bounding box and the true bounding box. IoU is the ratio of the intersection to the union of two bounding boxes. Assume is the predicted bounding box, is the true bounding box, and the IoU loss is expressed as: ; Therefore, the regression loss is expressed as: ; In actual training, the total loss is the weighted sum of classification loss and regression loss, expressed as: ; This total loss function only considers the case of a single-scale feature map. In actual calculation, considering that the model will calculate feature maps of different scales, the classification loss and regression loss corresponding to the feature map of each scale are calculated separately to form a total loss function, which is expressed as: ; in, and They are The classification loss and regression loss of feature maps at each scale.
[0043] ③ According to the total loss function, use the optimizer to optimize the model parameters so that the model can accurately detect various manhole cover hidden dangers.
[0044] Based on the above description, the model training process includes: Input: Dataset , For images, is the corresponding label, which contains the bounding box and categories ;Total number of training rounds required ; Model structure: backbone network (feature extractor )、Neck network (feature fusion ), head network (model prediction head ), auxiliary branch programmable gradient information module; Training process: ① Initialize the model model, initialize the optimizer optimizer, initialize the total loss function L, and define the data loader train_loader; ②Training cycle: for current training round t=1 to MAX_T (preset total training rounds) do: model is set to training mode; for the current number of iterations of the data set i=1 to MAX_I (the number of iterations required to iterate the data set) do: Clear the gradient retained by the optimizer; Take a sample , and get the category label of the sample With the ground-truth bounding box ; Get the multi-scale features of the sample through the backbone of the model ; Get the fused features through the neck of the model ; The final prediction results are obtained through the head of the model, including: predicting the object bounding box , Confidence and category probability distribution ; Calculate the classification loss, regression loss, and total loss, perform gradient backpropagation based on the total loss, and let the optimizer update the model parameters; End for; End for; ③Save the trained model model.
[0045] The embodiment of the present invention uses the target detection-based manhole cover hidden danger detection method provided by the present invention to perform detection based on the 2024 Fu Chuang Competition A03 manhole cover dataset; the 2024 Fu Chuang Competition A03 manhole cover dataset is a dataset specifically used for intelligent detection of manhole cover hidden dangers, containing more than 1,300 manhole cover images in various states, such as intact, damaged, missing, uncovered, and manhole ring problems; this dataset aims to achieve automatic detection of manhole cover hidden dangers through target detection technology to improve urban management efficiency and public safety. In the training phase, the stochastic gradient descent method is used with a decay rate of , momentum is 0.9; the model is trained for 100 rounds, and the batch size is 16. In order to enhance the training data, this embodiment uses the following data enhancement method: randomly select images, directly resize to 256×256, and then crop to 224×224; horizontally flip.
[0046] As shown in Table 1, the effectiveness of the method proposed in the present invention is verified in this embodiment based on YOLOv9; the prediction results are evaluated on the A03 manhole cover data test set of the 2024 Service Innovation Competition.
[0047] Table 1 Comparison of prediction results of YOLOv9 models with different functional modules deployed
[0048] Referring to Table 1, it can be seen that the mAP value of the manhole cover hidden danger detection method based on target detection provided by the present invention can reach 0.78, which can effectively detect the manhole cover hidden danger problem.
[0049] The YOLO model has been iterating rapidly. This embodiment compares the target detection model based on the YOLOv9 model with the improved target detection model based on the YOLOv5 model: when both AR downsampling and Dy upsampling are used, the map50 obtained by using YOLOv9-c as the benchmark model is 0.779, and when using YOLOv5-1, it is 0.753. Based on the YOLOv9 model, the present invention introduces a dynamic upsampling module and an inverted residual attention downsampling module, enhances the feature extraction capability and environmental adaptability of the YOLOv9 model, improves the accuracy and robustness of the model in detecting the hidden danger status of manhole covers in complex environments, and realizes efficient processing of a large amount of image data, meeting the needs of real-time monitoring. During training, the target detection model based on the YOLOv9 model uses a programmable gradient information module to generate reliable gradients through auxiliary reversible branches, and programs gradient information propagation at different semantic levels, so that deep features still maintain the key characteristics of performing target tasks, avoiding semantic loss that may be caused by traditional deep supervision processes, and improving model training accuracy. The network can better adapt to the more difficult data distribution of manhole cover problems. At the same time, it can also extract richer and more accurate local features, effectively reducing false detections.
[0050] Based on the above embodiment, in the embodiment of the present invention, when the model is deployed after training, this embodiment uses pytq5 to develop the application interface. During inference, the input image is first preprocessed in accordance with the training, and then the pretrained model is loaded, the image is input into the loaded model and the result is predicted, and then the prediction result is fed back to the user end. Fig. 9 The figure shows a user interface diagram; the user interface contains a variety of data input formats, and the boxed parts in the figure are from top to bottom: select file detection from file, start camera detection, and detect from network video stream. Click the button to select the file. Fig.10 The following is a schematic diagram of the user interface detection results. When a file is selected for detection, the text results will be displayed on the left after the detection is completed, and the image detection results will be displayed on the right side of the interface. Fig.11 As shown, it is a schematic diagram of user interface input selection; the user interface provides parameter settings and detection parameter settings, and supports input parameters such as IoU intersection over union ratio, confidence, and inter-frame delay.
[0051] Based on the manhole cover hidden danger detection model provided by the present invention, a variety of manhole cover images with different hidden danger categories are detected, and the detection results are as follows: Figures 12 to 16 See Fig.12 The following is a schematic diagram of the results of the manhole cover integrity inspection; Fig.13 The following is a schematic diagram of the detection result of the manhole cover not being covered; Fig.14 The figure shows the results of the manhole cover damage detection. Fig.15The figure shows the results of the wellbore problem detection. Fig.16 The figure shows the result of manhole cover missing detection.
[0052] The target detection-based manhole cover hidden danger detection method provided by the present invention can be deployed as a corresponding detection system, and the main functions of the system include but are not limited to real-time monitoring and detection, abnormal detection and early warning, user interface and interaction. Through the deployment and application of this system, the efficiency and level of urban manhole cover safety management can be improved, accidents caused by manhole cover hidden dangers can be reduced, and a safer living environment can be provided for urban residents.
[0053] The method for detecting hidden dangers of manhole covers based on target detection described in the present invention proposes a dynamic upsampling module and an inverted residual attention downsampling module. The dynamic upsampling module defines upsampling as a point sampling problem, and adjusts the sampling network by generating a sampling offset, which can more flexibly process the features corresponding to the manhole cover image, allowing the model to better capture the information of small targets or fine structures in the manhole cover image, improve the detection and segmentation accuracy of the details in the hidden dangers of manhole covers, and thus improve the accuracy of hidden danger detection for manhole covers of different shapes and materials under different environments; and the dynamic upsampling module still has fewer parameters, floating-point operations, GPU memory and latency while achieving stronger performance, has higher timeliness, and is more suitable for real-time monitoring of road manhole covers. After the channel-wise convolution, the inverted residual attention downsampling module introduces the attention mechanism, the depth-separable convolution and the residual mechanism, and extends the convolution-based inverted residual block to the attention-based model, so that the model combines the static short-range modeling ability of CNN and the dynamic long-range feature interaction ability of Transformer, which can not only effectively compress data, but also enhance the focus on key features while reducing information loss, improve the model's feature expression and detection capabilities for targets, and then accurately extract the features corresponding to the targets in the manhole cover image, and improve the accuracy of hidden danger detection in the manhole cover image. The present invention is based on the YOLOv9 model, introduces a dynamic upsampling module and an inverted residual attention downsampling module, enhances the feature extraction capability and environmental adaptability of the YOLOv9 model, improves the accuracy and robustness of the model's detection of the hidden danger status of manhole covers in complex environments, and still has a high accuracy when processing small targets, occluded targets and targets under complex backgrounds; at the same time, it realizes efficient processing of a large amount of image data, meeting the needs of real-time monitoring. During training, the target detection model based on the YOLOv9 model uses a programmable gradient information module to generate reliable gradients through auxiliary reversible branches, and programs gradient information propagation at different semantic levels, so that deep features still maintain the key characteristics of executing target tasks, avoiding the semantic loss that may be caused by the traditional deep supervision process, and improving the extraction accuracy of image features corresponding to hidden dangers in manhole covers, thereby improving the model training accuracy and thus improving the accuracy of manhole cover hidden danger detection.
[0054] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0055] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0056] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0057] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0058] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. A method for detecting hidden dangers of manhole covers based on target detection, characterized in that: include: The manhole cover image collected in real time is input into the backbone network of the target detection model, and multiple primary feature maps of different scales are extracted through multiple extraction units connected in series. Each extraction unit includes an inverted residual attention downsampling module and an efficient aggregation module connected in series. Input multiple primary feature maps of different scales into the neck network based on the FPN structure of the target detection model, and output the fused feature map corresponding to each scale; the top-down path in the neck network includes a dynamic upsampling module; the bottom-up path in the neck network includes an inverted residual attention downsampling module; The fused feature map corresponding to each scale is input into the corresponding detection head in the head network of the target detection model, and the predicted label, predicted bounding box and confidence of the fused feature map of each scale corresponding to the real-time collected manhole cover image are output; the predicted label is the category of manhole cover hidden danger, including intact manhole cover, damaged manhole cover, uncovered manhole cover, manhole ring problem and lost manhole cover; The inverted residual attention downsampling module performs average pooling on the input feature map and then divides it into two paths according to the number of channels. One path is convolved, and the other path is max pooled and convolved. The two paths are spliced by channel to obtain a spliced feature map. The attention feature map of the spliced feature map is extracted, and after depth-separable convolution, it is fused with the attention feature map and the spliced feature map to output a downsampled feature map. The dynamic upsampling module divides the input feature map into two paths, one path undergoes convolution, and the other path undergoes convolution and activation, and then multiplies the two paths to obtain a combined feature map; shuffles the pixels of the combined feature map to obtain a sampling offset; Add the sampling offset to the original sampling grid to obtain a new sampling grid, sample the input feature map, and output the upsampled feature map.
2. The method for detecting hidden dangers of manhole covers based on target detection according to claim 1 is characterized in that: The inverted residual attention downsampling module is specifically used for: After average pooling, the input feature map of the inverted residual attention downsampling module is evenly divided into two channel feature maps according to the number of channels; One channel feature map is sent to the splicing unit after 3×3 convolution; the other channel feature map is sent to the splicing unit after the maximum pooling layer and 1×1 convolution; The concatenation unit outputs a concatenated feature map, which is divided into two paths. One path is directly sent to the fusion unit; the other path is based on the self-attention mechanism. After obtaining the attention map, it is dot-multiplied with itself to obtain the attention feature map. The attention feature map is subjected to a 3×3 depth-separable convolution, fused with the attention feature map, and then subjected to a 1×1 convolution before being sent to the fusion unit. The feature maps in the fusion unit are fused and output as the downsampled feature map.
3. The method for detecting hidden dangers of manhole covers based on target detection according to claim 1, characterized in that: The backbone network includes a convolution unit, an efficient aggregation module and a plurality of extraction units connected in series.
4. The method for detecting hidden dangers of manhole covers based on target detection according to claim 1, characterized in that: Multiple primary feature maps of different scales are input into the neck network based on the FPN structure of the target detection model, and the fused feature maps corresponding to each scale are output, including: The primary feature map of the first scale is passed through the SP pooling module to obtain the pooled feature map of the first scale; the pooled feature map of the first scale is processed by the dynamic upsampling module and then spliced with the primary feature map of the second scale to output the upsampled spliced feature map of the second scale; The upsampled spliced feature map of the kth scale is passed through the SP pooling module to obtain the pooled feature map of the kth scale; the pooled feature map of the kth scale is processed by the dynamic upsampling module and then spliced with the primary feature map of the k+1th scale to output the upsampled spliced feature map of the kth scale; k=1, 2, ..., K-1; The upsampled spliced feature map of the Kth scale is passed through an efficient aggregation module to output a fused feature map of the Kth scale; The fused feature map of the Kth scale is processed by the inverted residual attention downsampling module and then spliced with the K-1th layer pooling feature map. After obtaining the K-1th layer downsampling splicing feature map, it is processed by the efficient aggregation module to output the K-1th scale fused feature map; the fused feature maps of the K-2th scale to the 1st scale are obtained in sequence.
5. The method for detecting hidden dangers of manhole covers based on target detection according to claim 1, characterized in that: The high-efficiency aggregation module is a generalized high-efficiency layer aggregation network, specifically used for: After convolution of the input feature map of the efficient aggregation module, it is evenly split into two paths according to the number of channels, and both feature maps are directly sent to the splicing unit of the efficient aggregation module; After randomly selecting one of the feature maps and passing through RepNCSP and convolution in sequence, it is divided into two paths. One of the feature maps is directly sent to the splicing unit of the efficient aggregation module, and the other feature map is sent to the splicing unit of the efficient aggregation module after passing through RepNCSP and convolution. After splicing the four feature maps in the splicing unit of the efficient aggregation module, the output of the efficient aggregation module is obtained through convolution. The RepNCSP is used to: divide the input feature map of the RepNCSP into two paths, one of which is sent to the splicing unit in the RepNCSP after convolution; the other path is divided into two paths again after convolution, and convolutions of different scales are performed and added, and then convolution is performed to obtain a convolution feature map, which is sent to the splicing unit in the RepNCSP; after splicing the two feature maps in the splicing unit in the RepNCSP, the output is convolved to obtain the output of the RepNCSP.
6. The method for detecting hidden dangers of manhole covers based on target detection according to claim 1, characterized in that: The target detection model is obtained by replacing the upsampling module and the downsampling module in the YOLOv9 model with the dynamic upsampling module and the inverted residual attention downsampling module respectively.
7. The method for detecting hidden dangers of manhole covers based on target detection according to claim 6 is characterized in that: The training process of the target detection model includes: Obtain manhole cover images corresponding to different manhole cover hazard categories as training images, and obtain the true labels and true prediction boxes corresponding to the training images; Input the training image into the programmable gradient information module and backbone network of the target detection model; Based on the feature maps of multiple different scales output by the backbone network, cross-branch fusion is performed with the feature maps of the programmable gradient information module that have passed through the efficient aggregation module and the inverted residual attention downsampling module, and the predicted label, predicted bounding box and confidence corresponding to the feature map of each scale are output; Calculate the cross entropy classification loss of the training image based on the predicted label and the true label; Using IoU loss, the regression loss of the training image is calculated based on the overlap between the true bounding box and the predicted bounding box. Construct a total loss function based on the weighted sum of the cross entropy classification loss and the regression loss; The target detection model is trained, and the gradient is returned based on the value of the total loss function. The optimizer updates the model parameters until the preset number of iterations is reached to obtain the pre-trained target detection model.
8. The method for detecting hidden dangers of manhole covers based on target detection according to claim 7 is characterized in that: The total loss function is expressed as: ; in, represents the total loss function; Represents the classification loss, expressed as , represents the true label, Indicates the category The predicted probability of Represents the regression loss, expressed as , represents the predicted bounding box, represents the ground-truth bounding box, Represents the intersection of the predicted bounding box and the true bounding box and union The ratio of is expressed as ; and is the loss weight hyperparameter.
9. The method for detecting hidden dangers of manhole covers based on target detection according to claim 8, characterized in that: The programmable gradient information module is specifically used for: The primary feature maps of multiple different scales output by the backbone network are respectively input into the cross-branch fusion module of the corresponding scale in the programmable gradient information module; After the training image passes through the convolution, efficient aggregation module and inverted residual attention downsampling module, it is input into the cross-branch fusion module of the Kth scale; Initialize k=K, update k=k-1 until k=1; pass the output of the cross-branch fusion module of the kth scale through the efficient aggregation module to output the aggregated feature map of the kth scale; pass the aggregated feature map of the kth scale through the inverted residual attention downsampling module and input it into the cross-branch fusion module of the k-1th scale; pass the aggregated feature map of the kth scale through the detection head output to obtain the category label, bounding box and confidence of the kth scale; Get the class labels, bounding boxes, and confidences at each scale as the output of the programmable gradient information module.
Citation Information
Patent Citations
Yolov5 target detection method based on cross-stage routing attention module and residual information fusion module
CN116721398A
Improved YOLOv4-based power grid infrastructure target detection method
CN117649514A
Well lid hidden danger integrated identification method and system fusing target detection and segmentation
CN119672418A
Multi-task joint perception network model and detection method for traffic road surface information
US20240420487A1
Cited By
Abrasive hardening detection method and system of deep network structure
CN121213571A
Deep network structure based abrasive plate bonding detection method and system
CN121213571B
Deep well block falling while drilling real-time detection method based on DA-YOLO
CN122115810A
Deep well while-drilling block shedding real-time detection method based on DA-YOLO
CN122115810B