Well lid damage detection method and device based on deep learning, electronic equipment and program product
By introducing a multi-scale residual global attention mechanism (MSRGA) submodule into the backbone network of manhole cover damage detection method, the problem of degradation of manhole cover damage detection accuracy in complex scenarios is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510440063.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-09
AI Technical Summary
When detecting damage to manhole covers, the prior art is affected by complex scenarios and diversified environments, resulting in a decrease in detection accuracy.
Using deep learning-based manhole cover damage detection method, multi-scale residual global attention mechanism (MSRGA) submodule is introduced into the backbone network to extract and fuse multi-scale features to reduce interference from complex backgrounds.
It significantly improves the accuracy and reliability of manhole cover damage detection, can more effectively capture multi-scale damage characteristics of different shapes and sizes, and improves detection performance in complex scenarios.
Smart Images

Figure CN119964013A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and in particular relates to a manhole cover damage detection method, a manhole cover damage detection device, an electronic device and a computer program product based on deep learning. Background Art
[0002] With the continuous advancement of urbanization, municipal infrastructure construction has been gradually improved. As an indispensable and important part of urban roads, the safety of manhole covers is directly related to the travel, life and property safety of citizens. Manhole covers not only carry important functions such as drainage and passage, but also play an important role in protecting underground pipelines, maintaining road flatness and ensuring urban operation. However, with the increase in the service life of manhole covers, changes in environmental conditions and the influence of external force damage, incidents such as manhole cover displacement, damage and even loss occur frequently, posing serious hidden dangers to pedestrians, vehicles and public safety.
[0003] Although deep learning can currently be used to provide new technical support for automated detection of manhole cover damage, since manhole covers are distributed in a variety of scenes such as urban roads, squares, parks, and around buildings, the images collected in different scenes are easily affected by the corresponding environment, reducing the accuracy of manhole cover damage detection. Summary of the invention The present application provides a manhole cover damage detection method, a manhole cover damage detection device, an electronic device and a computer program product based on deep learning, which can capture multi-scale damage features of different shapes and sizes, reduce interference from complex scenes, and thus significantly improve the accuracy and reliability of manhole cover damage detection.
[0004] In a first aspect, the present application provides a method for detecting manhole cover damage based on deep learning, comprising: The backbone network of the manhole cover detection model trained in advance performs feature extraction on the image to be detected to obtain image features; the image to be detected includes the manhole cover; The neck network based on the manhole cover detection model fuses the image features to obtain the target fusion features; The detection network based on the manhole cover detection model detects the target fusion features and obtains the detection result of the manhole cover in the image to be detected; Among them, the backbone network includes m feature extraction modules of different scales, m is a positive integer, and m≥2; the feature extraction module is provided with an MSRGA submodule, the MSRGA submodule extracts and fuses multi-scale features through the channel attention mechanism and the spatial attention mechanism, and the image features are obtained based on the fused multi-scale features.
[0005] In a second aspect, the present application provides a manhole cover damage detection device, comprising: An extraction module is used to extract features of an image to be detected based on a backbone network of a pre-trained manhole cover detection model to obtain image features; the image to be detected includes a manhole cover; A fusion module is used to fuse image features based on the neck network of the manhole cover detection model to obtain fusion features; A detection module is used to detect the fused features based on the detection network of the manhole cover detection model to obtain the detection result of the manhole cover in the image to be detected; Among them, the backbone network includes m feature extraction modules of different scales, m is a positive integer, and m≥2; the feature extraction module is provided with an MSRGA submodule, the MSRGA submodule extracts and fuses multi-scale features through the channel attention mechanism and the spatial attention mechanism, and the image features are obtained based on the fused multi-scale features.
[0006] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of the first aspect when executing the computer program.
[0007] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method of the first aspect are implemented.
[0008] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by one or more processors, it implements the steps of the method of the first aspect.
[0009] Compared with the prior art, the present application has the following beneficial effects: the backbone network includes m feature extraction modules of different scales, where m is a positive integer and m≥2. By introducing the Multi-scale Residual Global Attention (MSRGA) submodule in the feature extraction module, multi-scale features are extracted and fused through the channel attention mechanism and the spatial attention mechanism, the damage features of different sizes and shapes are highlighted and enhanced, and the representation ability of the backbone network is enhanced to reduce the interference of complex backgrounds on the identification of manhole cover damage. The image features extracted by this backbone network will contain rich and discriminative manhole cover damage features, so that the feature fusion based on it can more comprehensively integrate the effective information for identifying manhole cover damage, and significantly improve the detection accuracy of manhole cover damage in complex scenarios.
[0010] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 is a schematic diagram of the network structure of the manhole cover detection model provided in the embodiment of the present application; Figure 2 is an example graph in the MCD data set provided in the embodiment of the present application; Figure 3 It is a flowchart of a method for detecting manhole cover damage based on deep learning provided in an embodiment of the present application; Figure 4 is a schematic diagram of the network structure of the MSRGA submodule provided in an embodiment of the present application; Figure 5 Schematic diagram of the network structure of the MSWFF module provided in the embodiment of the present application; Figure 6 is a schematic diagram of the network structure of the CSFF module provided in an embodiment of the present application; Figure 7 is a structural schematic diagram of a manhole cover damage detection device provided in an embodiment of the present application; Figure 8 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0013] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0014] Damaged manhole cover identification faces many challenges in complex background environments, such as vehicle occlusion, pedestrian interference, vegetation coverage, shadow changes, and uneven lighting. These factors reduce image quality and increase the difficulty of feature extraction. In addition, the background texture and color are similar to the surface of the manhole cover, which can easily cause confusion in the damaged area and increase the difficulty of detection. The combination of these complex factors requires the model to have stronger robustness and accuracy to ensure accurate identification and classification of damaged manhole covers.
[0015] In order to solve this problem, this application proposes a manhole cover detection model, which includes a backbone network, a neck network and a detection head. Specifically, the backbone network is used to extract features from the input image to obtain image features; the neck network is used to perform feature fusion on the image features to obtain target fusion features; the detection head is used to detect the target fusion features to obtain the detection result of manhole cover damage in the image to be processed.
[0016] Among them, in order to ensure the accuracy of manhole cover damage detection in complex environments, the backbone network includes m feature extraction modules of different scales, where m is a positive integer and m≥2. By introducing the Multi-scale Residual Global Attention (MSRGA) submodule in the feature extraction module, multi-scale features are extracted and fused through the channel attention mechanism and the spatial attention mechanism, the damage features of different sizes and shapes are highlighted and enhanced, and the representation ability of the backbone network is enhanced to reduce the interference of complex background on manhole cover damage identification. The image features extracted by this backbone network will contain rich and discriminative manhole cover damage features, so that the feature fusion based on it can more comprehensively integrate the effective information for identifying manhole cover damage, and significantly improve the detection accuracy of manhole cover damage in complex scenes.
[0017] In some embodiments, the MSRGA submodule consists of a multi-scale channel attention mechanism (Efficient ChannelAttention, ECA) fusion block and a multi-scale spatial attention mechanism (Efficient SpatialAttention, ESA) fusion block, and is connected using a residual structure. By introducing multi-scale feature extraction in ECA and ESA, the model can effectively capture key information under different receptive fields. The residual connection not only enhances the stability of the model, but also retains the original input features when the two attention mechanisms fail to ensure information integrity.
[0018] In some embodiments, the multi-scale ECA fusion block is composed of a first convolutional layer, a multi-scale feature extraction structure, and an ECA structure. The multi-scale feature extraction structure includes at least two feature extraction layers of different scales, each of which is composed of an adaptive pooling and a depth-separable convolution in series. The structure effectively captures global and local information through the multi-scale feature extraction layer, and enhances the feature interaction between channels under the action of the ECA structure, thereby improving the feature expression capability.
[0019] In some embodiments, the multi-scale ESA fusion block is composed of a pooling structure, a first splicing layer, a parallel convolution structure, and an ESA structure. The pooling structure includes a multi-scale adaptive pooling layer, and the pooling layer of each scale is composed of average pooling and maximum pooling in series, which can efficiently obtain the spatial features of the dual channels. The parallel convolution structure is composed of at least two convolution layers with different convolution kernels, which can extract spatial information of different receptive fields from the spliced dual-channel spatial features, and further optimize the feature expression under the action of the ESA structure to improve the accuracy of spatial attention.
[0020] In some embodiments, the multi-scale ECA fusion block in series works with the multi-scale ESA fusion block. The former uses multi-scale adaptive pooling to efficiently capture global and local information, while the latter extracts deep feature information through convolutions of different scales, achieving complementary enhancement of global and local deep information and improving feature expression capabilities. Further combined with residual connections, the original information can be maintained and complex mapping relationships can be learned, thereby enhancing the robustness of the model in complex backgrounds and effectively improving the accuracy of manhole cover damage detection. The MSRGA submodule enables the model to focus on key damage features, while integrating global contextual information, reducing background interference, improving detection efficiency and reliability, and enabling it to have stronger generalization capabilities in complex environments.
[0021] In addition to complex environmental factors, the wide variety of manhole covers and various forms of damage also pose challenges to damage detection. Rainwater, sewage, power and communication manhole covers of different materials, shapes and sizes, as well as various types of damage such as cracks, fractures, corrosion, etc., require the model to have stronger adaptability and generalization capabilities. When complex environments are intertwined with diverse manhole covers and damage forms, the difficulty of detection will be further increased.
[0022] In some embodiments, to address this challenge, the neck network may include n feature fusion modules of different scales, where n is a positive integer not less than 2. Each feature fusion module is provided with a multi-scale weighted feature fusion (MSWFF) module, which is used to align and weight the initial fusion features output by the feature fusion module to generate target fusion features corresponding to each scale.
[0023] The MSWFF module combines features at different scales to enhance the ability to capture local details and global context information, thereby improving the model's recognition of various damage types. Under the combined influence of complex environments and diverse manhole covers and damage forms, this mechanism ensures that the detection model can accurately adapt to different scenarios and achieve more efficient and reliable manhole cover damage detection.
[0024] In some embodiments, the MSWFF module consists of a scale alignment layer, a convolutional splicing layer, an ECA layer, and an ESA layer. The scale alignment layer lays the foundation for the effective fusion of features of different scales, while the convolutional splicing layer further integrates multi-scale information to achieve full interaction of features. Through the adaptive adjustment of the ECA layer and the ESA layer, the module can accurately allocate channel weights and spatial weights, enhance the ability to capture local details and global context information, and ensure that all types of damage can be effectively identified.
[0025] The network structure design of the MSWFF module not only improves the robustness and accuracy of detection, but also can efficiently handle multi-scale problems from subtle defects to large-scale damage, making the model more adaptable and generalizable in complex environments.
[0026] The diversity of manhole cover damage makes the lack of clear details a major challenge in manhole cover damage identification, especially in complex real-world scenarios. Due to factors such as material differences, aging, stains, and lighting changes, damage features such as cracks and deformations on the manhole cover surface are often difficult to clearly capture and accurately distinguish, resulting in higher uncertainty in feature extraction and classification for detection algorithms.
[0027] In some embodiments, in order to deal with the uncertainty in feature extraction and classification, the backbone network introduces a convolution and space-to-depth feature fusion (CSFF) module. The CSFF module includes a convolution branch, an SPD convolution branch, a second concatenation layer, a channel attention mechanism (Effective Squeeze and Extraction, ESE) layer, and a second convolution layer.
[0028] Through the feature extraction of two branches, the CSFF module can capture local details and multi-scale contextual information at the same time, thereby effectively extracting subtle and complex damage features on the surface of the manhole cover. Even if there are interferences such as stains and lighting changes in the detection image, it can still maintain efficient and accurate feature extraction. The ESE layer highlights important features and suppresses redundant information through adaptive adjustment between channels, further enhancing the focus on subtle features. Ultimately, the fused features are mapped to a higher-dimensional space, increasing the richness and distinctiveness of the features, thereby highlighting the boundaries between different types of damage, which helps to accurately distinguish different types of manhole cover damage.
[0029] In some embodiments, the manhole cover detection model can be improved based on the YOLO series model in view of the following advantages of the YOLO series model: Both are end-to-end single-stage detection frameworks that can achieve efficient target detection. Compared with two-stage detectors (such as Faster R-CNN), YOLO directly regresses bounding boxes and categories through a single forward propagation, which improves the detection speed and is suitable for real-time applications. Its structure is constantly optimized, such as the introduction of Anchor-Free mechanism, feature pyramid (FPN, PAN), etc., which enhances the detection ability of small targets and multi-scale targets. In addition, the YOLO series models are continuously optimized in terms of lightweight, computational efficiency, and robustness, making them well adaptable in both embedded devices and cloud inference scenarios. YOLOv8 is a major upgrade of the YOLO series, supporting tasks such as target detection, image classification, and instance segmentation. Its architecture consists of a backbone network (Backbone), a neck network (Neck), and a detection head (Head). The backbone network uses the C2f module to improve feature extraction efficiency, the neck network uses the PANet structure to enhance multi-scale feature fusion, and the detection head introduces the Anchor-Free design and DFL loss to improve detection accuracy and flexibility. In addition, YOLOv8 provides five model variants: n / s / m / l / x, which are suitable for different scenarios. In the manhole cover damage detection task, YOLOv8s is selected for optimization to balance accuracy and real-time performance and improve the detection effect.
[0030] Based on this, YOLOv8s can be preferred when building a manhole cover detection model. For example, if YOLOv8s is used as the basic network for improvement, the network structure of the manhole cover detection model can be found in Figure 1 .
[0031] The manhole cover detection model is equipped with feature extraction modules of different scales in the backbone network. Each feature extraction module of each scale is equipped with an MSRGA submodule. The MSRGA submodule captures multi-scale damage features of different shapes and sizes based on two attention mechanisms, channel and space, and simultaneously considers context information, which can reduce background interference and improve the detection effect of manhole cover damage under complex backgrounds. The backbone network is also equipped with a CSFF module. The CSFF module simultaneously extracts local details and multi-scale context information through convolution branches and SPD convolution branches to effectively capture the subtle and complex damage features on the surface of the manhole cover, enhance feature representation, and improve the recognition accuracy and reliability of manhole cover damage. In addition, the neck network is equipped with an MSWFF module after the feature fusion module of each scale, which helps to combine the initial fusion features of different scales to enhance the local detail information and global context information of the local features, thereby improving the recognition ability of different types of manhole covers corresponding to various forms of damage.
[0032] In some embodiments, in order to ensure that the trained manhole cover detection model can meet expectations, accurately understand and detect manhole covers of different scales in different scenarios, a special dataset can be created for training the manhole cover detection model. Specifically, a dataset called manhole cover damage (MCD) is created. The dataset is collected by a vehicle-mounted camera under various environmental conditions and includes 2,000 images with a resolution of 1280×720, and each image is manually annotated. Figure 2 , Figure 2 An example image of a training sample is shown.
[0033] For example, in order to provide a high-quality data set, in addition to ensuring the number of images, images covering different scenes such as urban roads, squares, parks, and areas around buildings can also be collected.
[0034] For example, the image may contain common obstructions, such as vehicles, people, vegetation, warning objects (warning triangles), etc.
[0035] For example, the types of manhole covers in the image should cover as many manhole covers as possible, including rainwater, sewage, electricity, communications, etc. manhole covers of different materials, shapes, and sizes.
[0036] For example, the types of damage to the manhole cover in the image should cover as many types of damage as possible, such as cracks, fractures, corrosion, etc. of different shapes and sizes.
[0037] Exemplarily, each scene can include multiple time periods and weather conditions such as day, night, sunny days, rainy days and foggy days to increase the complexity and comprehensiveness of the scene. Each image can be configured with a corresponding txt label file that details the location and category of the manhole cover.
[0038] Such datasets can support research and applications in areas such as intelligent transportation systems, autonomous vehicles, and road safety monitoring to improve road safety and traffic efficiency.
[0039] In order to ensure the effectiveness of model training and the reliability of evaluation results, each image in the dataset of this application is divided into a training set and a validation set in a ratio of 8:2.
[0040] Preferably, in the process of dividing the data set, special attention is paid to maintaining the balance of sample distribution between the training set and the validation set, ensuring that the two have similar statistical characteristics in terms of manhole cover type, size, background environment, etc. This can avoid evaluation bias caused by uneven sample distribution and make the validation results more representative and credible.
[0041] In some embodiments, in order to comprehensively and accurately measure the performance of each version of the manhole cover detection model, after at least one version of the manhole cover detection model is converged based on the training set, the converged model can be evaluated through the validation set to avoid model overfitting, verify the generalization of the model, and ensure that the model can run stably after deployment.
[0042] Specifically, the performance of the manhole cover detection model on the validation set can be evaluated according to preset conditions.
[0043] Exemplarily, these preset conditions may include performance indicators such as the intersection-over-union ratio, detection accuracy and recall rate of the manhole cover damage by the manhole cover detection device.
[0044] IoU is an indicator to evaluate the degree of overlap between the predicted border and the true border. It calculates the ratio of the intersection and union of the predicted border and the true border. It plays a key role in determining whether it is a correct detection. The calculation formula of IoU can be written as:
[0045] Precision (p), also known as the precision rate, refers to the proportion of correct positive predictions to all positive predictions, as shown in the formula:
[0046] Recall (R), also known as recall rate, refers to the proportion of correctly predicted positives to all actual positives, as shown in the formula:
[0047] The average precision (AP) is calculated from the precision and recall. A line graph of the precision is drawn based on the recall value, and the area under the line is calculated, as shown in the formula:
[0048] The mean average precision (mAP) refers to the average of the average precision AP of C different manhole cover damage categories, as shown in the formula:
[0049] That is to say, after verifying each version of the manhole cover detection model through the verification set, the manhole cover detection model of each version can be comprehensively evaluated based on the above-mentioned indicators, so as to determine the manhole cover detection model with the best performance from each version as the trained manhole cover detection model.
[0050] In some embodiments, the model running environment includes an Intel Xeon Platinum 8255C processor, 314GB memory, NVIDIA Tesla V100 32 GB graphics card, and the operating system is CentOS 8.5.2 (64-bit). The deep neural network is built based on the PyTorch framework, the input image size is [640, 640], and a multi-scale training strategy is adopted. The experiment sets the batch size to 64, trains 300 rounds (epochs), uses the AdamW optimizer, the initial learning rate is 0.01, and is optimized in combination with the cosine decay strategy.
[0051] In some embodiments, the training of the manhole cover detection model is optimized by combining classification loss and regression loss. The classification loss uses binary cross entropy loss (BCE Loss) to determine the anchor box category; the regression loss consists of SIoU Loss and DFLLoss to measure the error between the predicted box and the true box. The positive and negative sample matching strategy uses TAL dynamic matching to optimize target allocation and improve detection accuracy.
[0052] Based on the network structure of the manhole cover detection model in the previous embodiment, the present application proposes a manhole cover damage detection method based on deep learning.
[0053] The deep learning-based manhole cover damage detection method provided in the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, vehicle-mounted equipment, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.
[0054] In order to illustrate the technical solution proposed in this application, each embodiment will be described below with an electronic device as the execution subject.
[0055] Figure 3 A schematic flow chart of a deep learning-based manhole cover damage detection method provided by the present application is shown, and the deep learning-based manhole cover damage detection method includes: Step 310: The electronic device extracts features of the image to be detected based on the backbone network of the pre-trained manhole cover detection model.
[0056] Step 320: The electronic device fuses the image features based on the neck network of the manhole cover detection model to obtain fused features.
[0057] Step 330: The electronic device detects the fused features based on the detection network of the manhole cover detection model to obtain the detection result of the manhole cover in the image to be detected.
[0058] In this embodiment, in order to ensure the accuracy of manhole cover damage detection in complex environments, the backbone network includes m feature extraction modules of different scales, where m is a positive integer and m≥2. By introducing the MSRGA submodule in the feature extraction module, multi-scale features are extracted and fused through the channel attention mechanism and the spatial attention mechanism, damage features of different sizes and shapes are highlighted and enhanced, and the representation ability of the backbone network is enhanced to reduce the interference of complex backgrounds on manhole cover damage identification. The image features extracted by this backbone network will contain rich and discriminative manhole cover damage features, so that the feature fusion based on it can more comprehensively integrate the effective information for identifying manhole cover damage, and significantly improve the detection accuracy of manhole cover damage in complex scenarios.
[0059] In some embodiments, the MSRGA submodule includes a multi-scale ECA fusion block and a multi-scale ESA fusion block, and the multi-scale ECA fusion block is connected with the multi-scale ESA fusion block based on a residual structure; for the first input feature input to the MSRGA submodule, the electronic device may perform the following operations: Step A1: The electronic device performs an ECA-based multi-scale weighted fusion operation on the first input feature through a multi-scale ECA fusion block to obtain a first fused feature.
[0060] Step A2: The electronic device performs an ESA-based multi-scale weighted fusion operation on the first fusion feature through a multi-scale ESA fusion block to obtain a second fusion feature.
[0061] Step A3: The electronic device concatenates the second fusion feature with the first input feature through a residual structure to obtain a first output feature corresponding to the first input feature.
[0062] The multi-scale ECA fusion block can capture global information and local detail features, and the multi-scale ESA fusion block of multi-scale convolution can capture deep feature information, so that the multi-scale ECA fusion block and the multi-scale ESA fusion block in series can realize the dynamic fusion of global and local deep features, focusing on capturing multi-scale damage features of different sizes and shapes. On this basis, the residual connection is further combined to maintain the original information and learn complex mapping relationships, which can ensure the detection accuracy of manhole cover damage under complex backgrounds.
[0063] In some embodiments, the multi-scale ECA fusion block includes a first convolutional layer, a multi-scale feature extraction structure, and an ECA structure; the multi-scale feature extraction structure includes feature extraction layers at least at two scales; for the first input feature, the electronic device may perform the following operations: Step A11: The electronic device performs a channel compression operation on the first input feature through a first convolutional layer to obtain a first convolutional feature.
[0064] The first convolution layer can use 1×1 convolution and combine it with SiLU (Sigmoid Linear Unit) activation function. 1×1 convolution is used to adjust the number of channels and fuse local information, while SiLU activation function can enhance gradient flow, improve nonlinear expression ability, and help improve the efficiency and effect of feature extraction.
[0065] Step A12: The electronic device sequentially performs adaptive pooling operations and depthwise separable convolution operations on the first convolutional features through the feature extraction layers at each scale to obtain convolutional features at each scale.
[0066] In order to extract global features of different scales, a feature extraction layer with different kernels can be used. The feature extraction layer consists of adaptive pooling and depth-wise separable convolution. Taking the extraction of features of three scales as an example, the kernel size of the feature extraction layer can be set to 1, 3, and 5 respectively, where the adaptive pooling and depth-wise separable convolution corresponding to each convolution kernel can extract the first convolution feature to obtain the convolution feature of the corresponding scale. Adaptive pooling is used to extract global features of different scales, while depth-wise separable convolution further enhances the expression of local features. After these two operations, the global and local fusion features of the corresponding scales (1, 3, 5) can be obtained respectively, thereby realizing the effective capture of multi-scale information. Through this multi-scale feature extraction method, global and local information can be fully captured, and the multi-scale feature expression ability of the model can be enhanced.
[0067] Step A13: The electronic device performs an ECA operation on the convolution features at each scale through the ECA structure, and performs weighted fusion on the first input features with the obtained channel weights to obtain a first fused feature.
[0068] The electronic device performs channel attention on convolution features of different scales through the ECA structure to calculate channel weights; then, the calculated channel weights are used to perform weighted fusion on the first input features to enhance the information expression of important channels and suppress the influence of secondary channels, thereby obtaining the optimized first fused features.
[0069] In this embodiment, the electronic device effectively captures global and local information through adaptive pooling of different scales and depth-separable convolution in the multi-scale feature extraction structure, and calculates channel weights under the action of the ECA structure to enhance the expression of important features. Finally, the first fusion feature after weighted fusion optimization is obtained to improve the recognition accuracy and robustness of the model for manhole cover damage.
[0070] In some embodiments, the multi-scale ESA fusion block includes a pooling structure, a first splicing layer, a parallel convolution structure, and an ESA structure; for the first fusion feature, the electronic device performs the following steps: Step A21: The electronic device performs an average pooling operation and a maximum pooling operation on the first fusion feature through a pooling structure to obtain a first pooling result and a second pooling result.
[0071] Average pooling can smooth features and retain overall information, while maximum pooling can highlight significant features and increase attention to key areas. Through these two pooling methods, the device obtains the first pooling result and the second pooling result, providing a basis for subsequent feature fusion.
[0072] Step A22: The electronic device splices the first pooling result and the second pooling result through a first splicing layer to obtain a first splicing feature.
[0073] This splicing operation can effectively combine the information extracted by different pooling methods, making the feature expression richer and providing diverse information input for subsequent convolution operations.
[0074] Step A23: The electronic device performs at least two convolution operations on the first splicing feature through a parallel convolution structure to obtain at least two corresponding convolution results.
[0075] The parallel convolution structure may include convolution layers with different convolution kernel sizes, and through different convolution operations, multi-level and multi-scale feature representations can be obtained. These convolution operations may include convolutions of different sizes, and the convolutions may be standard convolutions, depthwise separable convolutions, or dilated convolutions to enhance the ability to capture local details and global information, and ultimately obtain at least two different convolution results.
[0076] Step A24: The electronic device performs an ESA operation on each convolution result through the ESA structure, and performs weighted fusion on the first fusion feature with the obtained spatial weight to obtain a second fusion feature.
[0077] The electronic device performs ESA operation on each convolution result through the ESA structure to calculate the spatial weight, and uses the weight to weightedly fuse the first fusion feature to obtain the second fusion feature. The ESA structure improves the attention of key areas by adaptively adjusting the spatial attention, and the ESA mechanism can enhance the selectivity of channel information, thereby improving the expressiveness of the overall features and improving the accuracy and robustness of the detection task.
[0078] In this embodiment, the electronic device realizes multi-level feature extraction and optimization of the first fusion feature through operations such as pooling, splicing, parallel convolution and spatial attention fusion. The combination of average pooling and maximum pooling helps to retain the overall information while highlighting the key features; the splicing operation fuses different pooling results to improve the richness of feature expression; the parallel convolution structure obtains multi-scale features through multiple convolution operations to improve the ability to capture deep local details and global information; finally, the ESA structure combines the ESA mechanism to calculate the spatial weight, and performs weighted fusion on the first fusion feature to further enhance the focus on key information and obtain the second fusion feature to improve the feature expression ability and detection accuracy of the model.
[0079] In some embodiments, Figure 4 The network structure diagram of the MSRGA submodule is shown. Figure 4 The network structure of , first perform channel dimensionality reduction to reduce the amount of calculation and highlight important features to obtain the first convolution feature Y :
[0080] in, Is a 3×3 convolutional layer used to convert the number of input channels C Reduce to , σ is the SiLU activation function, which is used to enhance feature nonlinearity.
[0081] Extract structure pairs through multi-scale features Y Perform multi-scale processing, that is, use adaptive pooling of different scales and corresponding deep convolution to extract multi-scale context information:
[0082] The kernel size is Adaptive pooling operations are used to extract global and local context. It is a depth-wise separable convolution of the corresponding kernel size, which further captures the local features after pooling. These features have different receptive fields. , focusing on the local and global information of the features respectively, and obtaining the convolution features at each scale. After concatenating the convolution features at different scales, we obtain the concatenated features. , through 1×1 convolution Restore the original number of channels and perform the channel attention mechanism:
[0083] The channel attention mechanism determines the channel weights that contain the weights of each channel , and XThe weighted fusion operation is performed on each channel feature to obtain the first fusion feature X MSCA :
[0084] where ⊙ represents element-wise multiplication.
[0085] based on X MSCA Calculate the global average pooling and global maximum pooling to obtain the spatial features of the two channels. The first pooling result And the second pooling result :
[0086] Then, the two pooling results are concatenated to obtain the first concatenated feature :
[0087] exist Multi-scale convolution is applied to extract features of different receptive fields, that is, the convolution results , i= {1,3,5,7}:
[0088] After the convolution results are concatenated, a 1×1 convolution is performed to generate a spatial attention mechanism:
[0089] The spatial attention mechanism can weight each spatial position of the concatenated convolution results to obtain the second fusion feature :
[0090] Combine channel attention and spatial attention, and add residual connections, which can retain the original features. X , while alleviating the gradient vanishing problem and ensuring the effectiveness of the attention module.
[0091]
[0092] The MSRGA submodule can extract and capture multi-scale damage features of different sizes and shapes, while combining residual connections to retain original information and learn complex mapping relationships, thereby effectively detecting manhole cover damage in complex backgrounds. The global attention mechanism enables the model to focus on key damage features while considering the overall context, reducing background interference, and improving the robustness of feature expression, enabling it to achieve more efficient and reliable detection results in complex environments.
[0093] In some embodiments, the neck network includes n feature fusion modules of different scales, where n is a positive integer and n≥2; each feature fusion module corresponds to an MSWFF module, and the MSWFF modules corresponding to each scale are used to align and weightedly fuse the initial fusion features output by each feature fusion module to obtain the target fusion features corresponding to each scale.
[0094] Exemplarily, n may be consistent with the aforementioned m, that is, the number of feature extraction modules set in the backbone network is the same as the number of feature fusion modules set in the neck network.
[0095] In some embodiments, the MSWFF module includes a scale alignment layer, a convolutional splicing layer, an ECA layer, and an ESA layer. The MSWFF module performs the following steps on the initial fusion features output by each feature fusion module: Step B1: At the current scale, the electronic device aligns the scale of each initial fusion feature with the current scale through a scale alignment layer to obtain each aligned feature.
[0096] The electronic device adjusts the scale of each initial fusion feature through the scale alignment layer to keep it consistent with the current scale of the MSWFF module, and obtains each aligned feature. This alignment operation ensures that when each feature is subsequently fused, each aligned feature can be effectively matched in the spatial dimension, avoiding information loss or uneven fusion caused by scale differences.
[0097] Step B2: The electronic device performs a convolution operation and a splicing operation on each alignment feature through a convolution splicing layer to obtain a second splicing feature.
[0098] The convolution operation is used to extract local features and enhance the representation ability, while the splicing operation integrates features from different sources to improve the richness of feature expression, thereby obtaining a second splicing feature.
[0099] Step B3: The electronic device performs a weighted fusion operation based on a channel attention operation on the second splicing feature through the ECA layer to obtain a third fusion feature, and splices the second splicing feature with the third fusion feature to obtain a third splicing feature.
[0100] The electronic device performs a weighted fusion operation based on the channel attention mechanism on the second spliced feature through the ECA layer to dynamically adjust the importance of different channels, highlight key features, and suppress redundant information. Subsequently, the second spliced feature is spliced with the fused third fused feature to enrich the feature expression and provide a more comprehensive feature representation, thereby obtaining the third spliced feature.
[0101] Step B4: The electronic device performs a weighted fusion operation based on spatial attention on the third fusion feature operation through the ESA layer to obtain the target fusion feature at the current scale.
[0102] The electronic device performs a weighted fusion operation based on spatial attention on the third fusion feature through the ESA layer, using the spatial attention mechanism to enhance the model's attention to key areas and reduce the interference of irrelevant background. Finally, more discriminative target fusion features are extracted at the current scale, providing more accurate feature representation for subsequent detection tasks.
[0103] It can be considered that the MSWFF module includes MSWFF branches of different scales, which correspond one-to-one to each initial fusion module.
[0104] In this embodiment, the electronic device first adjusts the scale of the initial fused features through the scale alignment layer to match the current scale of the MSWFF module to ensure that the features are aligned in the spatial dimension. Subsequently, the convolutional splicing layer performs convolution and splicing operations on the aligned features to enhance the feature expression capability. Next, the ECA layer uses the channel attention mechanism to weightedly fuse the spliced features, highlight the key channel information, and further fuse them with the original spliced features to obtain a richer feature representation. Finally, the ESA layer weights the fused features through the spatial attention mechanism, so that the model focuses more on the key areas and reduces background interference, thereby obtaining more discriminative target fusion features at the current scale.
[0105] In some embodiments, the neck network can output target fusion features corresponding to each scale, and set a decoupled detection head for each scale to optimize the independence of classification and positioning tasks. Accordingly, referring to FIG1, the detection head processes target fusion features of different scales respectively, the classification branch improves the category discrimination ability, and the regression branch optimizes the bounding box prediction accuracy, thereby enhancing the model's detection ability for targets of different sizes and improving the accuracy and stability of detection.
[0106] In some embodiments, Figure 5 Figure 2 shows the network structure diagram of the MSWFF module. Figure 5 The network structure of MSWFF branches of different scales has the same subsequent operations except for the difference in scale alignment.
[0107] The initial fusion features for three different scales include: High-resolution initial fusion features: , the resolution is ; Medium resolution initial fusion features: , the resolution is ; Low-resolution initial fusion features: , the resolution is For ease of description, the MSWFF branches corresponding to the high, medium and low resolutions may be respectively recorded as P3 branch, P4 branch and P5 branch.
[0108] First, the features of different branches are resized and channel compressed through spatial downsampling, upsampling interpolation or convolutional layers to complete feature alignment:
[0109] For each initial fusion feature, the same feature alignment operation is different in different branches. Specifically: P3 branch:
[0110] P4 branch:
[0111] P5 branch:
[0112] The alignment features of each branch input can be expressed as , as well as , for different branches, i The value of is different. i ={3, 4, 5}.
[0113] Since the operations performed by subsequent branches are consistent, each branch is abstracted as a general branch for description, that is, each branch performs the operations performed by the general branch.
[0114] Concatenate the alignment features corresponding to the common branches to obtain the second concatenation feature F i ; Through the efficient channel attention mechanism ECA F i Perform weighted fusion operation to obtain the third splicing feature F ECA-i :
[0115] in, GAP is the global average pooling, is a 1×1 convolution, σ is Sigmoid Activation function.
[0116] Then, the spatial weights are generated by efficient spatial attention ESA W i :
[0117] in: It is a 1×1 convolution.
[0118] Finally, using spatial weights Fusion FECA , for the third fusion feature after fusion Perform dimension expansion and output the corresponding target fusion features.
[0119]
[0120] It can be understood that in the mathematical expression of the above general branch, i The parameters depend on the specific branch. i Different values indicate corresponding parameters of different branches.
[0121] In some embodiments, the backbone network further includes a CSFF module, the CSFF module includes a convolution branch, an SPD convolution branch, a second concatenation layer, an ESE layer, and a second convolution layer; for the second input feature of the input CSFF module: Step C1: The electronic device performs a convolution operation on the second input feature through a convolution branch to obtain a second convolution feature.
[0122] The electronic device performs a standard convolution operation on the second input feature through the convolution branch to extract local features and enhance feature expression capabilities, thereby obtaining a second convolution feature, which provides a basis for subsequent feature fusion.
[0123] Step C2: The electronic device performs a space-to-depth conversion operation and a convolution operation on the second input feature through the SPD convolution branch to obtain a rearranged feature.
[0124] The electronic device performs a space-to-depth conversion operation on the second input feature through the SPD convolution branch to rearrange the feature space distribution, improve the compactness of the feature expression, and extracts deep information in combination with the convolution operation to obtain the rearranged feature.
[0125] Step C3: The electronic device performs a splicing operation on the second convolutional feature and the rearrangement feature through the second splicing layer to obtain a fourth splicing feature.
[0126] The electronic device splices the second convolutional feature and the rearrangement feature through the second splicing layer to integrate the advantages of the two features, form a richer representation capability, and obtain a fourth splicing feature.
[0127] Step C4: The electronic device performs a global average pooling operation and a convolution operation on the fourth splicing feature through the ESE layer, and weights the fourth splicing feature with the calculated channel attention weight to obtain an initial loss feature.
[0128] The electronic device performs a global average pooling operation and a convolution operation on the fourth concatenated feature through the ESE layer, calculates the channel attention weight, and weights the fourth concatenated feature to highlight the key features, suppress redundant information, and obtain the initial loss feature.
[0129] Step C5: The electronic device convolves the initial loss feature through a second convolutional layer to obtain a second output feature corresponding to the second input feature.
[0130] The electronic device convolves the initial loss features through the second convolutional layer to further optimize the feature expression, obtains the second output features corresponding to the second input features, and finally constructs the image features based on the fused multi-scale features and the second output features to provide more accurate representation information for the detection task.
[0131] In this embodiment, the electronic device sequentially performs standard convolution and SPD transformation on the second input feature to extract local features and rearrange spatial information, thereby enhancing the feature expression capability. Subsequently, the features of different branches are fused through a splicing operation to enrich the feature representation capability. Next, the ESE layer uses global average pooling and channel attention weighting to highlight key information and optimize feature distribution. Finally, after further processing by the second convolutional layer, the final second output feature is obtained, and the image feature is constructed in combination with the multi-scale feature to provide a more accurate feature expression for the detection task.
[0132] In some embodiments, Figure 6 The network structure diagram of the CSFF module is shown. Figure 6 The network structure of , the convolution branch reduces the input second input feature through a 3×3 convolution operation with a stride of 2 X The size of the channel is compressed C comp , get the second convolution feature .
[0133]
[0134] At the same time, the SPD convolution branch performs X Perform space-to-depth conversion operations to avoid the loss of fine-grained spatial information, and use a 3×3 convolution to compress the number of channels and integrate features.
[0135]
[0136] Next, the output second convolution feature of the above two branches and rearrangement features Merge to get the fourth splicing feature Y ECA The ESE layer calculates the channel attention weights through global average pooling and a 1×1 convolution and applies the channel attention weights to Y ECA On top, we get the initial loss characteristics.
[0137]
[0138] Finally, a 3×3 convolutional layer is used to map the initial loss feature into a higher dimensional space to obtain the final second output feature .
[0139]
[0140] In some embodiments, the detection methods of the above embodiments are implemented based on a trained manhole cover detection model. The manhole cover detection model is provided with feature extraction modules of different scales in the backbone network, and each scale feature extraction module is provided with an MSRGA submodule. The MSRGA submodule captures multi-scale damage features of different shapes and sizes based on two attention mechanisms, channel and space, and simultaneously considers context information, which can reduce background interference and improve the detection effect of manhole cover damage under complex backgrounds. The backbone network is also provided with a CSFF module, which simultaneously extracts local details and multi-scale context information through convolution branches and SPD convolution branches to effectively capture subtle and complex damage features on the surface of the manhole cover, enhance feature representation, and improve the recognition accuracy and reliability of manhole cover damage. In addition, the neck network is provided with an MSWFF module after the feature fusion module of each scale, which helps to enhance the local detail information and global context information of the local features by combining the initial fusion features of different scales, thereby improving the recognition ability of various forms of damage of different types of manhole covers.
[0141] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0142] Corresponding to the manhole cover damage detection method based on deep learning in the above embodiment, Figure 7 A structural block diagram of a manhole cover damage detection device 7 provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0143] Reference Figure 7 , the manhole cover damage detection device 7 comprises: An extraction module 71 is used to extract features from an image to be detected based on a backbone network of a pre-trained manhole cover detection model to obtain image features; the image to be detected includes a manhole cover; A fusion module 72, used for fusing image features based on the neck network of the manhole cover detection model to obtain fused features; A detection module 73, used to detect the fused features based on the detection network of the manhole cover detection model to obtain the detection result of the manhole cover in the image to be detected; Among them, the backbone network includes m feature extraction modules of different scales, m is a positive integer, and m≥2; the feature extraction module is provided with an MSRGA submodule, the MSRGA submodule extracts and fuses multi-scale features through the channel attention mechanism and the spatial attention mechanism, and the image features are obtained based on the fused multi-scale features.
[0144] Optionally, the MSRGA submodule includes a multi-scale ECA fusion block and a multi-scale ESA fusion block, and the multi-scale ECA fusion block is connected with the multi-scale ESA fusion block based on a residual structure; the extraction module 71 includes a first extraction unit, which is used to: For the first input feature of the MSRGA submodule: Performing an ECA-based multi-scale weighted fusion operation on the first input feature through a multi-scale ECA fusion block to obtain a first fused feature; Performing an ESA-based multi-scale weighted fusion operation on the first fusion feature through a multi-scale ESA fusion block to obtain a second fusion feature; The second fusion feature is concatenated with the first input feature through a residual structure to obtain a first output feature corresponding to the first input feature.
[0145] Optionally, the multi-scale ECA fusion block includes a first convolutional layer, a multi-scale feature extraction structure and an ECA structure; the multi-scale feature extraction structure includes feature extraction layers at least at two scales; and the extraction unit is specifically used for: For the first input feature: Performing a channel compression operation on the first input feature through the first convolution layer to obtain a first convolution feature; The first convolution feature is sequentially subjected to adaptive pooling operations and depthwise separable convolution operations through the feature extraction layers at each scale to obtain convolution features at each scale; The ECA operation is performed on the convolution features at each scale through the ECA structure, and the first input feature is weightedly fused with the obtained channel weight to obtain the first fused feature.
[0146] Optionally, the multi-scale ESA fusion block includes a pooling structure, a first splicing layer, a parallel convolution structure and an ESA structure; the extraction unit is specifically used for: For the first fusion feature: Performing an average pooling operation and a maximum pooling operation on the first fusion feature through the pooling structure to obtain a first pooling result and a second pooling result; The first pooling result and the second pooling result are spliced together through a first splicing layer to obtain a first splicing feature; Perform at least two convolution operations on the first concatenated feature respectively through a parallel convolution structure to obtain at least two corresponding convolution results; The ESA structure performs an ESA operation on each convolution result, and uses the obtained spatial weight to perform weighted fusion on the first fusion feature to obtain the second fusion feature.
[0147] Optionally, the neck network includes n feature fusion modules of different scales, where n is a positive integer and n≥2; a MSWFF module is corresponding to each feature fusion module, and the MSWFF modules corresponding to each scale are used to align and weightedly fuse the initial fusion features output by each feature fusion module to obtain the target fusion features corresponding to each scale.
[0148] Optionally, the MSWFF module includes a scale alignment layer, a convolutional splicing layer, an ECA layer, and an ESA layer, and the fusion module 72 includes a fusion unit for: The scale of each initial fusion feature is aligned with the current scale of the MSWFF module through the scale alignment layer to obtain each aligned feature; The convolution operation and the concatenation operation are performed on each alignment feature through the convolution concatenation layer to obtain a second concatenation feature; Performing a weighted fusion operation based on a channel attention operation on the second spliced feature through the ECA layer to obtain a third fused feature, and splicing the second spliced feature with the third fused feature to obtain a third spliced feature; The third fusion feature operation is performed through the ESA layer. The weighted fusion operation based on spatial attention is performed on the third fusion feature operation to obtain the target fusion feature at the current scale.
[0149] Optionally, the extraction module 71 includes a second extraction unit, configured to: The backbone network also includes a CSFF module, which includes a convolution branch, an SPD convolution branch, a second concatenation layer, an ESE layer, and a second convolution layer; for the second input feature of the input CSFF module: Performing a convolution operation on the second input feature through the convolution branch to obtain a second convolution feature; Performing a space-to-depth conversion operation and a convolution operation on the second input feature through the SPD convolution branch to obtain a rearranged feature; Performing a concatenation operation on the second convolutional feature and the rearranged feature through the second concatenation layer to obtain a fourth concatenated feature; Perform global average pooling and convolution operations on the fourth concatenated feature through the ESE layer, and weight the fourth concatenated feature with the calculated channel attention weight to obtain the initial loss feature; The initial loss feature is convolved through the second convolution layer to obtain a second output feature corresponding to the second input feature, and the image feature is obtained based on the fused multi-scale feature and the second output feature.
[0150] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0151] Figure 8 This is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present application. Figure 8 As shown, the electronic device 8 of this embodiment includes: at least one processor 80 ( Figure 8 Only one processor is shown in the figure), a memory 81, and a computer program 82 stored in the memory 81 and executable on at least one processor 80. When the processor 80 executes the computer program 82, the steps in any of the above-mentioned embodiments of the method for detecting manhole cover damage based on deep learning are implemented, for example Figure 3 Steps 310-330 are shown.
[0152] The processor 80 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0153] In some embodiments, the memory 81 may be an internal storage unit of the electronic device 8, such as a hard disk or memory of the electronic device 8. In other embodiments, the memory 81 may also be an external storage device of the electronic device 8, such as a plug-in hard disk, a smart memory card (SmartMediaCard, SMC), a secure digital (SecureDigital, SD) card, a flash card (FlashCard), etc. equipped on the electronic device 8.
[0154] Furthermore, the memory 81 may include both an internal storage unit of the electronic device 8 and an external storage device. The memory 81 is used to store operating devices, applications, boot loaders, data, and other programs, such as program codes of computer programs, etc. The memory 81 may also be used to temporarily store data that has been output or is to be output.
[0155] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0156] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0157] An embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0158] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in source code form, object code form, executable file or some intermediate form. The above-mentioned computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.
[0159] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0160] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0161] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the above modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0162] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0163] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for detecting manhole cover damage based on deep learning, characterized in that: include: The backbone network based on the pre-trained manhole cover detection model extracts features from the image to be detected to obtain image features; The image to be detected includes a manhole cover; Based on the neck network of the manhole cover detection model, the image features are fused to obtain target fusion features; The detection network based on the manhole cover detection model detects the target fusion feature to obtain the detection result of the manhole cover in the image to be detected; Among them, the backbone network includes m feature extraction modules of different scales, m is a positive integer, and m≥2; the feature extraction module is provided with an MSRGA submodule, the MSRGA submodule extracts and fuses multi-scale features through a channel attention mechanism and a spatial attention mechanism, and the image features are obtained based on the fused multi-scale features.
2. The method for detecting manhole cover damage according to claim 1, characterized in that: The MSRGA submodule includes a multi-scale ECA fusion block and a multi-scale ESA fusion block, and the multi-scale ECA fusion block is connected with the multi-scale ESA fusion block based on a residual structure; for the first input feature input to the MSRGA submodule: Performing an ECA-based multi-scale weighted fusion operation on the first input feature through the multi-scale ECA fusion block to obtain a first fused feature; Performing an ESA-based multi-scale weighted fusion operation on the first fusion feature through the multi-scale ESA fusion block to obtain a second fusion feature; The second fusion feature is concatenated with the first input feature through a residual structure to obtain a first output feature corresponding to the first input feature.
3. The method for detecting damage to a manhole cover according to claim 2, characterized in that: The multi-scale ECA fusion block includes a first convolutional layer, a multi-scale feature extraction structure and an ECA structure; the multi-scale feature extraction structure includes feature extraction layers at least at two scales; for the first input feature: Performing a channel compression operation on the first input feature through the first convolution layer to obtain a first convolution feature; Performing adaptive pooling operations and depth-wise separable convolution operations on the first convolutional features in sequence through the feature extraction layers at each scale, to obtain convolutional features at each scale; The ECA operation is performed on the convolution features at each scale through the ECA structure, and the first input features are weightedly fused with the obtained channel weights to obtain the first fused features.
4. The method for detecting manhole cover damage according to claim 2, characterized in that: The multi-scale ESA fusion block includes a pooling structure, a first splicing layer, a parallel convolution structure and an ESA structure; for the first fusion feature: Performing an average pooling operation and a maximum pooling operation on the first fusion feature through the pooling structure to obtain a first pooling result and a second pooling result; splicing the first pooling result and the second pooling result through the first splicing layer to obtain a first splicing feature; Performing at least two convolution operations on the first concatenated features respectively through the parallel convolution structure to obtain at least two corresponding convolution results; The ESA operation is performed on each of the convolution results through the ESA structure, and the first fusion features are weightedly fused with the obtained spatial weights to obtain the second fusion features.
5. The method for detecting damage to a manhole cover according to any one of claims 1 to 4, characterized in that: The neck network includes n feature fusion modules of different scales, where n is a positive integer and n≥2; each feature fusion module corresponds to a MSWFF module, and the MSWFF module is used to align and weightedly fuse the initial fusion features output by each feature fusion module to obtain the target fusion features corresponding to each scale.
6. The method for detecting damage to a manhole cover according to claim 5, characterized in that: The MSWFF module includes a scale alignment layer, a convolutional splicing layer, an ECA layer, and an ESA layer. At each scale, the MSWFF module performs the following steps on the initial fusion features output by each feature fusion module: Aligning the scale of each of the initial fusion features with the current scale through the scale alignment layer to obtain each of the aligned features; The convolutional splicing layer performs a convolution operation and a splicing operation on each of the alignment features to obtain a second splicing feature; Performing a weighted fusion operation based on a channel attention operation on the second spliced feature through the ECA layer to obtain a third fused feature, and splicing the second spliced feature with the third fused feature to obtain a third spliced feature; A spatial attention-based weighted fusion operation is performed on the third fusion feature operation through the ESA layer to obtain the target fusion feature at the current scale.
7. The method for detecting damage to a manhole cover according to any one of claims 1 to 4, characterized in that: The backbone network further includes a CSFF module, which includes a convolution branch, an SPD convolution branch, a second concatenation layer, an ESE layer, and a second convolution layer; for the second input feature of the CSFF module: Performing a convolution operation on the second input feature through the convolution branch to obtain a second convolution feature; Performing a space-to-depth conversion operation and a convolution operation on the second input feature through the SPD convolution branch to obtain a rearranged feature; Performing a splicing operation on the second convolutional features and the rearranged features through the second splicing layer to obtain a fourth splicing feature; Performing a global average pooling operation and a convolution operation on the fourth concatenated feature through the ESE layer, and weighting the fourth concatenated feature with the calculated channel attention weight to obtain an initial loss feature; The initial loss feature is convolved through the second convolutional layer to obtain a second output feature corresponding to the second input feature, and the image feature is obtained based on the fused multi-scale feature and the second output feature.
8. A manhole cover damage detection device, characterized in that: include: An extraction module is used to extract features from the image to be detected based on the backbone network of the pre-trained manhole cover detection model to obtain image features; The image to be detected includes a manhole cover; A fusion module, used for fusing the image features based on the neck network of the manhole cover detection model to obtain fusion features; A detection module, used to detect the fused features based on the detection network of the manhole cover detection model to obtain a detection result of the manhole cover in the image to be detected; Among them, the backbone network includes m feature extraction modules of different scales, m is a positive integer, and m≥2; the feature extraction module is provided with an MSRGA submodule, the MSRGA submodule extracts and fuses multi-scale features through a channel attention mechanism and a spatial attention mechanism, and the image features are obtained based on the fused multi-scale features.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the deep learning-based manhole cover damage detection method as described in any one of claims 1 to 7 is implemented.
10. A computer program product, wherein the computer program product stores a computer program, characterized in that: When the computer program is executed by a processor, the deep learning-based manhole cover damage detection method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Road disease identification method and device based on RD-YOLO network
CN118429329A
Inspection well cover hidden danger detection system based on YOLOV8 improved algorithm
CN118823427A
Shape memory polymer printing defect detection method, device and system and medium
CN119722622A
Fire detection method, apparatus and device based on deep learning, and medium
WO2024109873A1
Cited By
Cross-gate operation identification method and device based on deep learning, electronic equipment and program product
CN120182720A
Magnetic mineral particle detection method and system fused with attention mechanism under microscope
CN120235882A
Method and system for detecting magnetic mineral particles under microscope fusion attention mechanism
CN120235882B
Fruit appearance defect detecting and grading system
CN121033021A