Remote sensing image well lid extraction method and system based on YOLOv11
The YOLOv11-based method for manhole cover extraction in remote sensing imagery addresses precision and robustness issues by employing data augmentation and a customized network structure, achieving improved detection accuracy and efficiency in complex urban environments.
Patent Information
- Application Number
- CN202510373841.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-15
AI Technical Summary
The existing manhole cover extraction method lacks accuracy and robustness in complex scenarios, making it difficult to meet the needs of efficient and accurate urban component inspection.
The manhole cover extraction method based on YOLOv11 is adopted. By labeling, data augmenting and model training of remote sensing image data, a manhole cover detection model is built, including backbone network, neck network and head network, and key features are captured using attention mechanism to extract the manhole cover bounding box and categories.
It significantly improves the accuracy and efficiency of manhole cover detection, can adapt to different lighting, occlusion and seasonal changes, has strong generalization capabilities, and is suitable for urban component inspection in complex scenarios.
Smart Images

Figure CN120318677A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ground object target extraction, and particularly relates to a method and system for extracting manhole covers from remote sensing images based on YOLOv11. Background Art
[0002] With the acceleration of the urbanization process, the census of urban components has become an inevitable requirement for the development of modern urban management. The accurate identification and positioning of urban components (such as manhole covers, street lights, traffic signs, etc.) are important contents of urban infrastructure management. However, traditional methods for investigating urban components mainly rely on primitive measurement tools such as total stations for manual collection. This method is not only time-consuming and laborious but also difficult to meet the needs of large-scale urban management. To improve efficiency, researchers have proposed various new measurement methods, such as unmanned aerial vehicle (UAV) aerial survey, mobile measurement collection vehicles, vehicle-mounted laser scanning, etc. These methods have reduced the fieldwork workload to a certain extent, but there are still problems such as a large amount of fieldwork tasks, complex indoor processing, and limited extraction scope.
[0003] In existing object detection methods based on aerial and satellite images, such as methods based on template matching, geometric information, image segmentation, machine learning feature extraction (such as HOG features, texture features, sparse representation, Haar-like features, etc.), although these methods have achieved certain success in specific tasks, they mainly rely on shallow features and are difficult to extract deeper features, resulting in insufficient detection accuracy and robustness in complex scenarios and being difficult to meet the requirements of efficient and accurate manhole cover target extraction. Summary of the Invention
[0004] Therefore, the present invention provides a method and system for extracting manhole covers from remote sensing images based on YOLOv11 to solve the problem that the existing accuracy and robustness of manhole cover extraction are not ideal.
[0005] According to the design scheme provided by the present invention, on the one hand, a method for extracting manhole covers from remote sensing images based on YOLOv11 is provided, including:
[0006] Collecting original remote sensing image data, and performing annotation and data augmentation processing on the manhole cover bounding boxes and manhole cover class labels in the original remote sensing image data to obtain remote sensing image sample data;
[0007] Using the remote sensing image sample data to train and optimize the YOLOv11 model to obtain a manhole cover detection model, where the YOLOv11 model includes: a backbone network for extracting remote sensing image features, a neck network for using an attention mechanism to capture key features of the remote sensing image, and a head network for classifying and regressing the captured features of the remote sensing image;
[0008] Input the remote sensing image of the area to be detected into the manhole cover detection model, and use the manhole cover detection model to extract the manhole cover bounding boxes and categories in the remote sensing image of the area to be detected.
[0009] As the method for extracting manhole covers from remote sensing images based on YOLOv11 of the present invention, further, label the manhole cover bounding boxes and manhole cover category labels in the remote sensing image data, including:
[0010] Preprocess the original remote sensing image data to obtain standardized remote sensing image data, and the preprocessing includes: geometric correction processing and radiometric correction processing;
[0011] Use the labeling tool to draw bounding boxes and label category labels for the manhole cover objects in the standardized remote sensing image data, and cross-check the drawn and labeled data through multiple humans or multiple devices;
[0012] Save the checked remote sensing image data in the form of one row of data for each manhole cover, and normalize the center coordinates and sizes of the manhole cover bounding boxes in the image during the saving process.
[0013] As the method for extracting manhole covers from remote sensing images based on YOLOv11 of the present invention, further, perform data augmentation processing on the remote sensing image data, including:
[0014] Use the image data augmentation processing method to perform data augmentation processing on the original remote sensing image data to obtain augmented remote sensing image data, and the image data augmentation processing method includes but is not limited to: one or a combination of geometric transformation, color transformation, noise addition, and occlusion simulation;
[0015] Fuse the original remote sensing image data and the augmented remote sensing image data to obtain remote sensing image sample data.
[0016] As the method for extracting manhole covers from remote sensing images based on YOLOv11 of the present invention, further, use the remote sensing image sample data to train and optimize the YOLOv11 model, including:
[0017] Divide the remote sensing image sample data into a training set, a validation set, and a test set according to a preset ratio;
[0018] Use the remote sensing image data in the training set to train the YOLOv11 model, and use the remote sensing image data in the validation set and the test set to perform performance evaluation and optimization on the trained YOLOv11 model. Among them, in the performance evaluation, use specified metrics to evaluate the model performance, and the specified metrics include model accuracy, recall rate, F1 score, and mean average precision.
[0019] As the manhole cover extraction method for remote sensing images based on YOLOv11 of the present invention, further, the backbone network includes: a number of convolutional layers, C3k2 modules stacked with each convolutional layer, and an SPPF module connected to the tail C3k2 module. The C3k2 module selects a corresponding path branch for feature extraction according to the C3k parameter setting. When the C3k parameter setting is true, the first path branch is selected for feature extraction. The first path branch consists of a head and tail convolutional layer and multiple C3ks arranged in the middle of the head and tail convolutional layers. When the C3k parameter setting is false, the second path branch is selected for feature extraction. The second path branch consists of a head and tail convolutional layer and multiple Bottlenecks arranged in the middle of the head and tail convolutional layers.
[0020] As the manhole cover extraction method for remote sensing images based on YOLOv11 of the present invention, further, the neck network includes: a C2PSA module connected to the SPPF module, and two feature fusion components respectively skip-connected to the middle two C3k2 modules of the backbone network. One of the feature fusion components is connected to the C2PSA module, and the output of this feature fusion component is the input of the other feature fusion component. The feature fusion component includes an upsampling layer, a splicing layer that performs feature connection between the sampled feature and the output of the C3k2 module in the backbone network, and a second C3k2 module connected to the splicing layer. The second C3k2 module adopts the same structure as the C3k2 module in the backbone network. The other feature fusion component is connected to two C3k2 components. The output of the other feature fusion component is the input of the first C3k2 component, and the output of the first C3k2 is the input of the second C3k2 component. The C3k2 component includes a convolutional layer, a splicing layer, and a second C3k2 module. The output of the other fusion component and the output of the second C3k2 module in the two C3k2 components are respectively connected to the head network.
[0021] As the manhole cover extraction method for remote sensing images based on YOLOv11 of the present invention, further, the C2PSA module includes: head and tail convolutions, and two parallel branch paths arranged between the head and tail convolutions. Among them, the first branch path uses a convolutional kernel to retain shallow feature information for direct feature transfer, and the second branch path uses stacked PSA modules to perform multiple attention calculations to gradually optimize the feature representation to extract deep features and perform feature transfer. The features transferred by the two parallel branch paths are spliced through a connection layer and then transferred to the tail convolution for processing.
[0022] On the other hand, the present invention also provides a manhole cover extraction system for remote sensing images based on YOLOv11, including: a sample construction module, a model training module, and a target extraction module, where
[0023] A sample construction module, which is used to collect original remote sensing image data, and label and perform data augmentation on the manhole cover bounding boxes and manhole cover category labels in the original remote sensing image data to obtain remote sensing image sample data;
[0024] A model training module, which is used to train and optimize the YOLOv11 model using the remote sensing image sample data to obtain a manhole cover detection model. The YOLOv11 model includes: a backbone network for extracting remote sensing image features, a neck network for capturing key features of the remote sensing image using an attention mechanism, and a head network for classifying and regressing the captured features of the remote sensing image;
[0025] An object extraction module, which is used to input the remote sensing image of the area to be detected into the manhole cover detection model, and use the manhole cover detection model to extract the manhole cover bounding box and category in the remote sensing image of the area to be detected.
[0026] Advantages of the present invention:
[0027] The present invention constructs a manhole cover extraction model based on YOLOv11. By improving the model structure, the manhole cover detection accuracy is improved, and the data samples are enhanced. By optimizing the training strategy, the model can adapt to different lighting, occlusion, and seasonal changes, has strong generalization ability, can process high-resolution remote sensing images in real time, significantly improves the detection efficiency and accuracy, especially in complex scenarios, and has broad application prospects in the field of urban component detection and recognition. Description of the drawings
[0028] Figure 1 It is a schematic diagram of the remote sensing image manhole cover extraction process based on YOLOv11 in the embodiment;
[0029] Figure 2 It is a schematic diagram of the manhole cover detection model structure in the embodiment;
[0030] Figure 3 It is a schematic diagram of the remote sensing image manhole cover annotation result in the embodiment;
[0031] Figure 4 It is a schematic diagram of the manhole cover extraction result in the embodiment. Detailed implementation manners
[0032] To make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and technical solutions.
[0033] As an important part of urban infrastructure, the accurate identification and positioning of manhole covers are crucial for urban safety management, underground pipe network maintenance, and smart city construction. Traditional methods are difficult to meet the requirements of efficient and accurate manhole cover extraction. Therefore, in the embodiments of the present invention, refer to Figure 1As shown in the figure, a manhole cover extraction method for remote sensing images based on YOLOv11 is provided, including:
[0034] S101. Collect the original remote sensing image data, and perform annotation and data augmentation processing on the manhole cover bounding boxes and manhole cover category labels in the original remote sensing image data to obtain remote sensing image sample data.
[0035] High-resolution remote sensing image data can be obtained from public remote sensing image databases (such as Google Earth, high-resolution satellite images) or through aerial photography by drones, ensuring that the images cover areas with dense distribution of manhole covers such as urban roads, sidewalks, and parks. The collected image data should include different times (such as spring, summer, autumn, and winter), different lighting conditions (such as sunny days, cloudy days, and nights), and different occlusion situations (such as vehicle and tree occlusions) to improve the generalization ability of the model. The image data is saved in common image formats (such as JPEG, PNG), and the resolution is not less than 0.2 meters / pixel.
[0036] Specifically, the annotation of the manhole cover bounding boxes and manhole cover category labels in the remote sensing image data can be designed to include:
[0037] Preprocess the original remote sensing image data to obtain standardized remote sensing image data. The preprocessing includes: geometric correction processing and radiometric correction processing;
[0038] Use annotation tools to draw bounding boxes for manhole cover objects in the standardized remote sensing image data and label the category labels, and perform cross-checking on the drawn and labeled data by multiple people or multiple devices;
[0039] Save the checked remote sensing image data in the form of one row of data for each manhole cover, and perform normalization processing on the center coordinates and sizes of the manhole cover bounding boxes in the image during the saving process.
[0040] Annotation tools such as LabelImg and LabelMe can be used to manually annotate the manhole covers in the remote sensing images. Draw a bounding box (Bounding Box) for each manhole cover and label the category label (such as "round manhole cover", "square manhole cover"). The annotation data is saved in the YOLO format, that is, one row of data for each manhole cover, including the category number, the center coordinates (x, y) of the bounding box, the width (w), and the height (h). The coordinates and sizes are all normalized values (range [0,1]). To ensure the annotation quality, the annotation results are cross-checked by multiple annotators. GIS software can also be used to annotate the manhole covers on the remote sensing images and store each feature in the shp format, such as Figure 3 The figure shows the annotation results of the manhole covers on the image.
[0041] Among them, the data augmentation processing of the remote sensing image data can include:
[0042] The original remote sensing image data is processed by an image data augmentation method to obtain augmented remote sensing image data. The image data augmentation method includes, but is not limited to, one or more combinations of geometric transformation, color transformation, noise addition, and occlusion simulation;
[0043] The original remote sensing image data and the augmented remote sensing image data are fused to obtain remote sensing image sample data.
[0044] The labeled data is enhanced, including geometric transformation (such as random rotation, scaling, translation, flipping), color transformation (such as adjusting brightness, contrast, and saturation), noise addition (such as Gaussian noise, salt-and-pepper noise), and occlusion simulation (such as randomly adding rectangular occlusion regions). Image processing libraries such as OpenCV and Albumentations are used to implement data augmentation, expanding the original dataset to 3 - 5 times its original size, significantly improving the generalization ability of the model.
[0045] Since most of the labeled result rectangular boxes are in an inclined state, a rectangular box correction operation needs to be performed. According to the four corner point coordinates (x1, y1), (x2, y2), (x3, y3), (x4, y4) of the labeled box, the coordinate range of the labeled box is recalculated, and the calculation method is shown in Equation (1);
[0046]
[0047] Based on the corrected labeled box, the calculation of the sample annotation cropping range is carried out. With a fixed sample size of 512×512, restricted by the position of the labeled box within the sample, a random cropping method is used to achieve sample cropping. The position information of the labeled box relative to the sample is calculated and converted, and the labeled box is recorded in the YOLO label data format. Data augmentation operations are performed on the cropped samples, and the samples are enhanced by means of flipping, rotation, noise, and color transformation to obtain the required sample dataset.
[0048] S102. The YOLOv11 model is trained and optimized using the remote sensing image sample data to obtain a manhole cover detection model. The YOLOv11 model includes: a backbone network for extracting remote sensing image features, a neck network for using the attention mechanism to capture key features of the remote sensing image, and a head network for classifying and regressing the captured features of the remote sensing image.
[0049] YOLOv11 uses CSPDarknet as the backbone network, which can effectively extract multi-scale features. The YOLOv11 weights pre-trained on the COCO dataset are used to initialize the model, accelerating the training process and improving the model performance. According to the characteristics of the manhole cover detection task, the input resolution (such as 640x640), anchor box size, and number of classes of the model are adjusted.
[0050] Specifically, the training and optimization of the YOLOv11 model using remote sensing image sample data can be designed to include:
[0051] Divide the remote sensing image sample data into a training set, a validation set, and a test set according to a preset ratio;
[0052] Use the remote sensing image data in the training set to train the YOLOv11 model, and use the remote sensing image data in the validation set and the test set to evaluate and optimize the trained YOLOv11 model. Among them, in the performance evaluation, specify metrics are used to evaluate the model performance, and the specified metrics include model accuracy, recall rate, F1 score, and mean average precision.
[0053] Divide the labeled manhole cover dataset into a training set, a validation set, and a test set according to a ratio of 7:2:1, ensuring that each subset contains image data with different scenarios, different lighting conditions, and different occlusion situations. Configure a high-performance GPU (such as NVIDIA RTX 3090) for training, install the CUDA and cuDNN acceleration libraries, and implement the YOLOv11 model based on the PyTorch framework. Set the initial learning rate to 0.001, use a cosine annealing scheduler to dynamically adjust the learning rate, set the batch size to 16 - 32 according to the GPU memory, and set the number of training epochs to 100 - 200. Use TensorBoard to monitor metrics such as the loss function value, accuracy, and recall rate during the training process, save the model weights every certain number of epochs, and select the model with the best performance on the validation set as the final model.
[0054] Among them, such as Figure 2The manhole cover detection model shown above, the backbone network can be designed to include: several convolutional layers, a C3k2 module stacked with each convolutional layer, and an SPPF module connected to the tail C3k2 module. The C3k2 module selects the corresponding path branch for feature extraction according to the C3k parameter setting. If the C3k parameter is set to true, the first path branch is selected for feature extraction. The first path branch consists of a head and tail convolutional layer and multiple C3ks arranged in the middle of the head and tail convolutional layers. If the C3k parameter is set to false, the second path branch is selected for feature extraction. The second path branch consists of a head and tail convolutional layer and multiple Bottlenecks arranged in the middle of the head and tail convolutional layers. The neck network can be designed to include: a C2PSA module connected to the SPPF module, and two feature fusion components that are skip-connected to the middle two C3k2 modules of the backbone network respectively. One of the feature fusion components is connected to the C2PSA module, and the output of this feature fusion component is the input of the other feature fusion component. The feature fusion component includes an upsampling layer, a splicing layer that performs feature connection between the sampled feature and the output of the C3k2 module in the backbone network, and a second C3k2 module connected to the splicing layer. The second C3k2 module adopts the same structure as the C3k2 module in the backbone network. The other feature fusion component is connected to two C3k2 components. The output of the other feature fusion component is the input of the first C3k2 component, and the output of the first C3k2 is the input of the second C3k2 component. The C3k2 component includes a convolutional layer, a splicing layer, and a second C3k2 module. The output of the other fusion component and the output of the second C3k2 module in the two C3k2 components are respectively connected to the head network. The C2PSA module can include: head and tail convolutions, two parallel branch paths arranged between the head and tail convolutions. Among them, the first branch path uses a convolutional kernel to retain shallow feature information for direct feature transmission, and the second branch path uses stacked PSA modules to perform multiple attention calculations to gradually optimize the feature representation to extract deep features and perform feature transmission. The features transmitted by the two parallel branch paths are spliced through a connection layer and then transmitted to the tail convolution for processing.
[0055] In the backbone network, convolutional layers, a Spatial Pyramid Pooling - Fast (SPPF) block, and a Cross - Stage Partial Spatial Attention (C2PSA) block are used to extract features from the input image. The SPPF block divides the image into grids and independently extracts features from the grids, using max - pooling to process multi - scale images, which helps to maintain speed. In the neck network, the PSA module in the C2PSA module processes position - sensitive attention while processing the input tensor, which helps to selectively focus on finer details.
[0056] As Figure 2In the shown network structure, the C3K2 module is used to replace the C2F module of YOLOv8. The C3K2 module has two convolutional layers at the start and end, and multiple C3K modules in the middle. It uses a smaller 3*3 kernel to capture basic features. It improves the early version of the CSP (Corss Stage Partial) bottleneck. The small kernel series is used to process independent feature maps and merge them after convolution, thus improving feature representation.
[0057] In the head network, feature maps of three scales can be used for multi-scale prediction - small scale, medium scale, and large scale, which helps detect objects of different sizes and finally outputs class predictions and bounding boxes.
[0058] In the embodiments of this case, the introduction of new components such as C3k2 blocks and C2PSA blocks helps improve feature extraction and processing. The C3k2 module is a computationally efficient implementation of the cross-stage partial (CSP) bottleneck, which replaces the C2f blocks in the backbone and neck and uses two smaller convolutions instead of one large convolution, thus reducing the processing time. The cross-stage partial spatial attention (C2PSA) module is introduced after the SpatialPyramid Pooling–Fast (SPPF) module to enhance spatial attention. This attention mechanism enables the model to more effectively focus on important regions in remote sensing images, thus potentially improving detection accuracy.
[0059] Run the trained model on the test set, calculate metrics such as precision, recall, F1 score, and mean average precision (mAP), and generate a confusion matrix. Analyze the performance of the model in complex scenarios such as occlusion, illumination changes, and low resolution, and identify the weak links of the model. According to the evaluation results, adjust the hyperparameters of the model (such as learning rate, batch size) or the weights of the loss function, improve the model structure (such as introducing an attention mechanism in the backbone network, adding a small object detection layer), optimize the data augmentation strategy (such as adding more occlusion simulations), and use techniques such as model pruning and quantization to compress the model size and improve the inference speed.
[0060] S103: Input the remote sensing image of the area to be detected into the manhole cover detection model, and use the manhole cover detection model to extract the manhole cover bounding box and category in the remote sensing image of the area to be detected.
[0061] Input the remotely sensed image to be detected into the trained manhole cover detection model, and the model outputs the bounding box coordinates, class labels, and confidence scores of the manhole covers. Post-process the detection results, including non-maximum suppression (NMS) to remove redundant boxes and optimize the bounding boxes, so that the detection boxes fit the edge of the manhole cover more precisely. Use the extracted manhole cover information for urban component census, smart city construction, and underground pipeline network management, support real-time monitoring and anomaly detection of the manhole cover status, and improve the urban safety management level.
[0062] Furthermore, based on the above method, an embodiment of the present invention also provides a manhole cover extraction system for remotely sensed images based on YOLOv11, including: a sample construction module, a model training module, and a target extraction module, where,
[0063] The sample construction module is used to collect the original remotely sensed image data, and perform annotation and data augmentation processing on the manhole cover bounding boxes and manhole cover class labels in the original remotely sensed image data to obtain remotely sensed image sample data;
[0064] The model training module is used to train and optimize the YOLOv11 model using the remotely sensed image sample data to obtain a manhole cover detection model. The YOLOv11 model includes: a backbone network for extracting the features of the remotely sensed image, a neck network for using the attention mechanism to capture the key features of the remotely sensed image, and a head network for classifying and regressing the captured features of the remotely sensed image;
[0065] The target extraction module is used to input the remotely sensed image of the area to be detected into the manhole cover detection model, and use the manhole cover detection model to extract the manhole cover bounding box and class in the remotely sensed image of the area to be detected.
[0066] To verify the effectiveness of the solution of this case, the following further explains with experimental data:
[0067] Obtain high-resolution remotely sensed image data through drone aerial photography, covering manhole cover distribution areas such as urban roads and sidewalks. Manually annotate the manhole covers in the remotely sensed images using annotation tools. Divide the annotated manhole cover dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1 to ensure that each subset contains image data with different scenes, different lighting conditions, and different occlusion situations. Configure a high-performance GPU (such as NVIDIA RTX 3090) for training, install the CUDA and cuDNN acceleration libraries, and implement the YOLOv11 model based on the PyTorch framework. Set the initial learning rate to 0.001, use the cosine annealing scheduler to dynamically adjust the learning rate, set the batch size to 16 - 32 according to the GPU memory, and set the number of training epochs to 100 - 200. Use TensorBoard to monitor indicators such as the loss function value, accuracy, and recall rate during the training process, save the model weights every certain number of epochs, and select the model with the best performance on the validation set as the final model.
[0068] Run the trained model on the test set, calculate metrics such as Precision, Recall, F1-score, and mean Average Precision (mAP), and generate a confusion matrix. Analyze the performance of the model in complex scenarios such as occlusion, illumination changes, and low resolution to identify the weak points of the model. According to the evaluation results, adjust the hyperparameters of the model (such as learning rate, batch size) or the weights of the loss function, improve the model structure (such as introducing an attention mechanism in the backbone network, adding a small object detection layer), optimize the data augmentation strategy (such as adding more occlusion simulations), and use techniques such as model pruning and quantization to compress the model size and improve the inference speed. Use the training set, test set, and validation set to train and optimize the performance evaluation of the model to obtain a model for manhole cover detection on remote sensing images.
[0069] Based on the trained manhole cover extraction model, use the anchor box mechanism to generate candidate target regions, and predict the target category and bounding box position through the classification branch and regression branch respectively. Combine the adaptive Non-Maximum Suppression (NMS) algorithm to optimize the screening process of overlapping targets. The automatic extraction result of the manhole cover target is as Figure 4 shown, which can significantly improve the detection accuracy and reduce false detections and missed detections. After the detection is completed, the target bounding box coordinates can be converted to the geographic coordinate system through the post-processing module, and spatial analysis and visualization can be carried out in combination with Geographic Information System (GIS) data to further realize the spatialization and practical application of the detection results. It not only solves the key technical problems such as multi-scale, multi-modal, and complex background in remote sensing image target detection, but also realizes the practical application value of the detection results through spatial projection and GIS integration.
[0070] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0071] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0072] The units and method steps of each example described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0073] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.
[0074] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A manhole cover extraction method for remote sensing images based on YOLOv11, characterized in that, Including: Collecting original remote sensing image data, and annotating and data augmenting the manhole cover bounding boxes and manhole cover class labels in the original remote sensing image data to obtain remote sensing image sample data; Training and optimizing the YOLOv11 model using the remote sensing image sample data to obtain a manhole cover detection model, where the YOLOv11 model includes: a backbone network for extracting features of remote sensing images, a neck network for capturing key features of remote sensing images using an attention mechanism, and a head network for classifying and regressing the captured features of remote sensing images; Inputting the remote sensing image of the area to be detected into the manhole cover detection model, and using the manhole cover detection model to extract the manhole cover bounding box and class in the remote sensing image of the area to be detected.
2. The method for extracting manhole covers from remote sensing images based on YOLOv11 according to claim 1, characterized in that, Annotating the manhole cover bounding boxes and manhole cover class labels in the remote sensing image data, including: Preprocessing the original remote sensing image data to obtain standardized remote sensing image data, where the preprocessing includes: geometric correction processing and radiometric correction processing; Using an annotation tool to draw bounding boxes and annotate class labels for manhole cover objects in the standardized remote sensing image data, and cross-checking the drawn and annotated data by multiple humans or multiple devices; Saving the checked remote sensing image data in the form of one line of data for each manhole cover, and normalizing the center coordinates and sizes of the manhole cover bounding boxes in the image during the saving process.
3. The manhole cover extraction method for remote sensing images based on YOLOv11 according to claim 1 or 2, characterized in that, Data augmenting the remote sensing image data, including: Data augmenting the original remote sensing image data using an image data augmentation method to obtain augmented remote sensing image data, where the image data augmentation method includes, but is not limited to, one or a combination of geometric transformation, color transformation, noise addition, and occlusion simulation; Fusing the original remote sensing image data and the augmented remote sensing image data to obtain remote sensing image sample data.
4. The method for extracting manhole covers from remote sensing images based on YOLOv11 according to claim 1, characterized in that, Training and optimizing the YOLOv11 model using the remote sensing image sample data, including: Dividing the remote sensing image sample data into a training set, a validation set, and a test set according to a preset ratio; Training the YOLOv11 model using the remote sensing image data in the training set, and evaluating and optimizing the trained YOLOv11 model using the remote sensing image data in the validation set and the test set. Among them, in the performance evaluation, the performance of the model is evaluated using specified metrics, and the specified metrics include model accuracy, recall rate, F1 score, and mean average precision.
5. The method for extracting manhole covers from remote sensing images based on YOLOv11 according to claim 1, characterized in that, The backbone network includes: several convolutional layers, a C3k2 module stacked with each convolutional layer, and an SPPF module connected to the tail C3k2 module. The C3k2 module selects the corresponding path branch for feature extraction according to the C3k parameter setting. If the C3k parameter setting is true, the first path branch is selected for feature extraction. The first path branch consists of a head and tail convolutional layer and multiple C3ks arranged in the middle of the head and tail convolutional layers. If the C3k parameter setting is false, the second path branch is selected for feature extraction. The second path branch consists of a head and tail convolutional layer and multiple Bottlenecks arranged in the middle of the head and tail convolutional layers.
6. The method for extracting manhole covers from remote sensing images based on YOLOv11 according to claim 5, characterized in that, The neck network includes: a C2PSA module connected to the SPPF module, and two feature fusion components respectively skip-connected to two C3k2 modules in the middle of the backbone network. One of the feature fusion components is connected to the C2PSA module, and the output of this feature fusion component is the input of the other feature fusion component. The feature fusion component includes an upsampling layer, a splicing layer for feature connection of the sampled feature and the output of the C3k2 module in the backbone network, and a second C3k2 module connected to the splicing layer. The second C3k2 module adopts the same structure as the C3k2 module in the backbone network. The other feature fusion component is connected with two C3k2 components. The output of the other feature fusion component is used as the input of the first C3k2 component, and the output of the first C3k2 is used as the input of the second C3k2 component. The C3k2 component includes a convolutional layer, a splicing layer and a second C3k2 module. The output of the other fusion component and the output of the second C3k2 module in the two C3k2 components are respectively connected to the head network.
7. The method for extracting manhole covers from remote sensing images based on YOLOv11 according to claim 6, wherein, The C2PSA module includes: head and tail convolutions, and two parallel branch paths arranged between the head and tail convolutions. Among them, the first branch path uses a convolution kernel to retain shallow feature information for direct feature transfer, and the second branch path uses stacked PSA modules to perform multiple attention calculations to gradually optimize the feature representation to extract deep features and perform feature transfer. The features transferred by the two parallel branch paths are spliced through a connection layer and then transferred to the tail convolution for processing.
8. A manhole cover extraction system for remote sensing images based on YOLOv11, characterized in that, It includes: a sample construction module, a model training module, and a target extraction module, where The sample construction module is used to collect original remote sensing image data, and label and perform data augmentation processing on the manhole cover bounding box and manhole cover category label in the original remote sensing image data to obtain remote sensing image sample data; The model training module is used to train and optimize the YOLOv11 model using the remote sensing image sample data to obtain a manhole cover detection model. The YOLOv11 model includes: a backbone network for extracting remote sensing image features, a neck network for capturing key features of the remote sensing image using an attention mechanism, and a head network for classifying and regressing the captured features of the remote sensing image; The target extraction module is used to input the remote sensing image of the area to be detected into the manhole cover detection model, and use the manhole cover detection model to extract the manhole cover bounding box and category in the remote sensing image of the area to be detected.
9. An electronic device, characterized in that, It includes: At least one processor, and a memory coupled to the at least one processor; Among them, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, it can implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
VHR remote sensing image road intersection detection method based on YOLO11
CN121725219A