A yolov1-based vhr remote sensing image road intersection detection method
Patent Information
- Application Number
- CN202511860110.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-12-10
AI Technical Summary
[0007](1)基于纯CNN架构的YOLO系列模型(YOLOv5、YOLOv8等模型)可以处理大范围遥感影像,但识别影像中道路交叉口的精度不高
[0037] (1) A RIS-42k dataset was constructed: A RIS-42k dataset containing 42,372 road intersection samples from remote sensing images of diverse scenes was constructed. Based on this dataset, extensive comparative adaptation and ablation studies were conducted to fully verify the significant advantages of this invention in terms of detection accuracy and efficiency at road intersections.
Smart Images

Figure CN121725219B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-resolution remote sensing image recognition technology, specifically to a method for detecting road intersections in VHR remote sensing images based on YOLO11. Background Technology
[0002] Road intersections are key hubs in urban road networks, serving as core areas for the convergence and dispersal of vehicle and pedestrian traffic. Rapid and accurate identification of road intersections from high-resolution remote sensing images (VHR images) is of significant practical importance for map updates, intelligent navigation, and traffic management. However, identifying intersections from high-resolution remote sensing images faces the following challenges: the diverse shapes and structures of road intersections (e.g., T-shaped, cross-shaped, Y-shaped, roundabouts, etc.), and significant differences in intersection dimensions across different regions (urban areas, suburbs, etc.); remote sensing images are also susceptible to interference from dynamic lighting changes, shadows, vegetation cover, building obstructions, and other factors.
[0003] In early research on road intersection recognition in remote sensing images, traditional image processing techniques and manually designed features were primarily relied upon. Cai Hongyue et al., in "Automatic Extraction of Road Intersections from High-Resolution Remote Sensing Images," proposed first extracting candidate intersection regions using gradient calculation and morphological operations, then combining angle, texture, and other information for further selection. This method achieved certain detection results in specific scenarios, but its limited generalization ability was constrained by the singularity of manually designed features and its sensitivity to changes in lighting and noise. Ravanbakhsh et al., in "Knowledge-based road junction extraction from high-resolution aerial images," proposed introducing geospatial prior information to construct a circular search area and extracting intersections using geometric rules. This method could improve detection efficiency by leveraging existing geographic data, but it still relied on manually set rules and lacked generalization adaptability in different regions. Wang Long et al., in "Road Intersection Detection and Recognition Based on DPM Model in High-Resolution Remote Sensing Images," used a Deformable Part Model (DPM) for intersection recognition. However, due to the complexity of its manually designed features and the large number of model parameters, its detection performance was limited and its computational cost was high. In summary, while these traditional methods can achieve intersection localization under specific conditions or with localized data, they are generally limited by the constraints of handcrafted features, rules, or models, making it difficult to adapt to diverse remote sensing image scenarios and large-scale application needs.
[0004] To overcome the limitations of traditional methods in feature representation, researchers have introduced two-stage object detection algorithms based on convolutional neural networks (CNNs). Among them, the Faster R-CNN method, a representative of two-stage object detection methods, divides the detection task into two steps: first, candidate regions are generated, and then these candidate regions are classified and regressed, enabling precise target localization and recognition. The Mask R-CNN method adds a fully convolutional branch (ROIAlign) to the Faster R-CNN method to further achieve instance segmentation. These two-stage object detection methods have achieved good results in general object detection. In the road intersection detection task, the Faster R-CNN and Mask R-CNN methods have achieved good generalization and accuracy. However, the high overhead of processing candidate regions one by one results in slow detection speed, making it difficult to meet the needs of real-time batch processing of remote sensing images.
[0005] To improve model performance while reducing complexity and parameter count, lightweight network architectures and attention mechanisms have become research hotspots. MobileNet uses depthwise separable convolutions, ShuffleNet uses grouped convolutions to reduce computation, and GhostNet performs linear transformations on a small number of generated feature maps; these lightweight models reduce computation and model parameters. Furthermore, attention mechanisms allow models to focus more on specific regions or channels. Channel attention modules (SENet), spatial attention modules (CBAM), and Transformer-based self-attention mechanisms all show potential for enhancing feature recognition. Vision Transformer (ViT) and Swin Transformer models demonstrate the powerful capabilities of the Transformer architecture in vision tasks. The Swin Transformer model, through window self-attention and hierarchical design, effectively reduces computational complexity, making it more suitable for processing high-resolution remote sensing images.
[0006] To address the speed bottleneck of two-stage detection methods, detection algorithms, represented by the YOLO (You Only Look Once) model, transform object detection into an end-to-end regression problem, directly predicting the object category and location simultaneously across the entire image, thus achieving faster detection speeds. It has also achieved good results in ground object recognition for autonomous driving and drones (general object detection tasks). Subsequent researchers have further improved the accuracy and efficiency of the YOLO series models by introducing Feature Pyramid Network (FPN), anchor-free mechanisms, and improving the loss function and attention mechanism. However, directly applying the general YOLO series models to road intersection detection in remote sensing images often fails to achieve optimal results, thus requiring targeted optimization of the network structure. Researchers have made numerous improvements to the YOLO series models. For example, Shao Xiaomei et al., in "Improved YOLOv3 Algorithm for Automatic Recognition of Road Intersections in Remote Sensing Images," introduced the CIoU loss function and CBAM attention to achieve road intersection detection, but this model has a large number of parameters, reaching 32M. In his paper "Road Intersection Detection in Remote Sensing Imagery Using a Lightweight YOLOv7 Model Improved with Gaussian Wasserstein Distance," Kang Chuanli reduced missed detections by introducing Gaussian Wasserstein distance and a 3D attention mechanism. However, this model suffers from excessively high parameter and computational costs. In summary, while the YOLO (YouOnly Look Once) series of models strikes a good balance between accuracy and speed and has been widely applied in object detection, directly applying the YOLO model to road intersection detection in high-resolution remote sensing imagery presents several problems that urgently need to be addressed.
[0007] (1) YOLO series models based on pure CNN architecture (YOLOv5, YOLOv8, etc.) can process large-scale remote sensing images, but their accuracy in identifying road intersections in the images is not high. The YOLO11 model, which adds a self-attention mechanism to the pure CNN architecture, improves the accuracy of identifying road intersections in the images, but the computational complexity of the self-attention module in the C2PSA module of the YOLO11 model increases quadratically with the feature map size, resulting in low efficiency in processing large-scale remote sensing images.
[0008] (2) The existing lightweight YOLO11 model has a contradiction between the number of parameters and accuracy, making it difficult to meet the needs of fast and accurate processing of large-scale remote sensing images;
[0009] (3) The positioning accuracy of the center of the road intersection in the remote sensing image identified by the general YOLO11 model is low. Summary of the Invention
[0010] This invention aims to propose a road intersection detection method based on YOLO11 in VHR remote sensing images, so as to improve the recognition accuracy while reducing the amount of computation.
[0011] A method for detecting road intersections in VHR remote sensing images based on YOLO11 includes the following steps:
[0012] Step S1: Construct an image dataset containing remote sensing images of road intersections with diverse scenes, and randomly divide the image dataset into training set, validation set and test set according to the proportions.
[0013] Step S2: Construct a GSEYOLO model based on the YOLO11 model, wherein the GSEYOLO model includes a BackBone network, a Neck network, and a Head network;
[0014] The Backbone network includes the GSConv module, GS3k2 module, SPPF module, and Swin-C2PSA module, which are used to extract multi-scale features from the input remote sensing image and generate a multi-scale feature map.
[0015] The Neck network includes the GSConv module, the GS3k2 module, upsampling and concatenation operations, which are used to fuse and enhance the multi-scale feature maps generated by the Backbone network to obtain multi-scale fused feature maps;
[0016] The Head network is used to receive the multi-scale fused feature map output by the Neck network, predict the features at each scale, and output the target category, confidence score, and bounding box coordinates. The bounding box coordinates include the x-coordinate of the bounding box center point, the y-coordinate of the bounding box center point, the bounding box width w, and the bounding box height h.
[0017] Step S3: Construct the EIoU loss function for model training;
[0018] Step S4: Using the data in the image dataset constructed in step S1, train the GSEYOLO model based on the EIoU loss function and verify the accuracy of the model extraction; use the trained GSEYOLO model to complete the detection of road intersections in the remote sensing image and obtain the detection results, which include the category, confidence level and bounding box coordinates of the target, i.e., the road intersection.
[0019] Furthermore, the Backbone network specifically includes a GSConv module, four groups of modules consisting of stacked GSConv modules and GS3k2 modules, an SPPF module, and a Swin-C2PSA module connected in sequence. The second GS3k2 module, the third GS3k2 module, and the Swin-C2PSA module output a high-resolution feature map B3, a medium-resolution feature map B4, and a low-resolution feature map B5, respectively, which are then input into the Neck network.
[0020] Furthermore, the GSConv module includes a Conv module, a depthwise separable convolution DWConv module, and a Shuffle module connected in sequence. The output of the Conv module is concatenated with the output of the depthwise separable convolution DWConv module, and the concatenated result is used as the input of the Shuffle module.
[0021] Furthermore, the GS3k2 module includes a GSConv module, a Split layer, n C3GS modules, and a GSConv module connected in sequence. The feature map generated by the GSConv module is divided into two parts, feature map Split_1 and feature map Split_2, after passing through the Split layer. Feature map Split_2 is then subjected to deep feature extraction by the n C3GS modules and concatenated with feature map Split_1. The concatenated result is used as the input of the last GSConv module.
[0022] Furthermore, the C3GS module uses two Conv modules as the main branch and the bypass branch, respectively. At least one GSBN module is stacked in the main branch for network depth expansion, and the output of the bypass branch directly participates in the concatenation. The output of the last GSBN module in the main branch is concatenated with the output of the bypass branch, and the concatenated result is then passed through a Conv module for feature fusion to obtain the output of the C3GS module.
[0023] Furthermore, the GSBN module is a residual structure composed of two GSConv modules. Through residual connection, the input features are added to the output features of the second GSConv module to obtain the final output of the GSBN module.
[0024] Furthermore, the Swin-C2PSA module includes a Conv module, a Split layer, n SW-PSA modules, and a Conv module connected in sequence. The feature map generated by the first Conv module is split into feature map Split_A and feature map Split_B after passing through the Split layer. Feature map Split_B is then processed by the n SW-PSA modules for deep feature extraction and concatenated with feature map Split_A. The concatenated result is used as the input of the last Conv module.
[0025] Furthermore, the SW-PSA module sequentially integrates an LN layer, an SW-MSA module, an LN layer, and an MLP module. The input of the first LN layer is added to the output of the SW-MSA module to serve as the input of the second LN layer. The input of the second LN layer is added to the output of the MLP module to serve as the final output of the SW-PSA module. The SW-MSA module is used to calculate self-attention features within a regular window or a shifted window. The MLP module is used for channel adjustment.
[0026] Furthermore, the process of the Neck network processing images is as follows:
[0027] a. After upsampling feature map B5, it is concatenated with feature map B4, and information is fused using the GS3k2 module to generate a medium-resolution feature map N4 that contains both deep semantic information and details.
[0028] b. After upsampling, feature map N4 is concatenated with feature map B3, and information is fused using the GS3k2 module to generate a high-resolution feature map N3 that contains both deep semantic information and details.
[0029] c. After feature map N3 is processed by the GSConv module, it is concatenated with feature map N4 and information is fused using the GS3k2 module to generate a medium-resolution feature map N4+ that contains both deep semantic information and details.
[0030] d. After the feature map N4+ is processed by the GSConv module, it is concatenated with the feature map B5 and then fused with information using the GS3k2 module to generate a low-resolution feature map N5 that contains both deep semantic information and details.
[0031] Furthermore, in step S3, the EIoU loss function is calculated using the following formula: ,
[0032] in, Represents the EIoU loss function;
[0033] This represents the standard crossover ratio loss. , This represents the intersection-union ratio (IUU) between the predicted bounding box and the ground truth bounding box. and These represent the predicted bounding box and the ground truth bounding box, respectively.
[0034] This indicates the loss due to the distance from the center point. , Indicates the center point of the prediction box. Represents the center point of the true bounding box. Indicates the calculation center point and center point The Euclidean distance between them It is the diagonal length of the smallest bounding rectangle that covers both the predicted bounding box and the ground truth bounding box;
[0035] This indicates the loss due to the difference in length and width. , Indicates the width of the prediction box. This represents the width of the actual bounding box. Indicates the height of the predicted bounding box. Represents the height of the true bounding box; length and width difference loss. Directly penalize the width of the prediction box Width of the actual frame The differences between them, and the height of the predicted box Height of the actual frame The differences between them; and These are the width and height of the smallest bounding rectangle, respectively.
[0036] The present invention has the following advantages:
[0037] (1) A RIS-42k dataset was constructed: A RIS-42k dataset containing 42,372 road intersection samples from remote sensing images of diverse scenes was constructed. Based on this dataset, extensive comparative adaptation and ablation studies were conducted to fully verify the significant advantages of this invention in terms of detection accuracy and efficiency at road intersections.
[0038] (2) A Swin-C2PSA module adapted to remote sensing images was designed: This invention integrates the shifted window attention of the Swin Transformer with the C2PSA module of the YOLO11 model, proposing the Swin-C2PSA module. This module restricts the computation of self-attention to within a window, allowing the model to directly process large-scale remote sensing images without the need for block processing. At the same time, this module also enhances the comprehensive perception capability of local details and global contextual features of intersections, improving recognition accuracy while controlling computational overhead.
[0039] (3) Constructing an efficient and lightweight network: The GSConv module was introduced and combined with the CSP structure of the YOLO11 model. The C3GS module and GS3k2 module were designed, which significantly reduced the number of network parameters and computational complexity, while enhancing the richness of feature representation, thereby achieving a simultaneous improvement in the efficiency and accuracy of remote sensing data processing.
[0040] (4) Improved the positioning accuracy of the center point of the road intersection: The EIoU loss function (Efficient IoU) is introduced. This function takes into account the overlap of the bounding box, the distance between the center points and the difference in length and width, which effectively improves the positioning accuracy of the center of the road intersection. Attached Figure Description
[0041] Figure 1 This is a schematic diagram of the structure of the GSEYOLO model constructed in the VHR remote sensing image road intersection detection method in this embodiment of the invention.
[0042] Figure 2 This is a schematic diagram of the structure of the GSConv module and GS3k2 module in the GSEYOLO model in this embodiment of the invention.
[0043] Figure 3 This is a schematic diagram of the structure of the SW-PSA module and the Swin-C2PSA module in the GSEYOLO model in this embodiment of the invention.
[0044] Figure 4 This is a graph showing the verification results of the GSEYOLO model of the present invention on the ability to identify road intersections in different environments. Detailed Implementation
[0045] To better understand the above-described objects, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Many specific details are set forth in the following description to provide a thorough understanding of the invention; however, the invention may be practiced in other ways different from those described herein, and therefore, the invention is not limited to the specific embodiments disclosed below.
[0046] The present invention proposes a method for detecting road intersections in VHR remote sensing images based on YOLO11, comprising the following steps:
[0047] Step S1, Dataset Construction:
[0048] Constructing the RIS-42k dataset: This invention constructs an image dataset containing 42,372 remote sensing images of road intersections in diverse scenes, named the RIS-42k dataset;
[0049] Remote sensing images of 400×400 pixels were cropped from the center point of road intersections on Google Earth. These images covered various environments, including urban, suburban, and rural areas, and included cross-shaped, T-shaped, X-shaped, Y-shaped, and multi-way intersections. A 100-pixel area around the center point was considered the intersection's extent. These cropped images were used to construct the road intersection dataset of this invention, named the RIS-42k dataset, containing 42,372 remote sensing image samples. The annotation information for the remote sensing image samples followed YOLO format requirements. Each bounding box contained a category label ("intersection") and four normalized coordinate values (x-coordinate of the bounding box center point, y-coordinate of the bounding box center point, bounding box width w, and bounding box height h).
[0050] The labeled RIS-42k dataset was randomly divided into training, validation, and test sets in an 8:1:1 ratio, which were used for model training, hyperparameter tuning, and final performance evaluation, respectively.
[0051] Step S2, construct the GSEYOLO model based on the YOLO11 model:
[0052] This invention is based on the YOLO11 model. First, it combines the Shift-Window Multi-Head Self-Attention (SW-MSA) module from the Swing Transformer with the C2PSA module from the YOLO11 model to design a Swing-C2PSA module, which replaces the C2PSA module in the YOLO11 model. Second, it introduces the GSConv module and designs the GS3k2 module. Finally, it replaces the CIoU loss function in the original YOLO11 with the EIoU loss function, constructing a YOLO11-based detection model for road intersections in VHR remote sensing images—the GSEYOLO model. The main structure of the GSEYOLO model is as follows: Figure 1 As shown, the GSEYOLO model mainly consists of three parts: the Backbone network, the Neck network, and the Head network.
[0053] In this invention, the Backbone network consists of the GSConv module, GS3k2 module, SPPF module, and Swin-C2PSA module, primarily used to extract multi-scale features from the input remote sensing image and obtain a multi-scale feature map. The Backbone network in the GSEYOLO model constructed in this invention is restructured into a lightweight and efficient feature extractor. Its core lies in introducing the GSConv module to replace the Conv module in the YOLO11 model. The GSConv module, through decomposition and recombination, significantly reduces computational costs while maintaining feature representation capabilities. Figure 1 As shown in (a), the Backbone network specifically includes a GSConv module, four groups of modules consisting of stacked GSConv modules and GS3k2 modules, an SPPF module, and a Swin-C2PSA module connected in sequence. The second GS3k2 module, the third GS3k2 module, and the Swin-C2PSA module output a high-resolution feature map B3, a medium-resolution feature map B4, and a low-resolution feature map B5, respectively, and are input into the Neck network.
[0054] like Figure 2 As shown in (a), the GSConv module includes a Conv module, a depthwise separable convolution DWConv module, and a Shuffle module connected in sequence. The output of the Conv module is concatenated with the output of the depthwise separable convolution DWConv module, and the concatenated result is used as the input to the Shuffle module. The computation process of the GSConv module can be decomposed as follows: a. Use a standard convolution (i.e., Conv module) to generate 1 / 2 channel intrinsic features; b. Based on inherent features, use the computationally inexpensive depthwise separable convolution DWConv to generate 1 / 2 channel features. c. The inherent features of 1 / 2 channel and the cheap features of 1 / 2 channel are concatenated, and the concatenated features are shuffled by channel mixing through the Shuffle module to promote information interaction; finally, the interactive and recombined feature map is obtained.
[0055] This invention builds upon the GSConv module to construct the GS3k2 module. A schematic diagram of the GS3k2 module is shown below. Figure 2 As shown.
[0056] like Figure 2As shown in (d), the GS3k2 module includes a GSConv module, a Split layer, n C3GS modules, and a GSConv module connected in sequence. The feature map generated by the GSConv module is divided into two parts, feature map Split_1 and feature map Split_2, after passing through the Split layer. Feature map Split_2 undergoes depth feature extraction by the n C3GS modules and is then concatenated with feature map Split_1. The concatenated result is used as the input to the last GSConv module. The image processing procedure of the GS3k2 module is as follows: a. The input features of the GS3k2 module first pass through a GSConv module for channel adjustment. The output of the GSConv module is split into two parts, feature map Split_1 and feature map Split_2, after passing through a Split layer. b. Feature map Split_1 is directly retained as a "shortcut"; c. Feature map Split_2 is obtained by deep feature extraction through a bottleneck structure composed of n stacked C3GS modules; d. Finally, feature map Split_1 and feature map Split_2+ are concatenated, and the concatenated result is then fused through a GSConv module to obtain the final output.
[0057] The structure of the C3GS module mentioned in the above processing is as follows: Figure 2 As shown in (c), the C3GS module uses two Conv modules as the main branch and the bypass branch, respectively. At least one GSBN module is stacked in the main branch for network depth expansion, and the output of the bypass branch directly participates in the concatenation. The output of the last GSBN module in the main branch is concatenated with the output of the bypass branch. Finally, the concatenated result is passed through a Conv module for feature fusion to obtain the output of the C3GS module.
[0058] The structure of the GSBN module mentioned in the above processing is as follows: Figure 2 As shown in (b), the GSBN module is a residual structure based on the GSConv module. This residual structure consists of two GSConv modules. Through residual connection, the input features are added to the output features of the second GSConv module, which enhances gradient propagation and reduces information loss.
[0059] The SPPF module in this invention is a built-in module of the YOLO11 model and is not an innovative point of this invention; therefore, no schematic diagram or detailed explanation is provided.
[0060] like Figure 3As shown in (b), the Swin-C2PSA module consists of a Conv module, a Split layer, n SW-PSA modules, and a Conv module connected in sequence. The feature map generated by the first Conv module is split into two parts, Split_A and Split_B, after passing through the Split layer. Split_B undergoes depth feature extraction by the n SW-PSA modules and is then concatenated with Split_A. The concatenated result is used as the input to the last Conv module. The image processing procedure of the Swin-C2PSA module is as follows: a. The input features of the Swin-C2PSA module first pass through a Conv module for channel adjustment. The output of the Conv module is then split into two parts, feature map Split_A and feature map Split_B, after passing through a Split layer. b. Feature map Split_A is directly retained as a "shortcut"; c. Feature map Split_B is processed through a bottleneck structure consisting of n stacked SW-PSA modules to extract deep features, resulting in feature map Split_B+. d. Finally, feature map Split_A and feature map Split_B+ are concatenated and then fused through a Conv module to obtain the final output.
[0061] The structure of the SW-PSA module mentioned in the above processing is as follows: Figure 3 As shown in (a) above, the SW-PSA module sequentially integrates an LN layer (layer normalization), a shifted window multi-head self-attention module (SW-MSA module), an LN layer, and an MLP module. The input of the first LN layer is added to the output of the SW-MSA module to serve as the input of the second LN layer; the input of the second LN layer is added to the output of the MLP module to serve as the final output of the SW-PSA module. The SW-MSA module, or Shifted Window Multi-head Self-Attention module, is used to calculate the self-attention features within a regular window or a shifted window. The MLP module is used for channel adjustment.
[0062] Through this computational method, the computational cost of the Backbone network in this invention is lower than that of the Backbone network in the original YOLO11 model, thus achieving a lightweight Backbone network in the entire model. Simultaneously, the introduction of window self-attention in the Swin-C2PSA module enables this invention to directly process large-size remote sensing images.
[0063] In this invention, the Neck network consists of a GSConv module, a GS3k2 module, an upsampling and stitching module, and adopts a fusion structure of a path aggregation network and a feature pyramid network. It is mainly used to fuse and enhance the multi-scale feature maps generated by the Backbone network. For example... Figure 1 As shown, after the input remote sensing image is processed by the Backbone network, a high-resolution feature map B3, a medium-resolution feature map B4, and a low-resolution feature map B5 are obtained. Figure 1 As shown in (b) above, the Neck network processes images as follows: a. After upsampling, feature map B5 is concatenated with feature map B4, and information is fused using the GS3k2 module to generate a medium-resolution feature map N4 that contains both deep semantic information and details. b. After upsampling, feature map N4 is concatenated with feature map B3, and information is fused using the GS3k2 module to generate a high-resolution feature map N3 that contains both deep semantic information and details. c. After feature map N3 is processed by the GSConv module, it is concatenated with feature map N4 and information is fused using the GS3k2 module to generate a medium-resolution feature map N4+ that contains both deep semantic information and details. d. After the feature map N4+ is processed by the GSConv module, it is concatenated with the feature map B5 and then fused with information using the GS3k2 module to generate a low-resolution feature map N5 that contains both deep semantic information and details.
[0064] After processing by the Neck section, we obtain high-resolution feature maps N3, medium-resolution feature maps N4+, and low-resolution feature maps N5, which contain both deep semantic information and details.
[0065] In this invention, the Head network receives high-resolution feature map N3, medium-resolution feature map N4+, and low-resolution feature map N5 from the Neck network, and predicts the features at each scale, outputting the target's category, confidence level, and bounding box coordinates, thereby completing the detection of road intersections in remote sensing images, such as... Figure 1 As shown in (c) of the diagram. In this invention, the Head section simply replaces the CIoU loss function in the original YOLO11 model with the EIoU loss function (Efficient IoU). The composition of the EIoU loss function will be detailed in step S3.
[0066] Step S3, construct the loss function for model training:
[0067] To improve detection accuracy, this invention introduces an EIoU loss function based on the YOLO11 model loss function, replacing the CIoU loss function in the original YOLO11 model. The EIoU loss function consists of three parts. The calculation formula is as follows: ,
[0068] in:
[0069] This represents the standard crossover-union ratio loss, focusing on the overlapping area. , This represents the intersection-union ratio (IUU) between the predicted bounding box and the ground truth bounding box. and These represent the predicted bounding box and the ground truth bounding box, respectively.
[0070] This represents the distance loss to the center point. , Indicates the center point of the prediction box. Represents the center point of the true bounding box. Indicates the calculation center point and center point The Euclidean distance between them It is the diagonal length of the smallest bounding rectangle that covers the predicted bounding box and the ground truth bounding box.
[0071] This indicates the loss due to the difference in length and width. , Indicates the width of the prediction box. This represents the width of the actual bounding box. Indicates the height of the predicted bounding box. Represents the height of the true bounding box; length and width difference loss. This is a key improvement of EIOU, which directly penalizes the width of the predicted bounding box. Width of the actual frame The differences between them, and the height of the predicted box Height of the actual frame The differences between them; and These are the width and height of the smallest bounding rectangle, respectively.
[0072] Compared to the CIoU loss function, which only considers aspect ratio as a blur penalty, the EIoU loss function more comprehensively measures the difference between the predicted bounding box and the ground truth bounding box. The EIoU loss function includes a loss based on aspect ratio difference. It can provide more direct and clear gradients to guide the model in optimizing the size of the bounding box, thereby significantly improving the localization accuracy.
[0073] Step S4: Use data from the RIS-42k dataset to train the GSEYOLO model; after the model training is complete, use large-size remote sensing images to extract road intersections and verify the accuracy of the model extraction.
[0074] Experimental verification:
[0075] Based on the RIS-42k dataset, extensive comparative adaptation and ablation studies have fully verified the significant advantages of the GSEYOLO model of this invention in terms of detection accuracy and efficiency.
[0076] Table 1. Comparison of different models on the RIS-42k dataset:
[0077]
[0078] As shown in Table 1, the experimental results of the present invention demonstrate that the detection accuracy and efficiency of the GSEYOLO model are superior to those of the existing yolov8n, yolov8s, yolo11n, yolo11s, yolo12n, and Rtdetr-l models. Specifically, the GSEYOLO model of the present invention has a precision of 0.886 and a recall of 0.853, indicating that the GSEYOLO model effectively reduces the number of model parameters while improving the extraction accuracy of road intersections in large-scale remote sensing images.
[0079] Table 2 Ablation experiments with different module combinations:
[0080]
[0081] As shown in Table 2, the experimental results demonstrate that the SwinC2PSA module effectively overcomes the image size limitation of the traditional YOLO11 model in remote sensing image feature extraction, while simultaneously improving the perception capability of road intersections. The introduction of the GSConv and GS3k2 modules reduces the number of parameters in the YOLO11 model while further improving the extraction accuracy of road intersections. The introduction of the EIoU loss function effectively improves the localization accuracy of the center point and bounding box of road intersections. Furthermore, the three modules exhibit good low coupling and synergistic enhancement effects, jointly improving the robustness and applicability of this invention.
[0082] To verify the ability of the GSEYOLO model of this invention to identify road intersections in different environments, VHR remote sensing images from four regions—Beijing, London, New York, and Paris—were selected for experiments. The identification results are as follows: Figure 4 The results are shown in (a), (b), (c), and (d). The accuracy and recall rates are shown in Table 3.
[0083] Table 3. Accuracy and recall of the model for identifying road intersections in different environments:
[0084]
[0085] Experimental results show that the GSEYOLO model maintains high precision when facing road networks with different city styles, but the recall rate varies significantly with the city, demonstrating its generalization ability. It can be seen that regardless of... Figure 4 (a) and Figure 4 The regular road network shown in (c) is still Figure 4 (b) and Figure 4 The radial road network shown in (d) can accurately identify various types of intersections.
[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for detecting road intersections using VHR remote sensing images based on YOLO11, characterized in that, Includes the following steps: Step S1: Construct an image dataset containing remote sensing images of road intersections with diverse scenes, and randomly divide the image dataset into training set, validation set and test set according to the proportions. Step S2: Construct a GSEYOLO model based on the YOLO11 model, wherein the GSEYOLO model includes a Backbone network, a Neck network, and a Head network; The Backbone network includes the GSConv module, GS3k2 module, SPPF module, and Swin-C2PSA module, which are used to extract multi-scale features from the input remote sensing image and generate a multi-scale feature map. The Neck network includes the GSConv module, the GS3k2 module, upsampling and concatenation operations, which are used to fuse and enhance the multi-scale feature maps generated by the Backbone network to obtain multi-scale fused feature maps; The Head network is used to receive the multi-scale fused feature map output by the Neck network, predict the features at each scale, and output the target's category, confidence level, and bounding box coordinates. Step S3: Construct the EIoU loss function for model training; Step S4: Using the data in the image dataset constructed in step S1, train the GSEYOLO model based on the EIoU loss function and verify the accuracy of the model extraction; use the trained GSEYOLO model to complete the detection of road intersections in the remote sensing image and obtain the detection results. The Backbone network specifically includes a GSConv module, four groups of modules consisting of stacked GSConv and GS3k2 modules, an SPPF module, and a Swin-C2PSA module connected in sequence. The second GS3k2 module, the third GS3k2 module, and the Swin-C2PSA module output a high-resolution feature map B3, a medium-resolution feature map B4, and a low-resolution feature map B5, respectively, which are then input into the Neck network. The GSConv module includes a Conv module, a depthwise separable convolution DWConv module, and a Shuffle module connected in sequence. The output of the Conv module is concatenated with the output of the depthwise separable convolution DWConv module, and the concatenated result is used as the input of the Shuffle module. The GS3k2 module includes a GSConv module, a Split layer, n C3GS modules, and a GSConv module connected in sequence. The feature map generated by the GSConv module is divided into two parts, feature map Split_1 and feature map Split_2, after passing through the Split layer. Feature map Split_2 is then processed by the n C3GS modules for deep feature extraction and concatenated with feature map Split_1. The concatenated result is used as the input to the last GSConv module. The Swin-C2PSA module includes a Conv module, a Split layer, n SW-PSA modules, and a Conv module connected in sequence. The feature map generated by the first Conv module is split into feature map Split_A and feature map Split_B after passing through the Split layer. Feature map Split_B is then processed by the n SW-PSA modules for deep feature extraction and concatenated with feature map Split_A. The concatenated result is used as the input to the last Conv module. The C3GS module uses two Conv modules as the main branch and the bypass branch, respectively. The main branch stacks at least one GSBN module for network depth expansion, and the output of the bypass branch directly participates in the concatenation. The output of the last GSBN module in the main branch is concatenated with the output of the bypass branch. Finally, the concatenated result is passed through a Conv module for feature fusion to obtain the output of the C3GS module.
2. The method according to claim 1, characterized in that, The GSBN module is a residual structure composed of two GSConv modules. Through residual connection, the input features are added to the output features of the second GSConv module to obtain the final output of the GSBN module.
3. The method according to claim 1, characterized in that, The SW-PSA module integrates an LN layer, an SW-MSA module, an LN layer, and an MLP module in sequence. The input of the first LN layer is added to the output of the SW-MSA module to serve as the input of the second LN layer. The input of the second LN layer is added to the output of the MLP module to serve as the final output of the SW-PSA module. The SW-MSA module is used to calculate the self-attention features within a regular window or a shifted window. The MLP module is used for channel adjustment.
4. The method according to claim 3, characterized in that, The process of image processing by the Neck network is as follows: a. After upsampling feature map B5, it is concatenated with feature map B4, and information is fused using the GS3k2 module to generate a medium-resolution feature map N4 that contains both deep semantic information and details. b. After upsampling, feature map N4 is concatenated with feature map B3, and information is fused using the GS3k2 module to generate a high-resolution feature map N3 that contains both deep semantic information and details. c. After feature map N3 is processed by the GSConv module, it is concatenated with feature map N4 and information is fused using the GS3k2 module to generate a medium-resolution feature map N4+ that contains both deep semantic information and details. d. After the feature map N4+ is processed by the GSConv module, it is concatenated with the feature map B5 and then fused with information using the GS3k2 module to generate a low-resolution feature map N5 that contains both deep semantic information and details.
5. The method according to claim 4, characterized in that, In step S3, the EIoU loss function is calculated using the following formula: , in, Represents the EIoU loss function; This represents the standard crossover ratio loss. , This represents the intersection-union ratio (IUU) between the predicted bounding box and the ground truth bounding box. and These represent the predicted bounding box and the ground truth bounding box, respectively. This indicates the loss due to the distance from the center point. , Indicates the center point of the prediction box. Represents the center point of the true bounding box. Indicates the calculation center point and center point The Euclidean distance between them It is the diagonal length of the smallest bounding rectangle that covers both the predicted bounding box and the ground truth bounding box; This indicates the loss due to the difference in length and width. , Indicates the width of the prediction box. This represents the width of the actual bounding box. Indicates the height of the predicted bounding box. Represents the height of the true bounding box; length and width difference loss. Directly penalize the width of the prediction box Width of the actual frame The differences between them, and the height of the predicted box Height of the actual frame The differences between them; and These are the width and height of the smallest bounding rectangle, respectively.
Citation Information
Patent Citations
Lightweight remote sensing image target detection method fusing dynamic upsampling
CN119027817A
Remote sensing image target detection method and system based on improved YOLOv11 network
CN120932121A