Road intersection detection method used in optical remote sensing image
By introducing the Swing Transformer and structural semantic vector perception loss term into the DETR model, the detection of road intersections in optical remote sensing images is optimized, which solves the problems of insufficient detection accuracy and adaptability to complex scenes in the existing technology and achieves high-precision and robust detection results.
Patent Information
- Application Number
- CN202511169693.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies are insufficient in terms of accuracy, structural modeling capabilities, and adaptability to complex scenes for detecting road intersections in optical remote sensing images. Traditional methods are prone to missed detections and false detections, and their reliance on anchor frame mechanisms increases system complexity.
In the DETR model framework, the Swing Transformer is introduced to replace the CNN backbone network, and a joint matching loss function is constructed by combining the structure semantic vector-aware loss term to optimize the target matching process. Multi-scale feature extraction and structure-aware matching are performed through the Swing Transformer.
It improves the accuracy and robustness of road intersection detection, enhances the model's ability to generalize to complex scenarios, and achieves high-precision and high-robustness intersection detection.
Smart Images

Figure CN120953816A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method for detecting road intersections in optical remote sensing images. Background Technology
[0002] Road intersections are key nodes in urban road networks, and their geometric structure and topological relationships are of great significance in improving traffic efficiency, supporting intelligent navigation, and enabling automatic map updates. With the widespread acquisition of high-resolution optical remote sensing imagery, automatically identifying intersection targets in remote sensing images using computer vision technology has become a research hotspot in the field of intelligent remote sensing interpretation. Traditional methods typically rely on edge extraction, shape analysis, or manual rule construction, which are sensitive to noise and scale changes and have limited generalization ability. In complex backgrounds, with occlusion interference, or with structurally similar targets, they are prone to missed detections and false detections, resulting in insufficient detection robustness.
[0003] Convolutional Neural Networks (CNNs) have made significant breakthroughs in image feature modeling, driving the development of object detection algorithms. Mainstream CNN-based detection frameworks (such as Faster R-CNN, YOLO, and RetinaNet) typically rely on anchor box mechanisms to generate candidate regions and perform classification and regression operations. However, for complex, spatially dense targets at intersections, these methods still face problems such as high false detection rates and inaccurate localization, and the post-processing operations they rely on further increase system complexity.
[0004] DETR is the first model to introduce Transformer into end-to-end object detection. It incorporates learnable target query vectors and the Hungarian matching algorithm, omitting steps such as candidate box generation and non-maximum suppression, demonstrating good modeling capabilities and a simple structural design. However, the DETR model still faces the following challenges in intersection detection in remote sensing images: 1) Insufficient structure awareness: traditional CNN backbone networks struggle to capture the spatial topological information of targets, resulting in weak ability to distinguish between structurally similar intersection types; 2) Limited position encoding methods: fixed or absolute position encoding is insensitive to local structural changes, limiting the model's spatial awareness; 3) Lack of structural constraints in the matching mechanism: the Hungarian matching algorithm only calculates based on category and position costs, ignoring the structural relationships between targets, easily leading to matching errors. Therefore, it is urgent to optimize feature representation, position modeling, and target matching mechanisms to construct a robust intersection detection method with structure awareness capabilities that adapts to the characteristics of optical remote sensing images. Summary of the Invention
[0005] This invention provides a method for detecting road intersections in optical remote sensing images, addressing the technical problem that existing technologies have significant shortcomings in terms of road intersection recognition accuracy, structural modeling capabilities, and adaptability to complex scenes.
[0006] To address the above technical problems, this invention provides a method for detecting road intersections in optical remote sensing images, comprising:
[0007] S1. By introducing the Swing Transformer into the DETR model framework to replace the CNN backbone network, a network architecture for the road intersection detection model is constructed.
[0008] S2. By introducing a structural semantic vector-aware loss term into the standard Hungarian matching algorithm, a joint matching loss function containing structural consistency constraints is constructed.
[0009] S3. The road intersection detection model is trained using the training set. During training, the optimal match is calculated based on the Hungarian matching algorithm, and the total loss is calculated using the joint matching loss function for backpropagation of the matching results.
[0010] S4. The trained road intersection detection model is used to perform forward inference on the optical remote sensing image to obtain the road intersection detection results.
[0011] Furthermore, the network architecture of the road intersection detection model includes a backbone network, an encoder, a decoder, and a prediction head. The backbone network is used to extract multi-scale spatial semantic features of the image and input them into the encoder. The encoder is used to encode the multi-scale spatial semantic features to obtain global semantic features, which are then input into the decoder. The decoder is used to introduce a learnable target query vector and generate a high-dimensional representation of the potential target by interacting with the global semantic features. The prediction head is used to output the category, bounding box parameters, and structural semantic vector of each intersection target based on the high-dimensional representation.
[0012] Furthermore, the backbone network is a Swing Transformer, followed by 1*1 convolutional layers;
[0013] The Swin Transformer employs a four-stage hierarchical architecture. In the first stage, the input image, with length H, width W, and 3 channels, is first divided into 4×4 non-overlapping blocks by a block-splitting layer. Then, it is mapped to a C-dimensional feature vector through a linear embedding layer. Finally, feature extraction is performed using two Swin Transformer blocks, with an output size of [missing information]. The input feature map has C channels. In the remaining nth stage, the input feature map first passes through a block merging layer to merge adjacent 2×2 blocks, achieving spatial downsampling and channel expansion; then it is mapped to a 2×2 feature map through a linear embedding layer. n-1 *C-dimensional feature vectors are extracted using two or more Swing Transformer blocks, resulting in a final output of size C. The number of channels is 2 n-1 *High-dimensional feature maps of C, n = 2, 3, 4;
[0014] The feature maps extracted by the Swin Transformer are reduced to 256 dimensions by the 1*1 convolutional layer. After flattening, a sequence of shape [B,N,256] is obtained, where B represents the batch size and N is the number of positions after flattening.
[0015] Furthermore, each stage of the Swin Transformer block alternately employs window multi-head self-attention and shifted window attention, combined with feedforward networks, layer normalization, and residual connections to achieve the fusion of local and global features.
[0016] Furthermore, the number of Swing Transformer blocks included in the first to fourth stages are 2, 2, 6, and 2, respectively.
[0017] Furthermore, the encoder employs a 3-layer Transformer encoder, each layer of which includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and a layer normalization module, and introduces a relative position bias.
[0018] Furthermore, the decoder employs a 6-layer Transformer decoder. The feature sequence output by the encoder is fed into the 6-layer Transformer decoder. Each layer of the decoder includes a target query self-attention, cross-attention, feedforward network, residual connection, and layer normalization module. The initialized target query vector set interacts with the encoder output to generate a high-dimensional representation of the potential target.
[0019] Furthermore, the joint matching loss function is constructed as L cls(i,j) L L1(i,j) L GIoU(i,j) L str(i,j) The weighted sum, L cls(i,j) L represents the cross-entropy loss between the predicted target class and the true target class. L1(i,j) L represents the L1 regression loss of the bounding box. GIoU(i,j) L represents the generalized IoU loss, used to evaluate the geometric overlap and spatial relationship between the predicted bounding box and the ground truth bounding box; str(i,j)The structure-aware loss measures the similarity between the predicted target's structural semantic vector and the true target's structural semantic vector.
[0020] Furthermore, the prediction head is composed of a feedforward neural network (FFN), which is a three-layer perceptron with a ReLU activation function, and outputs the prediction result through a linear projection layer. The prediction head includes a category prediction branch, a bounding box regression branch, and a target structure semantic vector prediction branch: the category prediction branch is composed of a linear layer and a softmax function; the bounding box branch is composed of a three-layer MLP, used to predict the normalized center coordinates, height, and width; the target structure semantic vector prediction branch is composed of a two-layer MLP, used to map the target query output by the decoder into the spatial topological feature vector of the intersection; the true target structure semantic vector is obtained by average pooling the labeled box regions in the fourth stage feature map of the backbone network.
[0021] Furthermore, when training the road intersection detection model, a remote sensing image dataset is first acquired and preprocessed, specifically including:
[0022] Acquire a remote sensing image dataset; the dataset consists of multiple high-resolution optical remote sensing images, and the location and type of the intersection targets are marked in the images;
[0023] Preprocessing of the raw remote sensing images includes image size standardization, annotation format unification, and pixel normalization.
[0024] This invention provides a method for detecting road intersections in optical remote sensing images. To overcome the shortcomings of existing network models in feature extraction and standard Hungarian matching mechanisms in structural modeling and target matching, a road intersection detection model is constructed. This model utilizes a Swing Transformer backbone network for feature extraction to achieve deep fusion of local details and global contextual information. It also combines a lightweight encoder structure with a structure-aware matching mechanism to improve spatial feature modeling efficiency while enhancing the structural representation capability of target recognition. Furthermore, by using structural semantic vectors in the joint matching loss optimization process, it further improves target representation accuracy and structural consistency, making the model more stable in complex intersection detection scenarios. Compared with traditional CNN detection methods and the standard DETR model, the road intersection detection model constructed in this invention significantly improves detection accuracy, structural discrimination ability, and robustness, making it suitable for high-precision remote sensing detection applications of various types of road intersections. Attached Figure Description
[0025] Figure 1 This is a flowchart of a method for detecting road intersections in optical remote sensing images provided by an embodiment of the present invention;
[0026] Figure 2This is a network architecture diagram of the road intersection detection model provided in an embodiment of the present invention;
[0027] Figure 3 This is a network structure diagram of the Swing Transformer provided in an embodiment of the present invention;
[0028] Figure 4 This is a flowchart illustrating the structure-aware matching mechanism provided in this embodiment of the invention. Detailed Implementation
[0029] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0030] To address the limitations of existing technologies in terms of intersection recognition accuracy, structural modeling capabilities, and adaptability to complex scenes, this invention proposes a road intersection detection method for optical remote sensing images. This method aims to improve the detection accuracy and structural recognition capabilities of various intersection targets, and enhance the model's generalization ability and robustness in complex scenes. Figure 1 As shown in the flowchart, the method specifically includes the following steps:
[0031] S1. By introducing the Swing Transformer into the DETR model framework to replace the CNN backbone network, a network architecture for the road intersection detection model is constructed.
[0032] S2. By introducing a structural semantic vector-aware loss term into the standard Hungarian matching algorithm, a joint matching loss function containing structural consistency constraints is constructed.
[0033] S3. The road intersection detection model is trained using the training set. During training, the optimal match is calculated based on the Hungarian matching algorithm, and the total loss is calculated using the joint matching loss function for backpropagation of the matching results.
[0034] S4. The trained road intersection detection model is used to perform forward inference on the optical remote sensing image to obtain the road intersection detection results.
[0035] This method enhances multi-scale feature representation by introducing a Swing Transformer to replace the CNN backbone network in the DETR model framework, constructs a structure-aware matching mechanism to optimize the matching process, and improves target representation capabilities through structural semantic vectors. Overall, it improves the model's performance in detecting road intersections in remote sensing images, achieving both high accuracy and robustness. This invention overcomes the shortcomings of traditional CNN feature extraction and standard Hungarian matching mechanisms in structural modeling and target matching. It integrates the multi-scale semantic representation capabilities of the Swing Transformer with a structure-aware matching strategy, introducing structural semantic vectors as an auxiliary matching basis to enhance the model's ability to discriminate complex road intersection topologies. The following provides a more detailed explanation of each step.
[0036] (1) Step S1: Construct the network architecture of the road intersection detection model
[0037] The network architecture of the road intersection detection model constructed in this invention is as follows: Figure 2 As shown, the system includes a backbone network, an encoder, a decoder, and a prediction head. The backbone network extracts multi-scale spatial semantic features of the image and inputs them into the encoder. The encoder encodes these multi-scale spatial semantic features to model the global semantic information of the image, obtaining global semantic features which are then input into the decoder. The decoder introduces a learnable target query vector and, through interaction with the global semantic features, generates a high-dimensional representation of the potential target. The prediction head outputs the category, bounding box parameters, and structural semantic vector of each target at each intersection based on the high-dimensional representation.
[0038] The backbone network is a Swing Transformer, followed by 1*1 convolutional layers (1*1 conv). The Swing Transformer is used to achieve deep fusion of local details and global context information, and the 1*1 convolutional layers are used to perform channel transformation on the features output by the backbone network, which are then flattened and input into the encoder.
[0039] The network structure of the Swin Transformer is as follows: Figure 3 As shown, the network employs a four-stage hierarchical architecture, with output channels numbering 96, 192, 384, and 768 respectively. In the first stage, the input image (image) with length H, width W, and 3 channels is first divided into 4×4 non-overlapping blocks through a patch partitioning layer. Then, it is mapped to a C=96-dimensional feature vector through a linear embedding layer. Feature extraction is then performed using two Swin Transformer blocks (SwinTransformer Block × 2), resulting in an output size of H. The feature map has C channels. In the remaining nth (n=2,3,4) stage, the input feature map first passes through a patch merging layer to merge adjacent 2×2 blocks, achieving spatial downsampling and channel expansion; then it is mapped to a 2n-1*C dimensional feature vector through a linear embedding layer, and features are extracted through two or more Swin Transformer blocks, with the final output size being... The number of channels is 2 n-1 *High-dimensional feature maps of C. Each stage's SwinTransformer block alternately employs window multi-head self-attention (W-MSA) and shifted window attention (SW-MSA), combined with a feedforward network (FFN), layer normalization (LN), and residual connections to achieve local and global feature fusion. The corresponding number of Swin Transformer blocks in each stage are 2, 2, 6, and 2, respectively.
[0040] The feature maps extracted by the Swin Transformer are reduced to 256 dimensions by a 1*1 convolutional layer. After flattening, a sequence of shape [B,N,256] is obtained, where B represents the batch size and N is the number of positions after flattening.
[0041] The encoder employs a 3-layer Transformer encoder. Feature maps extracted by 1x1 convolutional layers are input to the 3-layer Transformer encoder to model the global context information of the image. Each encoder layer includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and a layer normalization module, and introduces a relative position bias (RPB) to enhance the spatial structure modeling capability.
[0042] The decoder employs a 6-layer Transformer decoder. The feature sequence output from the encoder is fed into the 6-layer Transformer decoder, with each layer containing target query self-attention, cross-attention, a feedforward network, residual connections, and layer normalization modules. Initialized target query vectors (ObjectQueries) with dimensions [100, 256] interact with the encoder features layer by layer through the decoder, generating a high-dimensional representation of the potential target. Each target query vector ultimately outputs three types of information through the prediction head: target category, bounding box parameters, and structural semantic vector. The structural semantic vector, with 256 dimensions, represents the spatial topological features of the intersection targets, including road direction, connection patterns, and intersection modes. It serves as a key indicator for measuring structural similarity and is used to measure structural consistency during the structure-aware matching process in the training phase.
[0043] The number of layers in the Transformer encoder is reduced from 6 layers in the standard DETR (Detection Transformer) to 3 layers to reduce the computational cost and parameter overhead of the model. The encoder receives feature sequences from the Swing Transformer backbone network and introduces a relative position bias (RPB) mechanism to enhance the model's ability to model spatial position information.
[0044] The decoder ultimately outputs 100 candidate targets. Each candidate result contains two types of information: 1) intersection type, including three-way intersections, four-way intersections, roundabouts, grade-separated intersections, etc.; 2) bounding box parameters, including center point coordinates, width, and height.
[0045] During the inference phase, each target query generates a class probability distribution through the prediction header, and calculates the maximum class probability as the confidence level of the target through Softmax normalization. All candidate targets are sorted from high to low confidence, targets with confidence levels below 0.6 are discarded, and high-confidence predictions are retained as the final detection results to improve the accuracy and stability of inference.
[0046] The network architecture of the road intersection detection model designed in this invention integrates the multi-scale semantic expression capability of Swing Transformer with the structure-aware matching strategy, and introduces structural semantic vectors as auxiliary matching basis to enhance the model's ability to discriminate complex road intersection topologies.
[0047] (2) Step S2: Construct the loss function
[0048] Figure 4 This is a flowchart illustrating the structure-aware matching mechanism. Figure 4 As shown, this invention, based on the standard DETR category and location matching (standard Hungarian algorithm), introduces structural semantic vectors to enhance the topological consistency of intersection targets, which is called the structure-aware matching mechanism. During the training phase, each prediction result contains three types of outputs: target category probability, bounding box regression parameters (center point coordinates, width, and height), and predicted target structural semantic vector. The predicted target structural semantic vector is obtained by the decoder target query through a two-layer MLP mapping, used to characterize the spatial topological features of intersection targets, including road direction, connection form, and intersection method. The true target structural semantic vector extracts the channel features of the corresponding labeled region from the feature map of the fourth stage of the backbone network, and generates a structural representation with the same dimension as the predicted vector through average pooling.
[0049] The joint matching loss function containing a structure-aware term is constructed as follows:
[0050]
[0051] Among them, L cls(i,j) The cross-entropy loss represents the difference between the predicted target class and the true target class, and is used to measure the accuracy of the class prediction; L L1(i,j) L1 regression loss represents the bounding box loss, measuring the difference between the predicted and ground truth bounding boxes in terms of location parameters (center point, width, and height); L GIoU(i,j) L represents the generalized IoU loss, used to evaluate the geometric overlap and spatial relationship between the predicted bounding box and the ground truth bounding box; str(i,j) λ represents the structure-aware loss, which measures the similarity between the predicted structural semantic vector and the true structural semantic vector. cls , λ box , λ giou , λ str These are the weighting coefficients for each type of loss.
[0052] Among them, the structure-aware loss term L str(i,j)) The calculation is based on the Euclidean distance between the predicted structural semantic vector and the true structural semantic vector, specifically defined as follows:
[0053]
[0054] Among them, f i To predict the structural semantic vector of the target, g j Let f be the structural semantic vector of the real target, d be the dimension of the structural semantic vector, and f be the structural semantic vector. i (k) g j (k) These are the k-th components of the structural semantic vector.
[0055] This invention improves the accuracy and structural consistency of target representation by involving structural semantic vectors in the joint matching loss optimization process, making the model more stable in complex intersection detection scenarios.
[0056] (3) Step S3: Model Training
[0057] By designing a joint matching loss function, a one-to-one correspondence between the predicted target and the real target is achieved during the training phase.
[0058] First, it is necessary to acquire the remote sensing image dataset and perform preprocessing, specifically including:
[0059] Obtain a remote sensing image dataset; the dataset consists of multiple high-resolution optical remote sensing images, and the location and type of the intersection targets are marked in the images.
[0060] Preprocessing of raw remote sensing images mainly includes operations such as image size standardization, annotation format unification, and pixel normalization.
[0061] In practice, the process begins by acquiring optical remote sensing image data of road intersection targets labeled with publicly available datasets such as DIOR, DOTA, and HRRSD. Then, the images are scaled down to 800 pixels on the shorter side and no more than 1333 pixels on the longer side, maintaining the original aspect ratio. Next, the images are uniformly converted to COCO format, with annotations including category labels and rectangular bounding boxes (center point coordinates, width, and height). Finally, pixel normalization is performed to enhance the stability of the model input.
[0062] During model training, the optimal match is calculated based on the Hungarian algorithm. By minimizing the total cost matrix, the predicted results are optimally paired with the true labels, and a weighted total loss is calculated for backpropagation on the matching results. The weights of each loss term are adjusted according to the actual task requirements to balance classification accuracy and training efficiency. In this embodiment, it is set to λ. cls =1.0, λ box =5.0, λ giou =2.0, λ str =0.8.
[0063] Based on the input feature format and structural expression requirements of intersection targets, and combined with the model's structural characteristics and training stability requirements, an intersection detection model is trained. Training is based on the PyTorch 1.11 framework, using the AdamW optimizer. The initial learning rate is set to 0.0001, decaying to 0.1 times its original value every 50 epochs, with a total of 200 training epochs and a batch size of 4. The training process is end-to-end, improving the overall model convergence speed and modeling capability.
[0064] (4) Step S4: Model Application
[0065] Once trained, the model can directly perform single-stage forward inference on optical remote sensing images without the need for candidate box generation or multi-stage filtering.
[0066] Specifically, the image to be tested is input into the model and processed sequentially through the Swin Transformer backbone network, Transformer encoder, and decoder, ultimately outputting 100 candidate targets. Each candidate result contains two types of information: 1) intersection category, including three-way intersections, four-way intersections, roundabouts, grade-separated intersections, etc.; 2) bounding box parameters, including center point coordinates, width, and height. To improve the accuracy and stability of the inference stage, this method ranks the candidate targets based on category confidence and uses a threshold filtering strategy to extract high-confidence targets as the final detection results.
[0067] In summary, this invention provides a method for detecting road intersections in optical remote sensing images. To overcome the shortcomings of existing network model feature extraction and standard Hungarian matching mechanisms in structural modeling and target matching, this method constructs a road intersection detection model. This model utilizes a Swing Transformer backbone network for feature extraction to achieve deep fusion of local details and global contextual information. It also combines a lightweight encoder structure with a structure-aware matching mechanism to improve the efficiency of spatial feature modeling while enhancing the structural representation capability of target recognition. Furthermore, by using structural semantic vectors in the joint matching loss optimization process, it further improves the target representation accuracy and structural consistency, making the model more stable in complex intersection detection scenarios.
[0068] Compared to traditional methods that rely on convolutional features and fixed matching strategies (such as Faster R-CNN based on convolutional feature extraction, or YOLO series methods that rely on preset anchor boxes for fixed matching), this method combines the local modeling capabilities of the Swin Transformer with the end-to-end detection mechanism of DETR, effectively addressing challenges such as large scale variations and complex backgrounds in remote sensing images. By introducing structural semantic vectors and constructing a structure-aware matching mechanism, it achieves joint optimization of classification, regression, IoU, and topological structure, omitting candidate boxes and NMS operations. It possesses advantages such as high efficiency, accuracy, and robustness, and is suitable for fine-grained identification and structural analysis of various types of intersections (three-way intersections, four-way intersections, roundabouts, and overpasses, etc.). It has good structural generalization ability and target category discrimination ability, and is suitable for high-precision detection in other remote sensing scenarios with complex structures and similar categories.
[0069] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for detecting road intersections in optical remote sensing images, characterized in that, include: S1. By introducing the Swing Transformer into the DETR model framework to replace the CNN backbone network, a network architecture for the road intersection detection model is constructed. S2. By introducing a structural semantic vector-aware loss term into the standard Hungarian matching algorithm, a joint matching loss function containing structural consistency constraints is constructed. S3. The road intersection detection model is trained using the training set. During training, the optimal match is calculated based on the Hungarian matching algorithm, and the total loss is calculated based on the joint matching loss function. The model parameters are optimized through backpropagation. S4. The trained road intersection detection model is used to perform forward inference on the optical remote sensing image to obtain the road intersection detection results.
2. The method for detecting road intersections in optical remote sensing images according to claim 1, characterized in that: The network architecture of the road intersection detection model includes a backbone network, an encoder, a decoder, and a prediction head; the backbone network is used to extract multi-scale spatial semantic features of the image and input them into the encoder. The encoder is used to encode multi-scale spatial semantic features to obtain global semantic features, which are then input into the decoder. The decoder is used to introduce a learnable target query vector and generate a high-dimensional representation of the potential target by interacting with the global semantic features. The prediction head is used to output the category, bounding box parameters, and structural semantic vector of each intersection target based on the high-dimensional representation.
3. The method for detecting road intersections in optical remote sensing images according to claim 2, characterized in that: The backbone network is a Swing Transformer, followed by 1*1 convolutional layers; The Swin Transformer employs a four-stage hierarchical architecture. In the first stage, the input image, with length H, width W, and 3 channels, is first divided into 4×4 non-overlapping blocks by a block-splitting layer. Then, it is mapped to a C-dimensional feature vector through a linear embedding layer. Finally, feature extraction is performed using two Swin Transformer blocks, with an output size of [missing information]. The input feature map has C channels. In the remaining nth stage, the input feature map first passes through a block merging layer to merge adjacent 2×2 blocks, achieving spatial downsampling and channel expansion; then it is mapped to a 2×2 feature map through a linear embedding layer. n-1 *C-dimensional feature vectors are extracted using two or more Swing Transformer blocks, resulting in a final output of size C. The number of channels is 2 n-1 *High-dimensional feature maps of C, n = 2, 3, 4; The feature maps extracted by the Swin Transformer are reduced to 256 dimensions by the 1*1 convolutional layer. After flattening, a sequence of shape [B,N,256] is obtained, where B represents the batch size and N is the number of positions after flattening.
4. The method for detecting road intersections in optical remote sensing images according to claim 3, characterized in that: Each stage of the Swing Transformer block alternately employs multi-head window self-attention and shifted window attention, combined with feedforward networks, layer normalization, and residual connections to achieve the fusion of local and global features.
5. A method for detecting road intersections in optical remote sensing images according to claim 4, characterized in that: The number of Swing Transformer blocks included in the first to fourth stages are 2, 2, 6 and 2, respectively.
6. A method for detecting road intersections in optical remote sensing images according to claim 5, characterized in that: The encoder employs a 3-layer Transformer encoder. Each layer of the Transformer encoder includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and a layer normalization module, and introduces a relative position bias.
7. A method for detecting road intersections in optical remote sensing images according to claim 6, characterized in that: The decoder employs a 6-layer Transformer decoder. The feature sequence output by the encoder is fed into the 6-layer Transformer decoder. Each layer of the decoder includes a target query self-attention, cross-attention, feedforward network, residual connection, and layer normalization module. The initialized target query vector set interacts with the encoder output to generate a high-dimensional representation of the potential target.
8. A method for detecting road intersections in optical remote sensing images according to claim 7, characterized in that: The joint matching loss function is constructed as L cls(i,j) L L1(i,j) L GIoU(i,j) L str(i,j) The weighted sum, L cls(i,j) L represents the cross-entropy loss between the predicted target class and the true target class. L1(i,j) L represents the L1 regression loss of the bounding box. GIoU(i,j) L represents the generalized IoU loss, used to evaluate the geometric overlap and spatial relationship between the predicted bounding box and the ground truth bounding box; str(i,j) The structure-aware loss measures the similarity between the predicted target's structural semantic vector and the true target's structural semantic vector.
9. A method for detecting road intersections in optical remote sensing images according to claim 7, characterized in that: The prediction head consists of a feedforward neural network, which is a three-layer perceptron with a ReLU activation function, and outputs the prediction result through a linear projection layer. The prediction head includes a category prediction branch, a bounding box regression branch, and a target structure semantic vector prediction branch: the category prediction branch consists of a linear layer and a softmax function; the bounding box branch consists of a three-layer perceptron and is used to predict the normalized center coordinates, height, and width; the target structure semantic vector prediction branch consists of a two-layer perceptron and is used to map the target query output by the decoder into the spatial topological feature vector of the intersection; the true target structure semantic vector is obtained by average pooling the labeled box regions in the fourth stage feature map of the backbone network.
10. A method for detecting road intersections in optical remote sensing images according to claim 9, characterized in that: When training the road intersection detection model, a remote sensing image dataset is first acquired and preprocessed, specifically including: Acquire a remote sensing image dataset; the dataset consists of multiple high-resolution optical remote sensing images, and the location and type of the intersection targets are marked in the images; Preprocessing of the raw remote sensing images includes image size standardization, annotation format unification, and pixel normalization.
Citation Information
Cited By
VHR remote sensing image road intersection detection method based on YOLO11
CN121725219A
Multi-scale remote sensing target detection method and device based on Swin Transform and storage medium
CN121725222A