An end-to-end spatial-aware cross-view object segmentation method
Patent Information
- Application Number
- CN202610971881.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]然而,这种两阶段方法在处理道路等具有不规则、细长或蜿蜒几何结构的目标时,存在显著缺陷:
[0067](1)端到端优化:本发明摒弃了两阶段流程,避免了边界框预测误差对分割结果的干扰,实现了从输入图像到分割掩膜的直接映射。
Smart Images

Figure CN122820751A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and remote sensing image processing, and specifically relates to an end-to-end spatial perception cross-view target segmentation method. Background Technology
[0002] Cross-View Geo-Localization (CVGL) technology locates query images taken by ground-based or low-altitude UAVs by matching geotagged satellite reference images. It is a key auxiliary technology for Global Navigation Satellite Systems (GNSS) in signal-constrained scenarios.
[0003] To further improve the accuracy of localization, the Cross-View Object Segmentation (CVOS) task has emerged. Its purpose is to use visual information in the query image as a prompt to achieve pixel-level object segmentation in the corresponding satellite reference image.
[0004] Existing CVOS methods (such as the TROGeo algorithm) typically employ a two-stage cascading strategy:
[0005] The first stage is to train an object detection network to predict the bounding boxes of objects in satellite images.
[0006] The second stage involves using the predicted bounding boxes and their center points as cues to input into a pre-trained large segmentation model (such as the Segment Anything Model, SAM) to generate the final segmentation mask.
[0007] However, this two-stage approach has significant drawbacks when dealing with targets such as roads that have irregular, slender, or meandering geometries:
[0008] (1) Difficulty in bounding box regression: Roads usually have a long and thin curved shape, and traditional rectangular bounding boxes are difficult to tightly surround the target, resulting in the bounding boxes predicted in the first stage often containing a large amount of background area, which cannot accurately cover the target body.
[0009] (2) Error cascade propagation: The bounding box prediction error in the first stage will be directly propagated to the second stage, causing the segmentation results generated by SAM to be offset or missing.
[0010] (3) Redundancy of computational resources: Excessively large bounding box range will cause SAM to process a large number of irrelevant background pixels during inference, which significantly increases the computational burden and inference delay.
[0011] Therefore, there is an urgent need for a segmentation method that can extract multi-scale features end-to-end and explicitly establish cross-view spatial correspondences to overcome the limitations of existing technologies in the segmentation of slender targets. Summary of the Invention
[0012] To address the aforementioned issues, this invention provides an end-to-end spatially perceptive cross-view target segmentation method. By designing a cross-domain multi-scale fusion module (CDMS) and embedding adaptive gated feature modulation (AGFM) and coordinate-aware multi-scale feature interaction (CMFI) units, it achieves effective integration of multi-scale features and accurate capture of spatial information, significantly improving the segmentation accuracy of slender targets such as roads.
[0013] The specific steps of the end-to-end spatially aware cross-view target segmentation method are as follows:
[0014] Step 1: Collect query images containing the slender target to be tested. and reference image A cross-view dataset, and preprocess each image separately;
[0015] The query point is marked in the surrounding environment of the slender target in the query image, and the pixel-level segmentation ground truth value of the slender target is marked in the reference image.
[0016] Preprocessing includes resizing and data augmentation for each image individually.
[0017] Step 2: Construct a dual-branch encoder network to extract features from the query image and the reference image respectively, and perform mapping to obtain query mapping features and reference mapping features with consistent dimensions.
[0018] The dual-branch encoder network includes the ResNet network and the LSKNet network;
[0019] The first branch uses ResNet as its base network. After inputting a query image with incorporated cue points, it extracts four layers of feature maps at different scales. The last layer, a 512-channel feature map, is selected as the core query feature, denoted as […]. .
[0020] The second branch uses LSKNet as its base network, extracts four layers of feature maps at different scales from the input satellite image, and selects the third-layer channel feature map as the reference core feature, denoted as... .
[0021] Then, through two independent convolutional layers, the query core and reference core features are mapped to 512 channels respectively, resulting in query mapping features and reference mapping features with consistent dimensions.
[0022] Step 3: Use the DetGeo method to remap the query mapping features back to the geometric space of the reference mapping features, perform cross-view alignment, and obtain the aligned fused features. .
[0023] Specifically:
[0024] First, average pooling is performed on the query mapping features to generate a one-dimensional global feature vector and normalization is completed. At the same time, channel dimension normalization is performed on the reference mapping features.
[0025] Then, the reference mapping features normalized by the channel dimension are flattened and multiplied with the normalized global query features to obtain the similarity matrix;
[0026] Finally, the similarity matrix is normalized to generate an attention weight map, which is then reshaped and multiplied element-wise with the reference mapping features to obtain the cross-view fusion features. .
[0027] Step 4: Merge features The input is fed into the cross-domain multi-scale fusion CDMS module, where a progressive resolution enhancement strategy is employed to restore its resolution to match that of the reference image, thus obtaining the calibration features. ;
[0028] The CDMS module contains N cascaded processing flows, where N is a positive integer greater than or equal to 2; the specific processing flow is as follows:
[0029] Step 401: Analyze the fusion features Upsampling is performed to double the resolution;
[0030] Step 402: Concatenate the features with doubled resolution and the features of the reference image through channels to obtain the concatenated features. Using the Adaptive Gated Feature Modulation Unit (AGFM) to splice features Enhancement is performed to obtain spatial enhancement features. ;
[0031] Specifically:
[0032] Step 4021: Assess the splicing features Perform adaptive max pooling respectively Adaptive average pooling and standard deviation pooling .
[0033] Step 4022: Introduce a global feature-guided gating selection mechanism to generate three normalized weights. and satisfy .
[0034] Step 4023: Use three weights to perform a weighted fusion of the results of the three pooling operations to obtain the weighted fusion feature. :
[0035]
[0036] in, For dimensionality reduction convolution, It is the ReLU activation function. This is an up-dimensional convolution.
[0037] Step 4024: Activate using the Sigmoid function Weighted fusion features Generate channel attention weights and with the original features Add them together to obtain the enhanced features of the channels. :
[0038]
[0039]
[0040] Step 4025: Enhance the features of the channel Max pooling is performed again on the channel dimension. Average pooling and standard deviation pooling .
[0041] Step 4026: Generate spatial weights using a gating selection mechanism. The three pooling results are then weighted and concatenated to obtain the features. :
[0042]
[0043] Step 4027, through Convolution and Sigmoid activation for features Generate spatial attention weights and features enhanced by the channel Performing the Hadamard product yields spatially enhanced features. :
[0044] ,
[0045] Step 403: Enhance spatial features using the Coordinate-Aware Multi-Scale Feature Interaction Unit (CMFI). Perform spatial calibration to obtain calibration characteristics. This completes one cascading process;
[0046] Specifically:
[0047] Step 4031: Enhance spatial features pass Convolution expands the feature channels to obtain expanded features. Extended features are processed using three-way parallel depthwise convolution (DWC). The three outputs are added together and fused to obtain the feature. :
[0048] The kernel sizes are respectively , and .
[0049]
[0050] Step 4032: Perform feature analysis along both the horizontal and vertical directions. Perform 1D average pooling to generate feature vectors in two directions. and :
[0051]
[0052]
[0053] Step 4033, will and After splicing, pass The convolution is used for encoding, and then the result is split into two tensors. Attention weights are obtained by performing convolution and sigmoid activation on the tensors respectively. and .
[0054] Step 4034: Inject coordinate information into features through element-wise multiplication. , to obtain features :
[0055]
[0056] Step 4035, for features Perform channel shuffling operation and output calibration characteristics. .
[0057] Step 404: Return to step 401 and perform calibration on the features. Upsampling is performed to increase the resolution, followed by channel stitching, feature enhancement, and spatial calibration to obtain calibrated features. This completes the secondary cascading process;
[0058] Step 405: Repeat the above cascade process N times to obtain the calibration characteristics. ;
[0059] Step 5: Apply calibration features Upsampling is performed to obtain high-resolution features. The data is then input into a deep neural network to generate a prediction mask. The parameters of the deep neural network are optimized by calculating the loss function. The output corresponding to the optimal parameters of the deep neural network is the boundary result between the slender target and its surrounding environment.
[0060] Specifically:
[0061] First, high-resolution features Input a convolutional layer (segmentation head), which is mapped to a single-channel prediction map. Each pixel value in the map represents the probability that the location belongs to the slender target to be tested, thus generating a prediction mask.
[0062] Then, the binary cross-entropy loss (BCE Loss) between the predicted mask and the real label is calculated, and the parameters of the deep neural network are updated by using the backpropagation algorithm by minimizing the image segmentation Dice distance between the predicted result and the real label, thus achieving end-to-end training.
[0063] The total loss function is calculated as follows:
[0064]
[0065] It is a binary cross-entropy loss function; The Dice loss function is used for image segmentation.
[0066] The beneficial effects of this invention are as follows:
[0067] (1) End-to-end optimization: This invention abandons the two-stage process, avoids the interference of bounding box prediction error on the segmentation result, and realizes direct mapping from input image to segmentation mask.
[0068] (2) Multi-scale feature fusion: Through the CDMS module and the multi-scale convolution in CMFI, it is possible to capture the local details and global shape of slender targets such as roads at the same time, and adapt to the drastic scale changes of the targets.
[0069] (3) Fine feature enhancement: The AGFM module introduces standard deviation pooling and gating mechanisms, which can more finely and adaptively enhance effective features and suppress background noise compared to traditional CBAM and other modules.
[0070] (4) Precise spatial positioning: The CMFI module explicitly models the horizontal and vertical spatial position of the target through the coordinate attention mechanism, which significantly improves the accuracy of the segmentation boundary.
[0071] (5) Significantly improved performance: Experiments on the CVOGL-seg+ dataset show that our method significantly outperforms existing methods such as TROGeo in terms of mIoU, especially in the scenario of slender road segmentation. Attached Figure Description
[0072] Figure 1 This is a flowchart of an end-to-end spatial perception cross-view target segmentation method according to the present invention;
[0073] Figure 2 This is a schematic diagram of the principle architecture of an end-to-end spatial perception cross-view target segmentation method of the present invention;
[0074] Figure 3 This is a schematic diagram of the adaptive gated feature modulation module (AGFM) in this invention;
[0075] Figure 4 This is a schematic diagram of the coordinate-aware multi-scale feature interaction module (CMFI) in this invention.
[0076] Figure 5 This is a flowchart comparing the method of the present invention with existing two-stage methods. Detailed Implementation
[0077] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0078] This invention provides an end-to-end spatial perception cross-view target segmentation method, abbreviated as SCSeg; it is applicable to scenarios with slender targets such as roads; it eliminates the cumbersome two-stage process, such as... Figure 1 and Figure 2 As shown, the specific steps are as follows:
[0079] Step 1: Collect query images containing the slender target to be tested. and reference image A cross-view dataset, and preprocess each image separately;
[0080] Cross-view datasets such as CVOGL-seg+ include query points labeled in the surrounding environment of the slender target in the query image from the perspective of a drone, and pixel-level ground truth values of the slender target labeled in the reference image from the perspective of a satellite.
[0081] Preprocessing includes resizing and data augmentation for each image individually.
[0082] In practice, querying images The size was adjusted to Pixels, reference image The size was adjusted to Pixels. To enhance the robustness of the model and prevent overfitting, data augmentation operations are performed on the input image, including random horizontal flipping and random rotation.
[0083] Step 2: Construct a dual-branch encoder network to extract features from the query image and the reference image respectively, and perform mapping to obtain query mapping features and reference mapping features with consistent dimensions.
[0084] The dual-branch encoder network comprises ResNet and LSKNet networks. ResNet-18 is used as the backbone to extract features from the query image, while LSKNet tT is used as the backbone to extract features from the reference image. LSKNet (Large Selective Kernel Network) has a large receptive field, making it suitable for capturing global contextual information in remote sensing images. Its large selective kernel mechanism enables it to better adapt to large-scale context in remote sensing images.
[0085] Specifically, the first branch uses ResNet as its base network. After inputting the query image with integrated cue points, it extracts four layers of feature maps at different scales. The last layer, a 512-channel feature map, is selected as the core query feature, denoted as […]. .
[0086] The second branch uses LSKNet as its base network, extracts four layers of feature maps at different scales from the input satellite image, and selects the third-layer channel feature map as the reference core feature, denoted as... .
[0087] Then, through two independent convolutional layers, the query core and reference core features are mapped to 512 channels respectively, resulting in query mapping features and reference mapping features with consistent dimensions.
[0088] Step 3: Use the DetGeo method to remap the query mapping features back to the geometric space of the reference mapping features, perform cross-view alignment, and obtain the aligned fused features. .
[0089] By employing a feature fusion module that borrows from the DetGeo method, and through polar coordinate transformation or projection transformation, features from the UAV's perspective are mapped to the feature space of the satellite's perspective, establishing a preliminary correspondence. Specifically:
[0090] First, average pooling is performed on the query mapping features to generate a one-dimensional global feature vector and normalization is completed. At the same time, channel dimension normalization is performed on the reference mapping features.
[0091] Then, the reference mapping features normalized by the channel dimension are flattened and multiplied with the normalized global query features to obtain the similarity matrix;
[0092] Finally, the similarity matrix is normalized to generate an attention weight map, which is then reshaped and multiplied element-wise with the reference mapping features to obtain the cross-view fusion features. .
[0093] Step 4: Merge features The input is fed into the cross-domain multi-scale fusion CDMS module, where a progressive resolution enhancement strategy is employed to restore its resolution to match that of the reference image and enrich semantic details, thus obtaining the calibration features. ;
[0094] The CDMS module contains N cascaded processing flows, where N is a positive integer greater than or equal to 2; the processing flow of each stage is as follows:
[0095] 1. Upsampling: Upsample the feature map of the previous level to double its resolution.
[0096] 2. Feature Concat: This involves concatenating the upsampled features with the features at the corresponding scale of the satellite branches. Channel splicing is performed to supplement lost details.
[0097] 3. AGFM Processing: Adaptive Gated Feature Modulation (AGFM) units are used to enhance the spliced features. Within the AGFM unit, dynamic channel saliency extraction and context-aware spatial calibration are used to adaptively enhance the feature responses in both the channel and spatial dimensions.
[0098] 4. CMFI Processing: Spatial calibration of features is performed using a coordinate-aware multi-scale feature interaction unit. Within the CMFI unit, contextual information is captured through parallel multi-scale deep convolutions, spatial location information is injected using a coordinate attention mechanism, and finally, channel shuffling is performed.
[0099] A progressive resolution enhancement strategy is adopted: it gradually integrates detailed information from the UAV perspective and global context from the satellite perspective, restoring the feature resolution from low to high through iterative upsampling operations until it matches the size of the original image. In the processing at each scale, adaptive gated feature modulation (AGFM) and CMFI operations are performed sequentially.
[0100] The AGFM module aims to address the information loss problem caused by fixed pooling in traditional attention mechanisms; it adaptively adjusts the feature response by introducing standard deviation pooling and gating mechanisms. Figure 3 As shown, it includes two sub-steps:
[0101] Dynamic channel saliency extraction: Adaptive max pooling (capturing texture details), adaptive average pooling (capturing background information), and standard deviation pooling (capturing the dispersion of the feature distribution) are applied to the input features. Standard deviation pooling supplements the dispersion information of the features. Three weights are generated through a gating mechanism, and the three pooling features are weighted and fused. Channel attention weights are then generated through Sigmoid activation.
[0102] Context-aware spatial calibration: After channel enhancement, the features are subjected to three pooling operations again in the channel dimension. Spatial attention weights are generated through convolution and sigmoid to further enhance the spatial response of the features.
[0103] The Coordinate-Aware Multi-Scale Feature Interaction (CMFI) module aims to perceive the scale changes and precise location of targets, addressing the limited receptive field of a single convolutional kernel, and explicitly injecting the target's absolute spatial coordinate information, which is crucial for locating narrow roads. It comprises three sub-steps:
[0104] Multi-scale sensing: utilizing parallel , , Deep convolutional flow processes features to capture contextual information under different receptive fields.
[0105] Coordinate attention calibration: 1D pooling is performed on the features along the horizontal (H) and vertical (W) directions respectively to generate attention maps that can encode precise spatial coordinates, and these maps are then injected into the features.
[0106] Channel shuffling: Features are shuffled through channels to facilitate information flow between different groups of features and enhance the model's generalization ability.
[0107] The specific processing procedure is as follows:
[0108] Step 401: Analyze the fusion features Upsampling is performed to double the resolution;
[0109] Step 402: Concatenate the features with doubled resolution and the features of the reference image through channels to obtain the concatenated features. Using the Adaptive Gated Feature Modulation Unit (AGFM) to splice features Enhancement is performed to obtain spatial enhancement features. ;
[0110] like Figure 3 As shown, specifically:
[0111] Step 4021: Assess the splicing features Perform adaptive max pooling respectively Adaptive average pooling and standard deviation pooling .
[0112] Step 4022: Introduce a global feature-guided gating selection mechanism to generate three normalized weights. and satisfy .
[0113] Step 4023: Use three weights to perform a weighted fusion of the results of the three pooling operations to obtain the weighted fusion feature. :
[0114]
[0115] in, For dimensionality reduction convolution, It is the ReLU activation function. This is an up-dimensional convolution.
[0116] Step 4024: Activate using the Sigmoid function Weighted fusion features Generate channel attention weights and with the original features Add them together to obtain the enhanced features of the channels. :
[0117]
[0118] Step 4025: Enhance the features of the channel Max pooling is performed again on the channel dimension. Average pooling and standard deviation pooling .
[0119] Step 4026: Generate spatial weights using a gating selection mechanism. The three pooling results are then weighted and concatenated to obtain the features. :
[0120]
[0121] Step 4027, through Convolution and Sigmoid activation for features Generate spatial attention weights and features enhanced by the channel Performing the Hadamard product (element-wise multiplication) yields the spatially enhanced features. :
[0122] ,
[0123] Step 403: Enhance spatial features using the Coordinate-Aware Multi-Scale Feature Interaction Unit (CMFI). Perform spatial calibration to obtain calibration characteristics. This completes one cascading process;
[0124] like Figure 4 As shown, specifically:
[0125] Step 4031: Enhance spatial features pass Convolution expands the feature channels to obtain expanded features. Extended features are processed using three parallel depth-wise convolutions (DWC). The three outputs are added together and fused to obtain the feature. To capture multi-scale contextual information;
[0126] The kernel sizes are respectively , and .
[0127]
[0128] Step 4032: To preserve accurate positional information, the features are analyzed along both the horizontal (H-axis) and vertical (W-axis) directions. Perform 1D average pooling to generate feature vectors in two directions. and :
[0129]
[0130]
[0131] Step 4033, will and After splicing, pass The convolution is used for encoding, and then the result is split into two tensors. Attention weights are obtained by performing convolution and sigmoid activation on the tensors respectively. and .
[0132] Step 4034: Inject coordinate information into features through element-wise multiplication. , to obtain features :
[0133]
[0134] Step 4035, for features Perform a channel shuffle operation and output calibration characteristics. .
[0135] Rearranging feature channels allows information to flow between different groups, enhancing the model's generalization ability and improving its generalization capabilities. Perform channel shuffling operation and output the final features.
[0136] Step 404: Return to step 401 and perform calibration on the features. Upsampling is performed to increase the resolution, followed by channel stitching, feature enhancement, and spatial calibration to obtain calibrated features. This completes the secondary cascading process;
[0137] Step 405: Repeat the above cascade process N times to obtain the calibration characteristics. ;
[0138] Step 5: Apply calibration features Upsampling is performed to obtain high-resolution features. The data is then input into a deep neural network to generate a prediction mask. The parameters of the deep neural network are optimized by calculating the loss function. The output corresponding to the optimal parameters of the deep neural network is the boundary result between the slender target and its surrounding environment.
[0139] Specifically:
[0140] First, high-resolution features Input a convolutional layer (segmentation head), which is mapped to a single-channel prediction map. Each pixel value in the map represents the probability that the location belongs to the slender target to be tested, thus generating a prediction mask.
[0141] Then, the binary cross-entropy loss (BCE Loss) between the predicted mask and the real label is calculated, and the parameters of the deep neural network are updated by using the backpropagation algorithm by minimizing the image segmentation Dice distance between the predicted result and the real label, thus achieving end-to-end training.
[0142] The total loss function, designed to balance the imbalance between positive and negative samples and optimize the segmentation boundary, is calculated as follows:
[0143]
[0144] It is a binary cross-entropy loss function, responsible for restoring the classification accuracy of pixels; The Dice loss function is used for image segmentation, responsible for the overall shape integrity of the narrow road.
[0145] Example:
[0146] This embodiment is implemented based on the PyTorch deep learning framework, with a single NVIDIA GeForce RTX4090 GPU as the hardware environment. The optimizer used is the stochastic gradient descent (SGD) optimizer, and the hyperparameters are set to [specified parameters], with the initial learning rate set to [specified value]. Weight decay is set to The momentum was set to 0.9. The batch size was set to 4. The model was trained for 100 epochs. A cosine annealing strategy was used to dynamically adjust the learning rate.
[0147] Figure 5 This diagram compares the flow of the present invention with existing two-stage methods. To verify the effectiveness of the present invention, experiments were conducted on the constructed CVOGL-seg+ dataset. This dataset is an extension of CVOGL-seg, which for the first time includes the Road object category. The dataset contains: Training set: 965 image pairs; Validation set: 273 image pairs; Test set: 304 image pairs.
[0148] Comparative experimental results: mIoU (mean intersection-union ratio), acc@50, and acc@25 were used as evaluation metrics. The method of this invention (SCSeg) was compared with mainstream two-stage methods (such as TROGeo and DetGeo), and the results are shown in the table below:
[0149]
[0150] Experimental data show that the SCSeg model proposed in this invention significantly outperforms existing technologies in all metrics. Particularly in terms of mIoU, SCSeg achieves an 18.23% improvement compared to the state-of-the-art TROGeo method; this significant improvement is mainly attributed to:
[0151] 1. End-to-end design: avoids the cascading errors caused by inaccurate bounding box prediction.
[0152] 2. CMFI module: Through coordinate attention and multi-scale convolution, it accurately captures the slender shape and spatial position of the road.
[0153] 3. AGFM module: Through an adaptive gating mechanism, it effectively enhances target features and suppresses background noise.
Claims
1. An end-to-end spatially perceptual cross-view target segmentation method, characterized in that, The specific steps are as follows: Step 1: Collect query images containing the slender target to be tested. and reference image A cross-view dataset, and preprocess each image separately; Step 2: Construct a dual-branch encoder network to extract features from the preprocessed query image and reference image respectively, and perform mapping to obtain query mapping features and reference mapping features with consistent dimensions; Step 3: Using the DetGeo method, the query mapping features are remapped to the geometric space of the reference mapping features, performing cross-perspective alignment of spatial and semantic information to obtain the aligned fused features. ; Step 4: Merge features The input is fed into the cross-domain multi-scale fusion CDMS module, where a progressive resolution enhancement strategy is employed to restore its resolution to match that of the reference image, thus obtaining the calibration features. ; The CDMS module contains N cascaded processing flows, where N is a positive integer greater than or equal to 2: Step 401: Analyze the fusion features Upsampling is performed to double the resolution; Step 402: Concatenate the features with doubled resolution and the features of the reference image through channels to obtain the concatenated features. Using adaptive gated feature modulation units to splice features Enhancement is performed to obtain spatial enhancement features. ; Step 403: Enhance spatial features using coordinate-aware multi-scale feature interaction units. Perform spatial calibration to obtain calibration characteristics. This completes one cascading process; Step 404: Return to step 401 and perform calibration on the features. Upsampling is performed to increase the resolution, followed by channel stitching, feature enhancement, and spatial calibration to obtain calibrated features. This completes the secondary cascading process; Step 405: Repeat the above cascade process N times to obtain the calibration characteristics. ; Step 5: Apply calibration features Upsampling is performed to obtain high-resolution features. The data is then input into a deep neural network to generate a prediction mask. The parameters of the deep neural network are optimized by calculating the loss function. The output corresponding to the optimal parameters of the deep neural network is the boundary result between the slender target and its surrounding environment.
2. The end-to-end spatially perceptual cross-view target segmentation method as described in claim 1, characterized in that, In step one, query points are marked in the surrounding environment of the slender target to be tested in the query image, and the pixel-level segmentation ground value of the slender target is marked in the reference image. Preprocessing includes resizing and data augmentation for each image individually.
3. The end-to-end spatially perceptual cross-view target segmentation method as described in claim 1, characterized in that, In step two, the dual-branch encoder network includes a ResNet network and an LSKNet network; The first branch uses ResNet as its base network. After inputting a query image with incorporated cue points, it extracts four layers of feature maps at different scales. The last layer, a 512-channel feature map, is selected as the core query feature, denoted as […]. ; The second branch uses LSKNet as its base network, extracts four layers of feature maps at different scales from the input satellite image, and selects the third-layer channel feature map as the reference core feature, denoted as... ; Then, through two independent convolutional layers, the query core and reference core features are mapped to 512 channels respectively, resulting in query mapping features and reference mapping features with consistent dimensions.
4. The end-to-end spatial perception cross-view target segmentation method as described in claim 1, characterized in that, Step three specifically involves: First, average pooling is performed on the query mapping features to generate a one-dimensional global feature vector and normalization is completed. At the same time, channel dimension normalization is performed on the reference mapping features. Then, the reference mapping features normalized by the channel dimension are flattened and multiplied with the normalized global query features to obtain the similarity matrix; Finally, the similarity matrix is normalized to generate an attention weight map, which is then reshaped and multiplied element-wise with the reference mapping features to obtain the cross-view fusion features. .
5. The end-to-end spatially perceptual cross-view target segmentation method as described in claim 1, characterized in that, Step 402 specifically involves: Step 4021: Assess the splicing features Perform adaptive max pooling respectively Adaptive average pooling and standard deviation pooling ; Step 4022: Introduce a global feature-guided gating selection mechanism to generate three normalized weights. and satisfy ; Step 4023: Use three weights to perform a weighted fusion of the results of the three pooling operations to obtain the weighted fusion feature. : in, For dimensionality reduction convolution, It is the ReLU activation function. For higher-dimensional convolution; Step 4024: Activate using the Sigmoid function Weighted fusion features Generate channel attention weights and with the original features Add them together to obtain the enhanced features of the channels. : Step 4025: Enhance the features of the channel Max pooling is performed again on the channel dimension. Average pooling and standard deviation pooling ; Step 4026: Generate spatial weights using a gating selection mechanism. The three pooling results are then weighted and concatenated to obtain the features. : Step 4027, through Convolution and Sigmoid activation for features Generate spatial attention weights and features enhanced by the channel Performing the Hadamard product yields spatially enhanced features. : 。 6. The end-to-end spatially perceptual cross-view target segmentation method as described in claim 1, characterized in that, Step 403 specifically involves: Step 4031: Enhance spatial features pass Convolution expands the feature channels to obtain expanded features. Extended features are processed using three-way parallel depthwise convolution (DWC). The three outputs are added together and fused to obtain the feature. : The kernel sizes of the three-way parallel depthwise convolution (DWC) are respectively , and ; Step 4032: Perform feature analysis along both the horizontal and vertical directions. Perform 1D average pooling to generate feature vectors in two directions. and : Step 4033, will and After splicing, pass The convolution is used for encoding, and then the result is split into two tensors. Attention weights are obtained by performing convolution and sigmoid activation on the tensors respectively. and ; Step 4034: Inject coordinate information into features through element-wise multiplication. , to obtain features : Step 4035, for features Perform channel shuffling operation and output calibration characteristics. .
7. The end-to-end spatially perceptual cross-view target segmentation method as described in claim 1, characterized in that, Step five specifically involves: First, high-resolution features Input a convolutional layer (segmentation head), which is mapped to a single-channel prediction map. Each pixel value in the map represents the probability that the location belongs to the slender target to be tested, thus generating a prediction mask. Then, the binary cross-entropy loss between the predicted mask and the real label is calculated, and the parameters of the deep neural network are updated by using the backpropagation algorithm by minimizing the image segmentation Dice distance between the predicted result and the real label, thus achieving end-to-end training.