A Deep Learning-Based Remote Sensing Image Edge Detection Method
Patent Information
- Application Number
- CN202511493212.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-10-20
AI Technical Summary
S11、将遥感影像输入自适应大模型提取语义提示特征图;
本发明提供的一种基于深度学习的遥感影像边缘检测方法,通过自适应大模型语义提示与边缘检测大模型边缘特征的统一深度融合,实现跨场景泛化能力与几何细节的互补,显著减少光照、阴影、噪声干扰导致的虚检与漏检;通过多尺度融合增强,实现遥感影像边缘连续闭合,增强后续分割、变化检测、三维重建任务的边界可靠性;通过交叉注意力二次语义定位,实现像素级目标边缘轮廓与语义类别严格对齐。本发明通过端到端大模型全链路设计,实现无需传统后处理即可直接输出高精度、语义一致且闭合完整的目标边缘轮廓图,显著降低系统复杂度与部署成本。
Smart Images

Figure CN121458995B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge detection technology, and more particularly to a remote sensing image edge detection method based on deep learning. Background Technology
[0002] Driven by the rapid development of remote sensing technology, high spatial resolution, high spectral resolution, and multi-temporal remote sensing imagery have been widely applied in fields such as land surveying, urban planning, disaster monitoring, and agricultural surveys. The edge information of remote sensing images corresponds to the geometric boundaries, texture abrupt changes, and semantic transformation regions of ground features, serving as a crucial foundation for subsequent advanced applications such as target recognition, change detection, and 3D reconstruction. Therefore, how to accurately and robustly extract the edges of remote sensing images has become one of the core issues of long-term concern in the field of remote sensing image processing.
[0003] In recent years, deep learning technology has made groundbreaking progress in the field of computer vision. The powerful feature learning capabilities of convolutional neural networks (CNNs) have provided new solutions for edge detection in remote sensing images. CNN-based edge detection methods, through end-to-end training, can automatically learn multi-scale, multi-level edge features, significantly improving detection accuracy and robustness. However, directly transferring general-purpose deep edge detection networks to remote sensing scenarios still faces challenges. Remote sensing images suffer from large modal differences, indistinct features, and diverse semantic changes, making them difficult to effectively adapt to open-domain remote sensing change detection in new scenarios. How to maintain high detection accuracy while ensuring computational efficiency, and how to effectively integrate remote sensing-specific edge priors and semantic cues, remain pressing technical challenges that need to be addressed. Summary of the Invention
[0004] The purpose of this invention is to provide a remote sensing image edge detection method based on deep learning. By deeply fusing large model semantic priors with edge features and cross-attention relocalization, it can output a high-precision, semantically consistent and closed target edge contour map.
[0005] To achieve the above objectives, this invention provides a remote sensing image edge detection method based on deep learning, the method comprising: S11. Input the remote sensing image into the adaptive large model to extract semantic cue feature maps; S12. Extract edge features from remote sensing images using a large edge detection model; S13. Fusion of semantic cue feature map and edge feature to enhance the edge representation of remote sensing image, and output remote sensing enhanced image; S14. Semantic localization of remote sensing enhanced images is performed through a cross-attention mechanism, and the target edge contour map of the remote sensing image is output.
[0006] Furthermore, the adaptive large model includes at least a preprocessing module consisting of two convolutional layers. Each convolutional layer has a kernel size of 3, a stride of 2, and padding of 1, and includes a normalization layer and a GELU activation function layer.
[0007] Furthermore, the remote sensing imagery is input into the adaptive large model to extract semantic cue feature maps, specifically including: S21. Divide the remote sensing image to be processed into several image blocks without overlap according to a preset size, and perform radiometric correction and normalization on the image blocks to obtain standardized image blocks. S22. Use a large visual model pre-trained on a remote sensing dataset to extract features from standardized image patches and output the corresponding high-dimensional feature maps. S23. Construct a learnable Prompt Pool, wherein the Prompt Pool includes K trainable cue vectors of dimension d; S24. Based on the similarity between the high-dimensional feature map and the cue vectors in the Prompt Pool, soft cue is dynamically retrieved and weighted to obtain soft cue; S25. Input the soft prompts into the prompt encoder, and simultaneously perform the text generation task and the contrastive learning task to generate discrete text descriptions and continuous semantic embeddings. S26. Deduplication and fusion are performed on the discrete text descriptions and continuous semantic embeddings of all image blocks to obtain a semantic cue dictionary, and pooling is performed on all soft cues to obtain a semantic cue feature map.
[0008] Furthermore, the large edge detection model includes four basic modules composed of preset stacking rules and downsampling layers between the basic modules. Each basic module includes a deformable convolutional function, a normalization layer, a feedforward neural network, and an activation function.
[0009] Furthermore, the processing flow of deformable convolution functions specifically includes: S31. Divide the input feature map into several groups according to the number of channels, and traverse each group one by one, and then traverse several sampling points in each group. S32. After obtaining the predefined network coordinates of each sampling point, add the offset of the group to which the sampling point belongs to obtain the actual sampling position. Then, use bilinear interpolation to obtain the channel vector at the actual sampling position in the sliced feature map. S33. Input the channel vector into the preset modulation weight subnetwork to obtain the modulation weight of the corresponding group and the corresponding sampling point. Multiply each channel value of the channel vector by the modulation weight to obtain the weighted result of the sampling point. S34. Accumulate the weighted results of all sampling points in the same group to obtain the contribution of each group. Accumulate the contributions of each group to obtain the channel vector of the output feature map at the corresponding spatial location.
[0010] Furthermore, the preset stacking rules and guidelines specifically include: The expression for the number of channels in stage i is:
[0011] The set expression for the deformable convolution operator in the i-th stage is:
[0012] The number of basic modules stacked in stages 1, 2, and 4 is equal, as expressed by:
[0013] The number of stacks in phase 1 is no greater than the number of stacks in phase 3, as expressed by:
[0014] in, This represents the channel number of the i-th stage. This indicates the group number of the deformable convolution operator in period i. The preset number of channels, This represents the number of stacked basic modules in stage i.
[0015] Furthermore, the expression for the four-stage basic module of the large edge detection model is as follows:
[0016] in, This is the feature map output at the current stage. The feature map input for the current stage, This is a global edge probability map. () indicates splicing along the channel dimension. () represents a deformable convolution function under multiple mechanisms. The number of layers for stacking the basic modules in each stage () represents a multilayer perceptron composed of multiple fully connected layers. () represents layer normalization.
[0017] Furthermore, edge features of remote sensing images are extracted using a large edge detection model, specifically including: S41. Perform band selection, normalization, and tile segmentation on the remote sensing image to obtain a standardized image tile sequence; S42. Input the image tile sequence into the edge detection model. The model sequentially passes through the first basic module to extract primary edge features, and the first downsampling layer reduces the spatial resolution of the feature map to 1 / 2. The second basic module extracts intermediate edge features, and the second downsampling layer reduces the spatial resolution of the feature map to 1 / 4. The third basic module extracts high-level edge features, and the third downsampling layer reduces the spatial resolution of the feature map to 1 / 8. The fourth basic module extracts semantic-level edge features. S43. After reducing the dimension of the feature map output by the fourth basic module through 1×1 convolution, perform bilinear upsampling to restore it to the input size and obtain the tile-level edge logits. S44. Perform overlapping region fusion on all tile edge logits to generate an initial edge probability map of the remote sensing image. S45. Perform thresholding, edge thinning, and vectorization on the initial edge probability map to output the edge probability map of the remote sensing image.
[0018] Furthermore, the semantic cue feature map is fused with edge features to enhance the edge representation of remote sensing images, specifically including: S51. After stitching the edge probability map with the original remote sensing image, the edge encoder generates an edge feature map corresponding to the semantic cue feature map. The semantic cue feature map and the edge feature map at each scale are input into the cross-scale interaction unit at the corresponding scale. S52. The cross-scale interaction unit uses the edge feature map as the query and the semantic prompt feature map as the key and value to calculate multi-head cross attention and obtain semantic edge alignment features. S53. Based on learnable gating weights, semantic edge alignment features and edge feature maps are fused pixel-wise through bottom-up and top-down cross-layer interaction paths to obtain fused edge features. S54. The fused edge features at each scale are downsampled and upsampled respectively, and then added to the features at adjacent scales to output the remote sensing enhanced image.
[0019] Furthermore, semantic localization of remote sensing enhanced images is performed through a cross-attention mechanism, specifically including: S61. Using the enhanced remote sensing image as a visual feature map, obtain the target semantic category corresponding to the visual feature map, and obtain the corresponding semantic query vector based on the target semantic category. S62. Expand the visual feature map into a key matrix according to the spatial dimension, and use the semantic query vector as the query vector. Perform scaling dot product cross attention calculation on the query vector and the key matrix to obtain the semantic response score. S63. The fused edge features are compressed through a preset channel to obtain the edge prior. The semantic response score is restored to a semantic localization heatmap with spatial dimension. The edge prior is used to spatially constrain and enhance the semantic localization heatmap to generate a remote sensing image target edge contour map with edge enhancement.
[0020] Compared with the prior art, the beneficial effects of the present invention are: This invention provides a deep learning-based remote sensing image edge detection method. Through the unified deep fusion of adaptive large-model semantic prompts and edge features of the edge detection large-model, it achieves complementary cross-scene generalization capabilities and geometric details, significantly reducing false and missed detections caused by illumination, shadows, and noise interference. Multi-scale fusion enhancement ensures continuous closure of remote sensing image edges, improving the boundary reliability of subsequent segmentation, change detection, and 3D reconstruction tasks. Cross-attention secondary semantic localization achieves strict alignment of pixel-level target edge contours with semantic categories. This invention, through end-to-end large-model full-link design, enables the direct output of high-precision, semantically consistent, and completely closed target edge contour maps without traditional post-processing, significantly reducing system complexity and deployment costs. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort. Figure 1 A schematic diagram of a deep learning-based remote sensing image edge detection method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the process of extracting semantic cue feature maps from remote sensing images by inputting them into an adaptive large model, as provided in an embodiment of the present invention. Figure 3 A schematic diagram of the processing flow of the deformable convolution function provided in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the process of extracting edge features from remote sensing images using a large edge detection model, as provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of the process of fusing semantic cue feature maps and edge features to enhance the edge representation of remote sensing images, provided by an embodiment of the present invention. Figure 6 This is a schematic diagram illustrating the semantic localization process of remote sensing enhanced images using a cross-attention mechanism, as provided in an embodiment of the present invention. Detailed Implementation
[0022] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0023] Reference Figure 1 This embodiment provides a remote sensing image edge detection method based on deep learning, the method comprising: S11. Input the remote sensing image into the adaptive large model to extract semantic cue feature maps.
[0024] S12. Extract edge features from remote sensing images using a large edge detection model.
[0025] S13. The semantic cue feature map and edge features are fused to enhance the edge representation of the remote sensing image, and the enhanced remote sensing image is output.
[0026] S14. Semantic localization of remote sensing enhanced images is performed through a cross-attention mechanism, and the target edge contour map of the remote sensing image is output.
[0027] In this embodiment, the image is fed into an adaptive large-scale model to obtain a semantic cue feature map that is universal across scenes. This provides category priors for all subsequent processing. The adaptive large-scale model can learn stronger feature representations and prior knowledge from a small amount of labeled data, accelerating the construction of a large-scale remote sensing image representation dataset. An edge detection large-scale model is used to quickly delineate edge features, supplementing the semantic cue feature map with fine geometric details. Semantic-driven processing helps the adaptive large-scale model better understand the semantic information contained in the input image, correctly distinguish between multiple types of targets to be detected and irrelevant background factors in the remote sensing image, and enhance feature differences and target representation information.
[0028] This approach deeply integrates semantic cue feature maps and edge features within a unified framework, preserving both category accuracy and enhancing edge continuity to output a remote sensing-enhanced image with both semantic and edge enhancements. Using this enhanced image as visual input, cross-attention is employed to refocus the semantic query vector on the edge map, rapidly locating the target and achieving dense detection. The final output is a pixel-level target edge contour map that strictly corresponds to the category. This method ensures both geometric accuracy and accurate semantic labeling of edges, resulting in a continuous, closed, and correctly categorized target edge contour map. This achieves high-precision, low-complexity, and cross-scene-compatible edge extraction.
[0029] As a preferred embodiment, the adaptive large model includes at least a preprocessing module consisting of two convolutional layers. Each convolutional layer has a kernel size of 3, a stride of 2, and padding of 1, and includes a normalization layer and a GELU activation function layer.
[0030] In this embodiment, the adaptive large model acquires multi-level features of the input image. A preprocessing module consisting of two convolutional layers scales the features to an appropriate size. Each convolutional layer has a kernel size of 3, a stride of 2, and padding of 1, and includes a normalization layer and a GELU activation function layer. The adaptive large model reduces the pixel count of the original remote sensing image by a factor of 16 through two 3×3 downsampling operations, preserving the spectral-texture details of the 3×3 neighborhood. GELU and normalization ensure smooth deep gradients, enabling the adaptive large model to reconstruct high-fidelity semantic features even when processing lower-resolution remote sensing images. Simultaneously, 4×4 patch embedding is performed in one step, avoiding information loss due to additional mapping, resulting in halved GPU memory usage, faster training convergence, and almost no impact on the spatial-spectral accuracy of the semantic cue feature map.
[0031] As a preferred embodiment, remote sensing images are input into an adaptive large model to extract semantic cue feature maps, specifically including: S21. Divide the remote sensing image to be processed into several image blocks without overlap according to a preset size, and perform radiometric correction and normalization on the image blocks to obtain standardized image blocks.
[0032] S22. Use a large visual model pre-trained on a remote sensing dataset to extract features from standardized image patches and output the corresponding high-dimensional feature maps.
[0033] S23. Construct a learnable Prompt Pool, which includes K trainable cue vectors of dimension d.
[0034] S24. Based on the similarity between the high-dimensional feature map and the cue vectors in the Prompt Pool, soft cue is obtained by dynamically retrieving and weighting the data.
[0035] S25. Input soft prompts into the prompt encoder, and simultaneously perform text generation and contrastive learning tasks to generate discrete text descriptions and continuous semantic embeddings.
[0036] S26. Deduplication and fusion are performed on the discrete text descriptions and continuous semantic embeddings of all image blocks to obtain a semantic cue dictionary, and pooling is performed on all soft cues to obtain a semantic cue feature map.
[0037] In this embodiment, refer to Figure 2After standardizing large-scale remote sensing image blocks, high-dimensional features are extracted using a pre-trained remote sensing model. Then, soft cues are generated by dynamically retrieving and weighting the block content through a learnable Prompt Pool. This drives the cue encoder to simultaneously output discrete text descriptions and continuous semantic embeddings. Finally, after deduplication and fusion, a semantic cue dictionary and pooled feature maps of the entire image are obtained. Thus, the knowledge of a general large model is refocused on remote sensing tasks with very few parameters, taking into account zero-shot generalization, computational savings, and readability, achieving efficient semantic abstraction and global indexing of complex remote sensing images.
[0038] As a preferred embodiment, the large edge detection model includes four basic modules composed of preset stacking rules and downsampling layers between the basic modules. Each basic module includes a deformable convolutional function, a normalization layer, a feedforward neural network, and an activation function.
[0039] In this embodiment, the basic modules composed of four stages of preset stacking rules in the edge detection large model, along with the downsampling layers between these basic modules, are used to extract multi-scale features from the input remote sensing image. Long-distance dependent semantic information and adaptive spatial clustering are fundamental to accurate target feature extraction. Deformable convolution functions reduce the strong inductive bias of regular convolutions while utilizing 3×3 convolution kernels to achieve sparse sampling, improving computational and memory efficiency. The sampling offsets for long-distance dependencies and adaptive spatial clustering are learnable and subject to input constraints, effectively improving the extraction of rich semantic information from the receptive field.
[0040] As a preferred embodiment, the processing flow of the deformable convolution function specifically includes: S31. Divide the input feature map into several groups according to the number of channels, and traverse each group one by one, and then traverse several sampling points in each group.
[0041] S32. After obtaining the predefined network coordinates of each sampling point, add the offset of the group to which the sampling point belongs to obtain the actual sampling position. Then, use bilinear interpolation to obtain the channel vector at the actual sampling position in the sliced feature map.
[0042] S33. Input the channel vector into the preset modulation weight subnetwork to obtain the modulation weight of the corresponding group and the corresponding sampling point. Multiply each channel value of the channel vector by the modulation weight to obtain the weighted result of the sampling point.
[0043] S34. Accumulate the weighted results of all sampling points in the same group to obtain the contribution of each group. Accumulate the contributions of each group to obtain the channel vector of the output feature map at the corresponding spatial location.
[0044] In this embodiment, refer to Figure 3The position of each sampling point is determined by a predefined grid and learnable intra-group offsets. The network adaptively allows the receptive field to follow changes in the target's shape, size, or pose, significantly improving its ability to characterize irregular geometric structures. After channel grouping, each group has independent offset and modulation weights, essentially providing dedicated convolutional kernels for features with different semantics or spatial frequencies, reducing inter-group interference and improving discriminative power. The grouping mechanism and per-sample-point modulation allow for controllable expansion of the parameter count and receptive field; the number of groups and sampling points K can be adjusted, flexibly balancing accuracy and computation. Both offset and modulation are generated by lightweight bypass networks, resulting in minimal overhead.
[0045] Compared to using full-size convolutions on the entire feature map, grouped convolutions divide the input channels into G parts, with each group processing only C / G channels (C = total number of channels, G = number of groups), reducing the number of parameters and computational cost to approximately 1 / G. Furthermore, the shared offset / modulation structure further reduces redundancy. Grouping restricts the input dimension of each convolutional kernel, effectively introducing regularization in the channel dimension; simultaneously, the offset and modulation weights are explicitly constrained to prevent extreme sampling. This significantly enhances the characterization of irregular geometry while drastically reducing the number of parameters and computational cost through grouped convolutions, balancing accuracy, efficiency, and generalization.
[0046] The processing flow expression for deformable convolution functions is as follows:
[0047] in, For the current pixel, This represents the total number of polymeric groups. The number of polymer groups. The total number of sampling points. Number of sampling points For the k-th position of the predefined grid sampling, For the g-th group, This represents the location-independent projection weights. Let be the dimension of this group. This represents the modulation scalar of the k-th sample in the g-th group, normalized along the K-dimensional axis using the softmax activation function. For the input feature mapping of the slice, Sampling position of the g-th grid The corresponding offset.
[0048] As a preferred embodiment, the preset stacking rules specifically include: The expression for the number of channels in stage i is:
[0049] The set expression for the deformable convolution operator in the i-th stage is:
[0050] The number of basic modules stacked in stages 1, 2, and 4 is equal, as expressed by:
[0051] The number of stacks in phase 1 is no greater than the number of stacks in phase 3, as expressed by:
[0052] in, This represents the channel number of the i-th stage. This indicates the group number of the deformable convolution operator in period i. The preset number of channels, This represents the number of stacked basic modules in stage i.
[0053] In this embodiment, the preset stacking rules allow the network to progress in depth, width, and number of groups at different stages. This ensures that shallow layers efficiently extract local edges, middle layers fully mix semantics, and deep layers capture the global structure. Furthermore, the linear matching of the number of groups and channels controls the amount of parameters and computation, achieving a balance between performance, efficiency, and scalability.
[0054] As a preferred embodiment, the expression for the four-stage basic module of the large edge detection model is:
[0055] in, This is the feature map output at the current stage. The feature map input for the current stage, This is a global edge probability map. () indicates splicing along the channel dimension. () represents a deformable convolution function under multiple mechanisms. The number of layers in which the basic modules are stacked at each stage. () represents a multilayer perceptron composed of multiple fully connected layers. () represents layer normalization.
[0056] In this embodiment, a deformable convolution function under multiple mechanisms is first used to allow the convolution kernel to adaptively move with the edge shape, capturing fine local geometry; then, the original features and deformed features are concatenated along the channel dimension, preserving low-frequency semantics while enhancing high-frequency edges; a multilayer perceptron composed of multiple fully connected layers performs nonlinear fusion along the channel dimension, further expanding the receptive field. After each iteration, the receptive field is expanded exponentially, eliminating the need to deepen ordinary convolutional layers, resulting in a lightweight and efficient process. After each stage, the feature map output from the current stage is concatenated with the outputs from all previous stages, which is equivalent to accumulating edge cues at each scale layer by layer into the global edge probability map. This residual accumulation avoids the amplification of single-scale errors, making the final global edge probability map both detailed and with complete outlines.
[0057] As a preferred embodiment, edge features of remote sensing images are extracted using a large edge detection model, specifically including: S41. Perform band selection, normalization, and tile segmentation on the remote sensing image to obtain a standardized image tile sequence.
[0058] S42. Input the image tile sequence into the edge detection model. The first basic module extracts primary edge features, and the first downsampling layer reduces the spatial resolution of the feature map to 1 / 2. The second basic module extracts intermediate edge features, and the second downsampling layer reduces the spatial resolution of the feature map to 1 / 4. The third basic module extracts high-level edge features, and the third downsampling layer reduces the spatial resolution of the feature map to 1 / 8. The fourth basic module extracts semantic edge features.
[0059] S43. After reducing the dimension of the feature map output by the fourth basic module through 1×1 convolution, perform bilinear upsampling to restore it to the input size and obtain tile-level edge logits.
[0060] S44. Perform overlapping region fusion on all tile edge logits to generate an initial edge probability map of the remote sensing image.
[0061] S45. Perform thresholding, edge thinning, and vectorization on the initial edge probability map to output the edge probability map of the remote sensing image.
[0062] Thresholding, edge thinning, and vectorization are performed on the initial edge probability map to output the edge probability map of the remote sensing image.
[0063] In this embodiment, refer to Figure 4Addressing the challenges of numerous spectral bands, large image sizes, and complex edge types in remote sensing images, this study employs band selection and normalization to filter effective spectral information and eliminate radiometric interference. Tile segmentation solves the problem of directly inputting large-size images into the model, laying the foundation for accurate edge feature extraction. A multi-basic module and downsampling layer are used to progressively extract edge features from the primary to the semantic level. Downsampling allows the model to capture more global semantic relationships, while different levels of feature extraction adapt to the needs of edges at different scales in remote sensing images. 1×1 convolution dimensionality reduction simplifies parameters and improves efficiency, while bilinear upsampling accurately restores the original image size, ensuring that the tile-level edge logits precisely correspond to the input position. Overlapping region fusion avoids edge breaks caused by tile segmentation, ensuring the global continuity of the initial edge probability map. Finally, thresholding, edge refinement, and vectorization remove redundant noise and optimize edge detail accuracy, ensuring that the output edge probability map conforms to human visual perception of edges and meets the needs of subsequent remote sensing applications for precise edge information. Overall, this approach achieves high efficiency, accuracy, and practicality in remote sensing image edge extraction.
[0064] As a preferred embodiment, fusing semantic cue feature maps with edge features enhances the edge representation of remote sensing images, specifically including: S51. After stitching the edge probability map with the original remote sensing image, the edge encoder generates an edge feature map corresponding to the semantic cue feature map. The semantic cue feature map and the edge feature map at each scale are input into the cross-scale interaction unit at the corresponding scale.
[0065] S52. The cross-scale interaction unit uses the edge feature map as the query and the semantic prompt feature map as the key and value to calculate multi-head cross attention and obtain semantic edge alignment features.
[0066] S53. Based on learnable gating weights, semantic edge alignment features and edge feature maps are fused pixel-wise through bottom-up and top-down cross-layer interaction paths to obtain fused edge features.
[0067] S54. The fused edge features at each scale are downsampled and upsampled respectively, and then added to the features at adjacent scales to output the remote sensing enhanced image.
[0068] In this embodiment, refer to Figure 5By designing multiple stages, this approach precisely addresses the semantic ambiguity and lack of detail in remote sensing image edge representation, significantly improving the practicality and accuracy of edge information. It stitches edge probability maps with the original images to generate edge feature maps at corresponding scales, ensuring precise matching of edge features and semantic cue features in terms of dimension and spatial scale, laying the foundation for cross-modal interaction. The cross-scale interaction unit uses edge features as queries and semantic cue features as keys to calculate multi-head cross-attention, establishing a precise association between edge features and ground feature semantics, avoiding the semantic ambiguity problem in traditional edge extraction where only the outline is known, not the category.
[0069] Based on learnable gating weights, bottom-up and top-down cross-layer interactions dynamically adjust the fusion ratio of semantic information and edge details, avoiding the limitations of fixed fusion methods. Finally, by downsampling and upsampling edge features at each scale and adding them to adjacent scales, complementary edge information at different scales is achieved. The resulting remote sensing enhanced image has both clear spatial contour details and clear semantic attributes of ground features, providing more accurate and interpretable edge support for subsequent remote sensing tasks and significantly reducing the probability of edge misjudgment and category confusion in subsequent tasks.
[0070] As a preferred embodiment, semantic localization of remote sensing enhanced images is performed through a cross-attention mechanism, specifically including: S61. Using the enhanced remote sensing image as a visual feature map, obtain the target semantic category corresponding to the visual feature map, and obtain the corresponding semantic query vector based on the target semantic category.
[0071] S62. Expand the visual feature map into a key matrix according to the spatial dimension, and use the semantic query vector as the query vector. Perform scaling dot product cross attention calculation on the query vector and the key matrix to obtain the semantic response score.
[0072] S63. The fused edge features are compressed through a preset channel to obtain the edge prior. The semantic response score is restored to a semantic localization heatmap with spatial dimension. The edge prior is used to spatially constrain and enhance the semantic localization heatmap to generate a remote sensing image target edge contour map with edge enhancement.
[0073] In this embodiment, refer to Figure 6By combining cross-attention mechanisms with edge prior constraints, this method addresses the issues of semantic and spatial misalignment and blurred target contours in remote sensing image semantic localization, significantly improving localization accuracy and practicality. Using enhanced remote sensing imagery as a visual feature map, a dedicated semantic query vector is obtained by combining the target's semantic category, allowing the localization task to focus on specific targets and avoiding generalization or omissions caused by indiscriminate searching. The visual feature map is expanded into a key matrix along the spatial dimension and a scaling dot product cross-attention calculation is performed. Through precise matching of semantic queries and spatial features, a semantic response score reflecting the semantic relevance of the target at different spatial locations is generated, effectively solving the localization offset problem caused by the scattered distribution of similar targets and the interweaving of adjacent different targets in remote sensing imagery.
[0074] By compressing the fused edge features into edge priors, the semantic localization heatmap for restoring spatial dimensions is constrained and enhanced. This not only utilizes edge priors to clarify the spatial contour boundaries of targets but also strengthens the semantic response at the target edges. The resulting edge-enhanced target contour map possesses both accurate semantic orientation and clear spatial contour boundaries. It can directly provide precise localization support in three dimensions—semantics, space, and edge—for downstream tasks such as remote sensing target extraction, land feature boundary delineation, and change detection, significantly reducing the probability of target misdetection and boundary misjudgment in subsequent tasks.
[0075] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A remote sensing image edge detection method based on deep learning, characterized in that, The method includes: S11. Input the remote sensing image into the adaptive large model to extract semantic cue feature maps, specifically including: S21. Divide the remote sensing image to be processed into several image blocks without overlap according to a preset size, and perform radiometric correction and normalization on the image blocks to obtain standardized image blocks. S22. Use a large visual model pre-trained on a remote sensing dataset to extract features from standardized image patches and output the corresponding high-dimensional feature maps. S23. Construct a learnable Prompt Pool, wherein the Prompt Pool includes K trainable cue vectors of dimension d; S24. Based on the similarity between the high-dimensional feature map and the cue vectors in the Prompt Pool, soft cue is dynamically retrieved and weighted to obtain soft cue; S25. Input the soft prompts into the prompt encoder, and simultaneously perform the text generation task and the contrastive learning task to generate discrete text descriptions and continuous semantic embeddings. S26. Deduplication and fusion of discrete text descriptions and continuous semantic embeddings for all image blocks are performed to obtain a semantic cue dictionary, and pooling is performed on all soft cues to obtain a semantic cue feature map. S12. Extract edge features from remote sensing images using a large-scale edge detection model. This model comprises four basic modules based on pre-defined stacking rules, and downsampling layers between these modules. Each basic module includes a deformable convolutional function, a normalization layer, a feedforward neural network, and an activation function. The processing flow of the deformable convolutional function specifically includes: S31. Divide the input feature map into several groups according to the number of channels, and traverse each group one by one, and then traverse several sampling points in each group. S32. After obtaining the predefined network coordinates of each sampling point, add the offset of the group to which the sampling point belongs to obtain the actual sampling position. Then, use bilinear interpolation to obtain the channel vector at the actual sampling position in the sliced feature map. S33. Input the channel vector into the preset modulation weight subnetwork to obtain the modulation weight of the corresponding group and the corresponding sampling point. Multiply each channel value of the channel vector by the modulation weight to obtain the weighted result of the sampling point. S34. Accumulate the weighted results of all sampling points in the same group to obtain the contribution of each group. Accumulate the contributions of each group to obtain the channel vector of the output feature map at the corresponding spatial location. S13. Fusion of semantic cue feature map and edge feature to enhance the edge representation of remote sensing image, and output remote sensing enhanced image; S14. Semantic localization of remote sensing enhanced images is performed through a cross-attention mechanism, and the target edge contour map of the remote sensing image is output.
2. The remote sensing image edge detection method based on deep learning according to claim 1, characterized in that, The adaptive large model includes at least a preprocessing module consisting of two convolutional layers. Each convolutional layer has a kernel size of 3, a stride of 2, and padding of 1, and includes a normalization layer and a GELU activation function layer.
3. The remote sensing image edge detection method based on deep learning according to claim 1, characterized in that, Preset stacking rules and guidelines, specifically including: The expression for the number of channels in stage i is: The set expression for the deformable convolution operator in the i-th stage is: The number of basic modules stacked in stages 1, 2, and 4 is equal, as expressed by: The number of stacks in phase 1 is no greater than the number of stacks in phase 3, as expressed by: in, This represents the channel number of the i-th stage. This indicates the group number of the deformable convolution operator in period i. The preset number of channels, This represents the number of stacked basic modules in stage i.
4. The remote sensing image edge detection method based on deep learning according to claim 1, characterized in that, The expressions for the four basic modules of the large edge detection model are as follows: in, This is the feature map output at the current stage. The feature map input for the current stage, This is a global edge probability map. () indicates splicing along the channel dimension. () represents a deformable convolution function under multiple mechanisms. The number of layers for stacking the basic modules in each stage () represents a multilayer perceptron composed of multiple fully connected layers. () represents layer normalization.
5. The remote sensing image edge detection method based on deep learning according to claim 1, characterized in that, Edge features of remote sensing images are extracted using a large edge detection model, specifically including: S41. Perform band selection, normalization, and tile segmentation on the remote sensing image to obtain a standardized image tile sequence; S42. Input the image tile sequence into the edge detection model. The model sequentially passes through the first basic module to extract primary edge features, and the first downsampling layer reduces the spatial resolution of the feature map to 1 / 2. The second basic module extracts intermediate edge features, and the second downsampling layer reduces the spatial resolution of the feature map to 1 / 4. The third basic module extracts high-level edge features, and the third downsampling layer reduces the spatial resolution of the feature map to 1 / 8. The fourth basic module extracts semantic-level edge features. S43. After reducing the dimension of the feature map output by the fourth basic module through 1×1 convolution, perform bilinear upsampling to restore it to the input size and obtain the tile-level edge logits. S44. Perform overlapping region fusion on all tile edge logits to generate an initial edge probability map of the remote sensing image. S45. Perform thresholding, edge thinning, and vectorization on the initial edge probability map to output the edge probability map of the remote sensing image.
6. The remote sensing image edge detection method based on deep learning according to claim 5, characterized in that, The fusion of semantic cue feature maps and edge features enhances the edge representation of remote sensing images, specifically including: S51. After stitching the edge probability map with the original remote sensing image, the edge encoder generates an edge feature map corresponding to the semantic cue feature map. The semantic cue feature map and the edge feature map at each scale are input into the cross-scale interaction unit at the corresponding scale. S52. The cross-scale interaction unit uses the edge feature map as the query and the semantic prompt feature map as the key and value to calculate multi-head cross attention and obtain semantic edge alignment features. S53. Based on learnable gating weights, semantic edge alignment features and edge feature maps are fused pixel-wise through bottom-up and top-down cross-layer interaction paths to obtain fused edge features. S54. The fused edge features at each scale are downsampled and upsampled respectively, and then added to the features at adjacent scales to output the remote sensing enhanced image.
7. The remote sensing image edge detection method based on deep learning according to claim 6, characterized in that, Semantic localization of remote sensing augmented images using a cross-attention mechanism specifically includes: S61. Using the enhanced remote sensing image as a visual feature map, obtain the target semantic category corresponding to the visual feature map, and obtain the corresponding semantic query vector based on the target semantic category. S62. Expand the visual feature map into a key matrix according to the spatial dimension, and use the semantic query vector as the query vector. Perform scaling dot product cross attention calculation on the query vector and the key matrix to obtain the semantic response score. S63. The fused edge features are compressed through a preset channel to obtain the edge prior. The semantic response score is restored to a semantic localization heatmap with spatial dimension. The edge prior is used to spatially constrain and enhance the semantic localization heatmap to generate a remote sensing image target edge contour map with edge enhancement.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on semantic adaptive edge enhancement network
CN118781596A
Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment
WO2024230038A1