Remote sensing large model construction method fusing edge semantic knowledge

By constructing the GGB-SAMNet network and integrating global semantics with local boundary information, the problems of fracture and ambiguity in boundary areas of the remote sensing semantic segmentation model are solved, the segmentation accuracy and structural integrity of remote sensing images are improved, and the segmentation performance in complex scenarios is adapted.

CN120726320APending Publication Date: 2025-09-30ZHENGZHOU UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510771447.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing remote sensing semantic segmentation models have problems such as blurred edges, unclear contours, and prediction breaks when processing boundary areas of objects. It is difficult to balance the detail preservation of small-scale targets with the overall semantic expression of large-scale targets. The boundary information and semantic information are not tightly integrated, affecting the integrity and accuracy of the segmentation results.

Method used

A dual-branch architecture of semantic branch and boundary branch is adopted, and a boundary generation module, a multi-scale boundary refinement structure and a boundary semantic fusion mechanism are introduced. The global semantic and local boundary information are integrated through the GGB-SAMNet network to enhance the model's recognition accuracy and segmentation continuity of boundary areas.

Benefits of technology

It improves the expression accuracy and structural integrity of remote sensing image semantic segmentation in complex land feature boundary scenes, improves segmentation accuracy and the ability to adapt to complex scenes, especially showing higher continuity and integrity in boundary areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726320A_ABST
    Figure CN120726320A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing large model construction method fusing edge semantic knowledge, and the method comprises the following steps: S1, carrying out the data preprocessing and enhancement of a strategy for later use; s2, feature extraction and fusion network design based on boundary enhancement; s3, training and evaluating the data set processed in the step S1 through the model constructed in the step S2; according to the method, a double-branch architecture of semantic branches and boundary branches is adopted, global semantic information and local boundary information are fully integrated, and a segmentation framework with global guidance, boundary refinement and semantic alignment capabilities is constructed. By introducing a boundary generation module, a multi-scale boundary refining module and a boundary semantic fusion module, the recognition precision and segmentation continuity of the model to a boundary region are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image technology, and in particular to a method for constructing a large remote sensing model integrating edge semantic knowledge. Background Art

[0002] With the continuous advancement of remote sensing technology and high-resolution imaging equipment, remote sensing imagery is increasingly being used in fields such as agricultural monitoring, land use, urban planning, and ecological protection. Remote sensing semantic segmentation, as a crucial component of intelligent remote sensing interpretation, is becoming a key supporting technology for remote sensing information extraction and analysis. Its primary goal is to classify each pixel in a remote sensing image into semantically meaningful categories, thereby enabling refined identification and structured representation of ground objects.

[0003] Currently, remote sensing semantic segmentation tasks face numerous challenges, particularly the diverse variety of features, significant scale differences, blurred boundaries, and dense distribution of small objects, which greatly complicate modeling. Traditional methods often experience prediction discontinuities and gaps in boundary regions, leading to fragmented and incoherent semantic segmentation results, severely limiting the reliability and generalization of the models in practical applications. Therefore, enhancing the model's ability to perceive complex boundary regions while improving overall semantic modeling capabilities has become a key research topic in remote sensing semantic segmentation.

[0004] Remote sensing imagery, due to its wide coverage, complex ground structures, and dramatic scale variations, places high demands on the structural perception and boundary modeling capabilities of semantic segmentation models. However, most current mainstream remote sensing semantic segmentation models focus on constructing global semantics, ignoring the key role of boundary information in structural expression. This leads to the following potential problems:

[0005] First, existing models lack explicit boundary modeling mechanisms during feature extraction. Although large models such as Transformer have strong global modeling capabilities, they often exhibit problems such as blurred edges, unclear contours, and prediction breaks when processing boundary areas of objects, affecting the integrity and accuracy of segmentation results. References: He X, Zhou Y, Liu B, et al. Remote sensing image semantic segmentation via class-guided structural interaction and boundary perception [J]. Expert Systems with Applications, 2024, 252: 124019. Ma X, Wu Q, Zhao X, et al. Sam-assisted remote sensing imagery semantic segmentation with object and boundary constraints [J]. IEEE Transactions onGeoscience and Remote Sensing, 2024. Gao J, Zhou C, Xu G, et al.Multiscale sea-land segmentation networks for weak boundaries[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024.

[0006] Secondly, the scales of objects in remote sensing images vary significantly, making it difficult to balance the preservation of details in small-scale objects with the overall semantic expression of large-scale objects using a single-scale modeling approach. While existing methods incorporate multi-scale structures, semantic discontinuity and misclassification within regions often occur at boundaries, particularly at the intersections of urban buildings, roads, and water bodies. (See Zhang C, Atkinson PM, George C, et al. Identifying and mapping individual plants in a highly diverse high-elevation ecosystem using UAV imagery and deep learning [J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 169: 280-291.) Zheng Z, Zhong Y, Wang J, et al. Foreground-aware relation network forgeospatial object segmentation in high spatial resolution remote sensingimagery[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020:4096-4105. Long J,Shelhamer E,Darrell T.Ful ly convolutional networks forsemantic segmentation[C] / / Proceedings of the IEEE conference on computervision and pattern recognition.2015:3431-3440. Ronneberger O, Fischer P, Brox TU-net: Convolutional networks for biomedical image segmentation[C] / / Medical image computing and computer-assisted intervention–MICCAI 2015:18th international conference,Munich,Germany,October 5-9,2015,proceedings,part III 18. Springer InternationalPublishing,2015:234-241.

[0007] Furthermore, the fusion of boundary information and semantic information is not close enough. Although some methods introduce boundary auxiliary modules or supervision signals, they fail to achieve deep coordination with the backbone network. As a result, boundary features are gradually weakened in high-level representations, making it difficult to effectively constrain the final segmentation results, and reducing the model's ability to express structurally sensitive areas. (The following is the reference: Yeung CC, Lam K M. Attentive boundary-aware fusion for defect semantic segmentation using transformer [J]. IEEE Transactions on Instrumentation and Measurement, 2023, 72: 1-13.) Zhang J, Shao M, Wan Y, et al. Boundary-aware spatial and frequency dual-domain transformer for remote sensing urban images segmentation[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024.

[0008] Therefore, a remote sensing large model construction method integrating edge semantic knowledge is provided to solve the above technical problems. Summary of the Invention

[0009] The purpose of this invention is to provide a method for constructing a large remote sensing model that integrates edge semantic knowledge. This method employs a dual-branch architecture of semantic and boundary branches to fully integrate global semantics and local boundary information, constructing a segmentation framework with global guidance, boundary refinement, and semantic alignment capabilities. By introducing a boundary generation module, a multi-scale boundary refinement structure, and a boundary semantic fusion mechanism, the model enhances the recognition accuracy and segmentation continuity of boundary regions.

[0010] The object of the present invention is achieved like this: A method for constructing a large remote sensing model integrating edge semantic knowledge includes the following steps: S1, data preprocessing and enhancement strategy for backup; S2. Feature extraction and fusion network design based on boundary enhancement: GGB-SAMNet structure and boundary guidance mechanism design. After completing remote sensing image preprocessing, standardization, and multi-source enhancement operations, the semantic modeling network structure construction phase begins. By constructing a fusion network architecture with global perception and boundary detail modeling capabilities, the expression accuracy and structural integrity of remote sensing image semantic segmentation in complex land feature boundary scenarios are improved. The structure is based on the SegmentAnything Model (SAM) as the backbone, introduces a boundary enhancement path, and embeds a boundary perception module in the semantic branch backbone to achieve coordinated optimization of semantic modeling and boundary supervision. The overall network is named GGB-SAMNet. S3. Train and evaluate the data set processed in step S1 using the model constructed in step S2, focusing on selecting the optimal model and performing comparative analysis on the performance of each model under different indicators.

[0011] The specific process of S1 is as follows: S1.1. Establish a standardized data input process for remote sensing image preprocessing and enhancement. During the data preparation phase, representative public remote sensing semantic segmentation datasets, including LoveDA, GID, and ISPRS Potsdam, were selected to cover a variety of typical remote sensing scenes, ensuring diversity and generalization in model training. To ensure uniformity in input data, all original remote sensing images were cropped or scaled to 512×512 pixels to maintain consistent image resolution. They were then normalized using fixed statistical parameters for the three image channels, with the mean values ​​for the red, green, and blue channels being [0.485, 0.456, 0.406] and the standard deviations being [0.229, 0.224, 0.225], respectively. This normalization process eliminates the effects of brightness, color, and sensor variability in remote sensing images, improving the model's ability to perceive the image's essential characteristics.

[0012] S1.2. To construct a representative set of training samples, a random stratified sampling strategy was used to divide the original data into a training set, a validation set, and a test set, with a division ratio of approximately 7:1.5:1.5. This ensured that each subset was balanced in terms of the number of categories, spatial distribution, and sample characteristics, and avoided the impact of data bias on model training. In addition, to enhance the model's robustness to factors such as scale changes, morphological changes, and imaging perturbations in remote sensing images, a data augmentation strategy was implemented on the samples in the training set. The enhancement methods included, but were not limited to, a combination of rotation, scaling, mirroring, color perturbation, and random cropping. Geometric transformation operations helped improve the model's adaptability to changes in the target space's morphology, and color perturbation operations enhanced the model's performance under different lighting and seasonal conditions. Overall, the distribution range of training samples in the feature space was significantly expanded.

[0013] S1.3. Considering that the GGB-SAMNet structure proposed in the present invention has a dual-branch design of semantics and boundary, the input image is processed differently according to the branch feature extraction requirements; specifically, the semantic branch needs to capture global semantic information, and its input image size is set to 256×256 pixels, while the boundary branch focuses on the recovery of detailed structures and edge information extraction, retaining the higher resolution original input of 512×512 pixels to enhance spatial boundary details; before the image is input into the network, normalization and standardization preprocessing are performed on the images of the two branches respectively to ensure the consistency and comparability of the image input features in the distribution dimension.

[0014] S1.4. To provide reasonable prior information support for the boundary modeling branch, a boundary label map generation mechanism is introduced during the training sample preparation phase. The Canny edge detection algorithm is used to extract edges from the original RGB remote sensing imagery, generating a binary boundary map as the initial boundary feature map B1. This binary boundary map preserves boundary information in areas of significant variation between features and provides a stable supervisory signal for the boundary branch in the early stages of training, preventing overfitting to noisy areas. Furthermore, the binary boundary map serves as input features for the subsequent refinement module of the boundary branch, supporting the refined modeling process.

[0015] The specific process of S2 is as follows: S2.1. A globally guided boundary enhancement strategy is proposed to compensate for the broken or blurred boundary regions caused by the SAM encoder's serialized segmentation of the image. From the perspective of accurate extraction and fusion of boundary information, this paper proposes a globally guided boundary enhancement strategy. The globally guided boundary enhancement strategy obtains a refined boundary representation by introducing independent boundary branches and deeply integrates them with the backbone network features at various stages to achieve global enhancement of boundary information. S2.2. Using a boundary generation module; in the boundary generation module, the Canny operator is used to generate boundary features B1 for the input image. Canny is a gradient operator in image processing that identifies the boundaries between different objects in the image by calculating the grayscale change along a specific direction. S2.3. Construct a boundary refinement module based on multi-scale feature fusion; In order to further improve the accuracy and fine-grained expression ability of boundaries, this paper proposes a boundary refinement module based on multi-scale feature fusion. The boundary refinement module based on multi-scale feature fusion enhances the true expression of boundaries through multi-scale feature compression and reconstruction; specifically as follows: The boundary refinement module adopts an encoder-decoder architecture and combines jump connections to effectively improve segmentation performance. Figure 1 As shown in the boundary refinement module in the figure, the encoder and decoder of the boundary refinement module based on multi-scale feature fusion are composed of three convolution blocks, each of which contains two convolution operations, a batch normalization (BatchNorm) layer and a ReLU activation function. In the encoder part, the network gradually increases the number of convolution channels and combines the maximum pooling layer to reduce the spatial resolution, thereby extracting and compressing the local information of the boundary features, providing a clearer boundary representation for the subsequent refinement process. The decoder part restores the spatial resolution through upsampling operations and accurately restores the boundary details in the image through skip connections and fusion with low-level features. The boundary refinement module can effectively fuse features of different scales, significantly improving the accuracy and completeness of boundary recovery, which not only enhances the ability to recover boundary details, but also effectively improves the accuracy of segmentation results. S2.4. Construct a boundary-semantic adaptive fusion module. Traditional feature fusion modules usually use the same weights when fusing semantic features and boundary features at different stages. This method fails to fully consider the differences between features at different stages, which leads to unbalanced information fusion. Especially when processing remote sensing images with complex boundaries, it cannot effectively enhance the expression of boundary information. In order to make full use of the boundary features extracted from the boundary refinement module and further improve the model performance, this paper proposes a boundary-semantic adaptive fusion module. The boundary-semantic adaptive fusion module effectively integrates the explicit boundary features from the boundary branch by using the attention mechanism, aligns them with the intermediate semantic features extracted by the main branch, and dynamically adjusts the weights, significantly enhancing the role of boundary information in the global scope, thereby further improving the performance of the model. S2.5. Constructing a Transformer module based on boundary enhancement. To improve the local boundary perception capability of the SAM encoder in remote sensing image analysis, this paper embeds a boundary-semantic adaptive fusion module (BSAF) in each Transformer block, thereby enhancing the segmentation performance of the model in complex remote sensing scenes. Specifically, while maintaining the basic image encoder architecture of the original SAM model, all original network parameters are frozen. In each Transformer module, fine boundary features are introduced into the network, and multiple simple and efficient adapters are embedded in parallel to enhance the model's adaptability to remote sensing imagery. S2.6. Construct an enhanced mask decoder. In order to improve the accurate modeling of boundaries and semantics in the semantic segmentation task of remote sensing images, this paper proposes an enhanced mask decoder. The enhanced mask decoder fully utilizes the important role of boundary information in semantic modeling by fusing the semantic features with accurate boundary information extracted by the network encoder and the hint embedding information, significantly improving the segmentation accuracy and detail recovery ability. Specifically: The output features with good boundary modeling generated by the encoder and the hint embedding features are jointly input to the prediction head for processing. The overall architecture is as follows: Figure 4As shown in the figure; first, two bidirectional Transformer modules are used to efficiently model all embedded features to fully explore the rich information in the embedded features and significantly enhance the expressiveness of the features. The self-attention mechanism is used to model the hint embedding features to effectively capture their contextual relevance; secondly, the bidirectional cross-attention mechanism establishes an efficient interactive relationship between the semantic features with precise boundary information and the hint embedding, effectively strengthening the long-range dependency between boundary details and semantic information, thereby significantly improving boundary clarity and the accuracy of semantic understanding; then, the class-level convolution module efficiently upsamples the fused features to further refine the boundary information and generate fine semantic masks for multiple category channels to ensure the accurate expression and coherence of the boundary area in the high-resolution mask; finally, the prediction head efficiently combines the category information mask with the tokens generated by the multi-layer perceptron (MLP) through the dot product operation, and outputs a high-resolution multi-category semantic segmentation mask, which achieves accurate classification of each pixel while maintaining the detailed structure and coherence of complex boundaries; S2.7. Construct a loss function. For data with imbalanced categories, use the weighted cross entropy loss function, as follows: In the formula, N represents the total number of categories; i is the index of the category, from 1 to N; w i is the weight value of the i-th category, which is used to balance the importance between categories and is often used to deal with category imbalance problems; i is the true label of the i-th category; w i is the original prediction value of the model for the i-th type of output (that is, the output without activation function); σ(logit i ) represents the logit i Apply the Sigmoid function to calculate the probability that the prediction belongs to this category. The Sigmoid function is defined as In addition to the main semantic segmentation loss function, this paper also defines boundary loss to supervise the generation of boundary maps. The boundary pixels of objects are significantly less than the non-boundary pixels, so boundary prediction is a problem of imbalanced class distribution. Binary cross entropy loss is used for joint optimization of boundary learning. Where, and p i Represents the true boundary and the predicted boundary; The total loss is calculated as follows: L total =L seg +α×L bound (14) Among them, L total is the total loss, Lseg and L bound Represent segmentation loss and boundary loss respectively; L bound The hyperparameter α before represents its proportion in the total loss; after experimental analysis, the best segmentation effect is achieved when α is set to 16.

[0016] The specific process of S2.2 is as follows: the image is first smoothed by a Gaussian filter to reduce noise; the two-dimensional convolution kernel of the Gaussian filter is: Represents the weight value at position (x, y) in a two-dimensional Gaussian filter. It is a distance-based weighting function used for image smoothing. In this formula, x and y represent the horizontal and vertical offsets of the pixel in the current convolution kernel relative to the center point, used to construct a discrete convolution kernel; σ is the standard deviation, which controls the degree of blurring of the filter. It is a normalization factor that ensures that the sum of the weights of the entire convolution kernel is 1 to maintain the overall brightness of the image unchanged; is the exponential part of the Gaussian function The image I(x,y) is convolved with the Gaussian kernel G(x,y) to obtain a smoothed image I': I′(x,y)=I(x,y)*G(x,y) (2) The formula I′(x, y) = I(x, y) * G(x, y) represents a two-dimensional convolution operation between an image I(x, y) and a Gaussian kernel G(x, y), yielding a smoothed image I(x, y). Here, I(x, y) is the pixel value of the input image at coordinate (x, y), and G(x, y) is the convolution kernel defined by a Gaussian function, with a weight at each location used to perform a weighted average on the image. The convolution symbol "*" indicates that the Gaussian kernel is aligned with the image at a certain center position and the sum of the products of the pixels in that region of the image and the corresponding weights in the kernel are calculated.

[0017] Calculate the gradient strength and direction of the image using the Sobel operator; horizontal gradient G x and vertical gradient G y , obtained by convolution: G x =I′*S x , G y =I′*S y (3) Among them, S x and S y They are the Sobel operators in the horizontal and vertical directions, respectively, and are initialized as: The gradient strength G and direction θ are: formula Represents the gradient intensity G and gradient direction θ of the image at a certain point. x Indicates the gradient of the image in the horizontal direction (x-axis), reflecting the speed of the grayscale value change of the image in the horizontal direction, usually calculated by the Sobel operator or other differential filters; G y It represents the gradient of the image in the vertical direction (y-axis), reflecting the degree of grayscale change in the vertical direction; the gradient intensity G is the Euclidean norm of the gradients in these two directions, indicating the significance of the edge at that point. The larger the value, the more drastic the grayscale change at that point, and the more likely it is an edge; the gradient direction θ is the direction angle of the gradient at that point, indicating the direction of the edge normal, that is, the direction in which the grayscale changes fastest. It is the direction angle perpendicular to the edge calculated by the inverse tangent function and is generally expressed in degrees or radians.

[0018] Finally, the gradient intensity values ​​of the image are divided into three categories by non-maximum suppression and double thresholding: strong edge, weak edge, and non-edge. Finally, edge connection is performed. After obtaining the preliminary binary boundary map B1, it is further enhanced using the boundary enhancement module. Specifically, the boundary enhancement module uses four sequentially connected convolutional layers to extract features, and each convolutional layer is followed by batch normalization (BatchNormalization) and ReLU activation function. The convolution kernel size of the first three convolutional layers is 3x3, and the convolution kernel size of the last convolutional layer is 1x1 to compress the number of channels and extract local details. It is worth noting that no downsampling or pooling operations are used in the entire process to maximize the retention of spatial details of the boundary. The final output boundary feature map can be recorded as: B f =ConvBNReLU 1×1 (ConvBNReLU 3×3 (...(ConvBNReLU 3×3 (B1))) (6) Among them ConvBNReLU k×k Represents a composite operation of applying k×k convolution, batch normalization, and ReLU activation in sequence; B1 is the initial binary boundary map obtained after detection using the Canny operator, which is used to represent the contour position of the object in the image; ConvBNReLU 3×3 Represents a composite operation module consisting of a 3×3 convolutional layer, batch normalization, and ReLU activation function. This module is used to extract local boundary features and enhance edge continuity and clarity. This composite module is stacked three times in sequence, allowing the network to gradually capture more complex boundary structure information. The following ConvBNReLU 1×1It is a composite module that uses a 1×1 convolution kernel to compress the number of channels, integrate local features, and further highlight edge areas. It also includes BN and ReLU operations to ensure numerical stability and nonlinear expression capabilities.

[0019] The specific process of S2.4 is as follows: First, although the semantic features extracted by the Transformer Block have strong discriminative ability, the resolution is low. To this end, this paper first expands the spatial resolution of the encoder intermediate feature Ff; Ff is input to a 1x1 convolution layer and upsampled to generate features with higher resolution. Then, through three 1x1 convolutional layers, With boundary feature B f Together they are converted into new feature maps Q, K, and V; then, the attention mask of the intermediate features is calculated by matrix multiplication and Softmax function normalization, which is defined as follows: Among them, S(i, j) represents the influence of the i-th pixel on the j-th pixel in the semantic feature map. is the calculated affinity matrix, and The intermediate features The query feature map and key feature map generated, C′ is the number of feature channels, the size is 512, H and W are the height and width of the feature map respectively; After the affinity matrix S is generated, it is further processed by a 1×1 convolutional layer. The Softmax operation is applied to the boundary feature map V to generate a weighted matrix to balance the degree of fusion between boundary information and context information. Finally, the semantic features and the context attention map are summed element by element to obtain the output feature map F. o ; The specific aggregation operation is expressed as: Among them, F oj and F fj Represent the output feature map and input semantic feature map respectively, is a boundary feature map with a single channel, p s is the projection matrix, used to optimize S j Through the above process, boundary features and semantic features are deeply adaptively fused to improve the performance of the model.

[0020] In S2.5, the specific structure of each Transformer Block module is as follows Figure 3As shown in the figure, it mainly includes a multi-head self-attention layer, a boundary-semantic adaptive fusion module and an MLP layer. The semantic features in the encoder backbone are processed by layer normalization (Layer Norm) and then input into the multi-head self-attention layer; the multi-head self-attention layer receives the query (Q), key (K) and value (V) matrices and calculates the attention weight. Its mathematical expression is: Among them, d head Represents the dimension of each attention head; effectively captures the global dependencies of the input sequence through adaptive weight distribution; The semantic features processed by Transformer are passed to the feature fusion module together with the boundary features. In the boundary-semantic adaptive fusion module, the boundary features and intermediate features are weighted and jointly processed. Assuming that the output intermediate feature of the self-attention mechanism is f mid , first f mid Going into the first adapter and connecting it with the initial feature residual, it can be expressed as: F o =FFM(x i , f Boundary ) (10) Among them, F o represents the fused output feature map, and FFM(·) is the joint weighted fusion of intermediate features and boundary features by the feature fusion module. Through this fusion method, the model can efficiently combine boundary information and context information, thereby significantly improving the perception of local boundary areas and enhancing segmentation accuracy and robustness in complex scenarios.

[0021] In S3, the multiple models or parameter configurations obtained in the training and validation stages are first screened through a unified evaluation on the test set. The data in the test set has not participated in model training or parameter adjustment. The diversity and complexity of the data in the test set reflect the generalization performance of the model in actual scenarios.

[0022] The beneficial effects of the present invention are as follows: by constructing a dual-branch structure, introducing a boundary modeling path based on edge detection, and combining a multi-scale feature refinement mechanism with a semantic boundary fusion strategy, the present invention effectively enhances the model's structural perception of complex boundary regions without changing the backbone network structure. This not only improves segmentation accuracy, particularly achieving greater continuity and integrity in boundary regions, but also possesses good adaptability and engineering feasibility, offering significant advantages in practical applications such as intelligent interpretation of remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a diagram of the overall network architecture of the present invention; Figure 2This is a diagram of the boundary-semantic adaptive fusion module of the present invention; Figure 3 This is a diagram of the Transformer Block structure based on boundary enhancement of the present invention; Figure 4 This is a structural diagram of the enhanced mask decoder of the present invention; Figure 5 This is a visualization result diagram of the comparative experiment on the Vaihingen dataset of the present invention; Figure 6 This is a visualization result of the comparative experiment on the Postdam dataset of the present invention; Figure 7 This is a visualization result diagram of the comparative experiment of the present invention on the Loveda dataset. DETAILED DESCRIPTION

[0024] The present invention will be further described below with reference to the accompanying drawings and examples.

[0025] A method for constructing a large remote sensing model integrating edge semantic knowledge includes the following steps: S1, data preprocessing and enhancement strategy for backup; The specific process of S1 is as follows: S1.1. Establish a standardized data input process for remote sensing image preprocessing and enhancement. During the data preparation phase, representative public remote sensing semantic segmentation datasets, including LoveDA, GID, and ISPRS Potsdam, were selected to cover a variety of typical remote sensing scenes, ensuring diversity and generalization in model training. To ensure uniformity in input data, all original remote sensing images were cropped or scaled to 512×512 pixels to maintain consistent image resolution. They were then normalized using fixed statistical parameters for the three image channels, with the mean values ​​for the red, green, and blue channels being [0.485, 0.456, 0.406] and the standard deviations being [0.229, 0.224, 0.225], respectively. This normalization process eliminates the effects of brightness, color, and sensor variability in remote sensing images, improving the model's ability to perceive the image's essential characteristics.

[0026] S1.2. To construct a representative set of training samples, a random stratified sampling strategy was used to divide the original data into a training set, a validation set, and a test set, with a division ratio of approximately 7:1.5:1.5. This ensured that each subset was balanced in terms of the number of categories, spatial distribution, and sample characteristics, and avoided the impact of data bias on model training. In addition, to enhance the model's robustness to factors such as scale changes, morphological changes, and imaging perturbations in remote sensing images, a data augmentation strategy was implemented on the samples in the training set. The enhancement methods included, but were not limited to, a combination of rotation, scaling, mirroring, color perturbation, and random cropping. Geometric transformation operations helped improve the model's adaptability to changes in the target space's morphology, and color perturbation operations enhanced the model's performance under different lighting and seasonal conditions. Overall, the distribution range of training samples in the feature space was significantly expanded.

[0027] S1.3. Considering that the GGB-SAMNet structure proposed in the present invention has a dual-branch design of semantics and boundary, the input image is processed differently according to the branch feature extraction requirements; specifically, the semantic branch needs to capture global semantic information, and its input image size is set to 256×256 pixels, while the boundary branch focuses on the recovery of detailed structures and edge information extraction, retaining the higher resolution original input of 512×512 pixels to enhance spatial boundary details; before the image is input into the network, normalization and standardization preprocessing are performed on the images of the two branches respectively to ensure the consistency and comparability of the image input features in the distribution dimension.

[0028] S1.4. To provide reasonable prior information support for the boundary modeling branch, a boundary label map generation mechanism is introduced during the training sample preparation phase. The Canny edge detection algorithm is used to extract edges from the original RGB remote sensing imagery, generating a binary boundary map as the initial boundary feature map B1. This binary boundary map preserves boundary information in areas of significant variation between features and provides a stable supervisory signal for the boundary branch in the early stages of training, preventing overfitting to noisy areas. Furthermore, the binary boundary map serves as input features for the subsequent refinement module of the boundary branch, supporting the refined modeling process.

[0029] S2. Feature extraction and fusion network design based on boundary enhancement: GGB-SAMNet structure and boundary guidance mechanism design; After completing the preprocessing, standardization and multi-source enhancement operations of remote sensing images, the structure construction stage of the semantic modeling network is entered. By constructing a fusion network architecture with global perception ability and boundary detail modeling ability, the expression accuracy and structural integrity of remote sensing image semantic segmentation in complex land feature boundary scenes are improved; the structure is based on the SegmentAnything Model (SAM) as the backbone skeleton, introduces the boundary enhancement path, and embeds the boundary perception module in the semantic branch backbone to achieve the coordinated optimization of semantic modeling and boundary supervision; the overall network is named GGB-SAMNet, and its structure is as follows Figure 1 shown.

[0030] The specific process of S2 is as follows: S2.1. We propose a globally guided boundary enhancement strategy to compensate for the broken or blurred boundary regions caused by the SAM encoder's serialized segmentation of the image. Starting from the perspective of accurate extraction and fusion of boundary information, we propose a globally guided boundary enhancement strategy. This strategy introduces independent boundary branches to obtain refined boundary representations and deeply integrates them with the backbone network features at various stages to achieve global enhancement of boundary information.

[0031] S2.2. Using a boundary generation module; In the boundary generation module, the Canny operator is used to generate boundary features B1 for the input image. Canny is a gradient operator in image processing that identifies the boundaries between different objects in the image by calculating the grayscale changes along a specific direction. The specific process of S2.2 is as follows: the image is first smoothed by a Gaussian filter to reduce noise; the two-dimensional convolution kernel of the Gaussian filter is: Represents the weight value at position (x, y) in a two-dimensional Gaussian filter. It is a distance-based weighting function used for image smoothing. In this formula, x and y represent the horizontal and vertical offsets of the pixel in the current convolution kernel relative to the center point, used to construct the discrete convolution kernel; v is the standard deviation, which controls the degree of blurring of the filter. It is a normalization factor that ensures that the sum of the weights of the entire convolution kernel is 1 to maintain the overall brightness of the image unchanged; is the exponential part of the Gaussian function The image I(x,y) is convolved with the Gaussian kernel G(x,y) to obtain a smoothed image I': I′(x,y)=I(x,y)*G(x,y) (2) The formula I′(x, y) = I(x, y) * G(x, y) represents a two-dimensional convolution operation between an image I(x, y) and a Gaussian kernel G(x, y), yielding a smoothed image I(x, y). Here, I(x, y) is the pixel value of the input image at coordinate (x, y), and G(x, y) is the convolution kernel defined by a Gaussian function, with a weight at each location used to perform a weighted average on the image. The convolution symbol "*" indicates that the Gaussian kernel is aligned with the image at a certain center position and the sum of the products of the pixels in that region of the image and the corresponding weights in the kernel are calculated.

[0032] Calculate the gradient strength and direction of the image using the Sobel operator; horizontal gradient G x and vertical gradient G y , obtained by convolution: G x =I'*S x , G y =I'*S y (3) Among them, S x and S y They are the Sobel operators in the horizontal and vertical directions, respectively, and are initialized as: The gradient strength G and direction θ are: formula Represents the gradient intensity G and gradient direction θ of the image at a certain point. x Indicates the gradient of the image in the horizontal direction (x-axis), reflecting the speed of the grayscale value change of the image in the horizontal direction, usually calculated by the Sobel operator or other differential filters; G y It represents the gradient of the image in the vertical direction (y-axis), reflecting the degree of grayscale change in the vertical direction; the gradient intensity G is the Euclidean norm of the gradients in these two directions, indicating the significance of the edge at that point. The larger the value, the more drastic the grayscale change at that point, and the more likely it is an edge; the gradient direction θ is the direction angle of the gradient at that point, indicating the direction of the edge normal, that is, the direction in which the grayscale changes fastest. It is the direction angle perpendicular to the edge calculated by the inverse tangent function and is generally expressed in degrees or radians.

[0033] Finally, the gradient intensity values ​​of the image are divided into three categories through non-maximum suppression and double thresholding: strong edge, weak edge, and non-edge. Finally, edge connection is performed. After obtaining the preliminary binary boundary map B1, it is further enhanced using the boundary enhancement module. Specifically, the boundary enhancement module uses four sequentially connected convolutional layers to extract features, and each convolutional layer is followed by batch normalization and ReLU activation function. The convolution kernel size of the first three convolutional layers is 3x3, and the convolution kernel size of the last convolutional layer is 1x1 to compress the number of channels and extract local details. It is worth noting that no downsampling or pooling operations are used in the entire process to maximize the retention of spatial details of the boundary. The final output boundary feature map can be recorded as: B f =ConvBNReLU 1×1 (ConvBNReLU 3×3 (...(ConvBNReLU 3×3 (B1))) (6) Among them ConvBNReLU k×k Represents a composite operation of applying k×k convolution, batch normalization, and ReLU activation in sequence; B1 is the initial binary boundary map obtained after detection using the Canny operator, which is used to represent the contour position of the object in the image; ConvBNReLU 3×3 Represents a composite operation module consisting of a 3×3 convolutional layer, batch normalization, and ReLU activation function. This module is used to extract local boundary features and enhance edge continuity and clarity. This composite module is stacked three times in sequence, allowing the network to gradually capture more complex boundary structure information. The following ConvBNReLU 1×1 It is a composite module that uses a 1×1 convolution kernel to compress the number of channels, integrate local features, and further highlight edge areas. It also includes BN and ReLU operations to ensure numerical stability and nonlinear expression capabilities.

[0034] S2.3. Construct a boundary refinement module based on multi-scale feature fusion; In order to further improve the accuracy and fine-grained expression ability of boundaries, this paper proposes a boundary refinement module based on multi-scale feature fusion. The boundary refinement module based on multi-scale feature fusion enhances the true expression of boundaries through multi-scale feature compression and reconstruction; specifically as follows: The boundary refinement module adopts an encoder-decoder architecture and combines jump connections to effectively improve segmentation performance. Figure 1As shown in the boundary refinement module in the figure, the encoder and decoder of the boundary refinement module based on multi-scale feature fusion are composed of three convolution blocks, each of which contains two convolution operations, a batch normalization (BatchNorm) layer and a ReLU activation function; in the encoder part, the network gradually increases the number of convolution channels and combines the maximum pooling layer to reduce the spatial resolution, thereby extracting and compressing the local information of the boundary features, providing a clearer boundary representation for the subsequent refinement process; the decoder part restores the spatial resolution through upsampling operations, and accurately restores the boundary details in the image through the fusion of jump connections and low-level features; the boundary refinement module can effectively fuse features of different scales, significantly improving the accuracy and completeness of boundary recovery, which not only enhances the ability to recover boundary details, but also effectively improves the accuracy of the segmentation results.

[0035] S2.4, build boundary-semantic adaptive fusion module, such as Figure 2 As shown in Figure 1, traditional feature fusion modules typically use the same weights when fusing semantic features and boundary features from different stages. This approach fails to fully consider the differences between features from different stages, leading to unbalanced information fusion. This is particularly true when processing remote sensing images with complex boundaries, where the representation of boundary information cannot be effectively enhanced. To fully utilize the boundary features extracted from the boundary refinement module and further improve model performance, this paper proposes a boundary-semantic adaptive fusion module. The boundary-semantic adaptive fusion module effectively integrates explicit boundary features from the boundary branch using an attention mechanism, aligns them with the intermediate semantic features extracted from the main branch, and dynamically adjusts their weights. This significantly enhances the global role of boundary information, further improving model performance.

[0036] The specific process of S2.4 is as follows: First, although the semantic features extracted from the Transformer Block have strong discriminative ability, the resolution is low. Therefore, this paper first performs the encoder intermediate feature F f The spatial resolution of F f It is input into a 1x1 convolutional layer and upsampled to generate features with higher resolution. Then, through three 1x1 convolutional layers, With boundary feature B f Together they are converted into new feature maps Q, K, and V; then, the attention mask of the intermediate features is calculated by matrix multiplication and Softmax function normalization, which is defined as follows: Among them, S(i, j) represents the influence of the i-th pixel on the j-th pixel in the semantic feature map. is the calculated affinity matrix, and The intermediate features The query feature map and key feature map generated, C′ is the number of feature channels, the size is 512, H and W are the height and width of the feature map respectively; After the affinity matrix S is generated, it is further processed by a 1×1 convolutional layer. The Softmax operation is applied to the boundary feature map V to generate a weighted matrix to balance the degree of fusion between boundary information and context information. Finally, the semantic features and the context attention map are summed element by element to obtain the output feature map F. o ; The specific aggregation operation is expressed as: Among them, F oj and F fj Represent the output feature map and input semantic feature map respectively, is a boundary feature map with a single channel, p s is the projection matrix, used to optimize S j Through the above process, boundary features and semantic features are deeply adaptively fused to improve the performance of the model.

[0037] S2.5. Construct a Transformer module based on boundary enhancement. In order to improve the local boundary perception ability of the SAM encoder in remote sensing image analysis, this paper embeds a boundary-semantic adaptive fusion module (BSAF) in each Transformer Block, thereby enhancing the segmentation performance of the model in complex remote sensing scenes. Specifically, while keeping the basic image encoder architecture of the original SAM model unchanged, all original network parameters are frozen. In each Transformer module, fine boundary features are introduced into the network, and multiple simple and efficient adapters are embedded in parallel to enhance the adaptability of the model to remote sensing images.

[0038] In S2.5, the specific structure of each Transformer Block module is as follows Figure 3 As shown in the figure, it mainly includes a multi-head self-attention layer, a boundary-semantic adaptive fusion module and an MLP layer. The semantic features in the encoder backbone are processed by layer normalization (Layer Norm) and then input into the multi-head self-attention layer; the multi-head self-attention layer receives the query (Q), key (K) and value (V) matrices and calculates the attention weight. Its mathematical expression is: Where dhead represents the dimension of each attention head; the global dependency of the input sequence is effectively captured through adaptive weight distribution; The semantic features processed by Transformer are passed to the feature fusion module together with the boundary features. In the boundary-semantic adaptive fusion module, the boundary features and intermediate features are weighted and jointly processed. Assuming that the output intermediate feature of the self-attention mechanism is f mid , first f mid Going into the first adapter and connecting it with the initial feature residual, it can be expressed as: F o =FFM(x i , f Boundary ) (10) Among them, F o represents the fused output feature map, and FFM(·) is the joint weighted fusion of intermediate features and boundary features by the feature fusion module. Through this fusion method, the model can efficiently combine boundary information and context information, thereby significantly improving the perception of local boundary areas and enhancing segmentation accuracy and robustness in complex scenarios.

[0039] S2.6. Construct an enhanced mask decoder. In order to improve the accurate modeling of boundaries and semantics in the semantic segmentation task of remote sensing images, this paper proposes an enhanced mask decoder. The enhanced mask decoder fully utilizes the important role of boundary information in semantic modeling by fusing the semantic features with accurate boundary information extracted by the network encoder and the hint embedding information, significantly improving the segmentation accuracy and detail recovery ability. Specifically: The output features with good boundary modeling generated by the encoder and the hint embedding features are jointly input to the prediction head for processing. The overall architecture is as follows: Figure 4 As shown in the figure; first, two bidirectional Transformer modules are used to efficiently model all embedded features to fully explore the rich information in the embedded features and significantly enhance the expressiveness of the features. The self-attention mechanism is used to model the hint embedding features to effectively capture their contextual relevance; secondly, the bidirectional cross-attention mechanism establishes an efficient interactive relationship between the semantic features with precise boundary information and the hint embedding, effectively strengthening the long-range dependency between boundary details and semantic information, thereby significantly improving the boundary clarity and the accuracy of semantic understanding; then, the class-level convolution module efficiently upsamples the fused features to further refine the boundary information and generate fine semantic masks for multiple category channels to ensure the accurate expression and coherence of the boundary area in the high-resolution mask; finally, the prediction head efficiently combines the category information mask with the tokens generated by the multi-layer perceptron (MLP) through the dot product operation, and outputs a high-resolution multi-category semantic segmentation mask, which achieves accurate classification of each pixel while maintaining the detailed structure and coherence of complex boundaries.

[0040] S2.7. Construct a loss function. For data with imbalanced categories, use the weighted cross entropy loss function, as follows: In the formula, N represents the total number of categories; i is the index of the category, from 1 to N; w i is the weight value of the i-th category, which is used to balance the importance between categories and is often used to deal with category imbalance problems; i is the true label of the i-th category; w i is the original prediction value of the model for the i-th type of output (that is, the output without activation function); σ(logit i ) represents the logit i Apply the Sigmoid function to calculate the probability that the prediction belongs to this category. The Sigmoid function is defined as In addition to the main semantic segmentation loss function, this paper also defines boundary loss to supervise the generation of boundary maps. The boundary pixels of objects are significantly less than the non-boundary pixels, so boundary prediction is a problem of imbalanced class distribution. Binary cross entropy loss is used for joint optimization of boundary learning. Where, and p i Represents the true boundary and the predicted boundary; The total loss is calculated as follows: L total =L seg +α×L bound (14) Among them, L total is the total loss, L seg and L bound Represent segmentation loss and boundary loss respectively; L bound The hyperparameter α before represents its proportion in the total loss; after experimental analysis, the best segmentation effect is achieved when α is set to 16.

[0041] S3: Train and evaluate the dataset processed in step S1 using the model constructed in step S2, focusing on selecting the optimal model and performing comparative analysis of the performance of each model under different metrics. First, multiple models or parameter configurations obtained during the training and validation phases are screened through a unified evaluation on a test set. The data in the test set has not been used in model training or parameter adjustment, and its diversity and complexity reflect the model's generalization performance in real-world scenarios.

[0042] The experiments in this step were run on a Windows 10 operating system and trained and validated using the PyTorch framework on a single Nvidia RTX 3090 GPU. All experiments used the AdamW optimizer and a cosine annealing learning rate strategy, with an initial learning rate of 1e-4. During training, to ensure stability and computational efficiency, a batch size of 1 was used, and 200 training epochs were used.

[0043] During model training, three data augmentation techniques—random vertical flips, random horizontal flips, and random brightness—were used to improve the model's adaptability to different scenarios. All compared methods were trained and validated on the same data benchmark to ensure fairness. To verify the effectiveness of our method, we systematically validated it on three benchmark datasets: Vaihingen, Potsdam, and Loveda, provided by the International Society for Photogrammetry and Remote Sensing (ISPRS).

[0044] GGB-SAMNet further improves the model's ability to extract details by incorporating a globally guided edge enhancement strategy. This strategy demonstrates particularly strong performance in the segmentation of small objects, such as vehicles. In the small vehicle category of the Vaihingen dataset, the F1 score reached 84.62%, significantly outperforming the second-ranked CWSAM by 0.65%. This result demonstrates that edge enhancement strategies can effectively improve the segmentation accuracy of small objects in complex scenarios, particularly in accurately locating edge details between objects. Table 1 Accuracy of different models on the Vaihingen dataset

[0045] In order to intuitively demonstrate the segmentation effects of different algorithms, this paper selects some remote sensing images with occlusion and shadows for visualization. Figure 5The figure shows some segmentation results from the Vaihingen dataset. As can be seen in the red-boxed areas of the first and third groups, traditional CNN methods, due to the limitations of the convolution kernels, are unable to effectively capture long-range contextual information, resulting in incomplete segmentation of low vegetation and buildings. Transformer-based networks, on the other hand, effectively extract global contextual information through their self-attention mechanism, effectively handling large objects and achieving relatively complete extraction results. However, when dealing with objects with complex edges, some edge localization issues arise. CWSAM, one of the most advanced models currently, achieves relatively accurate extraction results compared to the aforementioned models. However, due to its architecture still relying on the limitations of the Transformer, edge localization accuracy still needs improvement. In contrast, GGB-SAMNet not only inherits the Transformer's powerful global context extraction capabilities, but also further improves the localization accuracy of object edges through an edge enhancement strategy. In the red-boxed area of ​​the second group, GGB-SAMNet accurately captures the subtle boundary between houses and vegetation.

[0046] Overall, we can see that for the large, irregularly edged buildings in the dataset, GGB-SAMNet can more completely segment them, with clearer edge contours between objects. In scenes with small vehicles, GGB-SAMNet can better locate the edges between densely packed vehicles, demonstrating its ability to accurately segment objects of varying scales.

[0047] (2) Analysis of experimental results of Postdam dataset Table 2 Accuracy of different models on the Postdam dataset Experimental results on the ISPRS Potsdam dataset are shown in Table 2. These results further demonstrate the superior performance of GGB-SAMNet, particularly in terms of mF1, mIoU, and OA metrics, where GGB-SAMNet surpasses all other methods. Compared to the best-performing CNN-based method, PspNet, GGB-SAMNet achieves a 2.77% improvement in average F1, a 4.52% improvement in mIoU, and a 3% improvement in OA. Compared to the best-performing Transformer-based method, UNetFormer, GGB-SAMNet achieves a 1.21% improvement in average F1, a 2.63% improvement in mIoU, and a 2.37% improvement in OA.

[0048] In order to more intuitively compare the effects of different algorithms in remote sensing image semantic segmentation, this paper shows some visual segmentation results of the test set, such as Figure 6 As shown. This paper selects some challenging samples from the Postdam dataset for analysis. These sets of data contain a variety of complex texture features, such as urban buildings, forests, roads, and impervious surfaces. The background is usually complex and the boundaries between objects are fuzzy, which poses a great challenge to the segmentation task. For example, the texture of the impervious surface in the red frame area in the first and third groups of images is similar to the surrounding low vegetation. The buildings in the second group of images are more difficult to accurately segment due to their shadows and occlusion between different categories.

[0049] Due to the limitations of convolution operations, traditional CNN methods struggle to extract sufficient global contextual information, resulting in incomplete segmentation of object edges and details. For example, in the first set of images, traditional methods blur the building boundaries due to interference from surrounding vegetation. Although Transformer-based networks can effectively capture long-range contextual information, these methods still suffer from inaccurate edge localization when dealing with objects with complex shapes and complex backgrounds. In contrast, GGB-SAMNet demonstrates higher accuracy in this set of images, especially at the junction of buildings and vegetation, accurately locating edges and effectively avoiding missegmentation between objects. In the second and third sets of images, blurred boundaries with surrounding ground and vegetation are effectively avoided. Through its enhanced edge strategy, GGB-SAMNet significantly improves the overall accuracy of buildings, validating the effectiveness of the global guidance and edge enhancement strategy proposed in this paper, especially in complex scenes.

[0050] (3) Analysis of Loveda dataset experimental results The experimental results on the Loveda dataset are shown in Table 3. Our method achieves the best performance in terms of mIoU and achieves excellent segmentation performance on roads, water bodies, and forests. Compared with the second-best method CWSAM, the average IoU is improved by 1.27%, indicating the effectiveness of the global boundary enhancement mechanism in optimizing target contours and reducing mis-segmentation and missed segmentation. Table 3 Accuracy of different models on the Loveda dataset

[0051] In order to more intuitively compare the performance of different algorithms in remote sensing image semantic segmentation tasks, Figure 7 Visual segmentation results of several test set samples are shown.

[0052] Because the spectral characteristics of roads and surrounding objects (such as farmland and bare land) are highly similar, existing methods often experience blurring or breakage at road boundaries, resulting in incoherent segmentation results. In the road area shown in the red box in the first row, GGB-SAMNet effectively extracts the contour information of the road through a global boundary enhancement strategy, making the segmentation results more accurate and coherent. In the building cluster area in the second row, GGB-SAMNet relies on the boundary refinement module to achieve more precise boundary constraints, making the contours of the extracted results clearer and more complete than other methods, and improving the segmentation accuracy of the building boundaries. In the forest area in the third row, GGB-SAMNet can effectively distinguish between forest and background areas, improving the robustness and generalization ability of the model in complex terrain environments. Overall, GGB-SAMNet shows better segmentation results in different typical scenarios, demonstrating its excellent performance in the semantic segmentation task of remote sensing imagery.

Claims

1. A method for constructing a large remote sensing model integrating edge semantic knowledge, characterized in that: The following steps are involved: S1, data preprocessing and enhancement strategy for backup; S2. Feature extraction and fusion network design based on boundary enhancement: GGB-SAMNet structure and boundary guidance mechanism design. After completing the preprocessing, standardization and multi-source enhancement operations of remote sensing images, the semantic modeling network structure construction phase is entered. A fusion network architecture with global perception and boundary detail modeling capabilities is constructed to improve the expression accuracy and structural integrity of remote sensing image semantic segmentation in complex land feature boundary scenes. The GGB-SAMNet structure is based on the SegmentAnything Model as the backbone, introduces a boundary enhancement path, and embeds a boundary perception module in the semantic branch backbone to achieve coordinated optimization of semantic modeling and boundary supervision. S3. Train and evaluate the data set processed in step S1 using the model constructed in step S2, focusing on selecting the optimal model and performing comparative analysis on the performance of each model under different indicators.

2. The method for constructing a large remote sensing model integrating edge semantic knowledge according to claim 1, characterized in that: The specific process of S1 is as follows: S1.

1. Establish a standardized data input process for remote sensing image preprocessing and enhancement. To ensure uniformity of input data, all original remote sensing images are cropped or scaled to 512×512 pixels to maintain consistent image resolution. Fixed three-channel statistical parameters are used for standardization to enhance the model's ability to perceive the essential characteristics of the images. S1.

2. To construct a representative set of training samples, a random stratified sampling strategy was used to divide the original data into a training set, a validation set, and a test set, with a division ratio of approximately 7:1.5:1.

5. This ensured that each subset was balanced in terms of the number of categories, spatial distribution, and sample characteristics, and avoided the impact of data bias on model training. In addition, to enhance the model's robustness to factors such as scale changes, morphological changes, and imaging disturbances in remote sensing images, a data augmentation strategy was implemented on the samples in the training set. S1.

3. Input images are processed differently based on the branch feature extraction requirements. Specifically, the semantic branch captures global semantic information, and its input image size is set to 256×256 pixels. The boundary branch focuses on recovering detailed structures and extracting edge information, retaining the higher-resolution original input of 512×512 pixels to enhance spatial boundary details. Before the image is input to the network, the images of the two branches are subjected to normalization and standardization preprocessing to ensure consistency and comparability of the image input features in the distribution dimension. S1.

4. To provide the boundary modeling branch with reasonable prior information support, a boundary label map generation mechanism is introduced in the training sample preparation stage. The Canny edge detection algorithm is used to extract edges from the original RGB remote sensing image to generate a binary boundary map as the initial boundary feature map B1. The binary boundary map retains the boundary line information of areas with significant changes between objects and provides a stable supervision signal for the boundary branch in the early training stage to avoid overfitting of noisy areas in the early training stage. The binary boundary map also serves as the input feature of the subsequent refinement module of the boundary branch to support the refined modeling process.

3. The method for constructing a remote sensing large model integrating edge semantic knowledge according to claim 1, characterized in that: The specific process of S2 is as follows: S2.

1. We propose a globally guided boundary enhancement strategy to compensate for the broken or blurred boundary regions caused by the SAM encoder's serialized segmentation of the image. This globally guided boundary enhancement strategy introduces independent boundary branches to obtain refined boundary representations and deeply integrates them with the backbone network features at various stages to achieve global enhancement of boundary information. S2.

2. Construct a boundary generation module. In the boundary generation module, use the Canny operator to generate boundary features B1 for the input image. Canny is a gradient operator in image processing that identifies the boundaries between different objects in the image by calculating the grayscale change along a specific direction. S2.

3. Construct a boundary refinement module based on multi-scale feature fusion; the boundary refinement module based on multi-scale feature fusion enhances the true expression of the boundary through multi-scale feature compression and reconstruction; specifically as follows: the boundary refinement module adopts an encoder-decoder architecture and combines jump connections to effectively improve the segmentation performance. The encoder and decoder of the boundary refinement module based on multi-scale feature fusion are composed of three convolution blocks, each of which contains two convolution operations, a batch normalization BatchNorm layer and a ReLU activation function; in the encoder part, the network gradually increases the number of convolution channels and combines the maximum pooling layer to reduce the spatial resolution, thereby extracting and compressing the local information of the boundary features, providing a clearer boundary representation for the subsequent refinement process; the decoder part restores the spatial resolution through upsampling operations, and accurately restores the boundary details in the image through jump connections and fusion with low-level features; the boundary refinement module can effectively fuse features of different scales, significantly improving the accuracy and completeness of boundary recovery, which not only enhances the ability to recover boundary details, but also effectively improves the accuracy of the segmentation results; S2.

4. Construct a boundary-semantic adaptive fusion module. This module uses an attention mechanism to effectively integrate explicit boundary features from the boundary branch, aligning and dynamically weighting them with the intermediate semantic features extracted by the main branch. This significantly enhances the global role of boundary information, further improving model performance. S2.

5. Constructing a Transformer module based on boundary enhancement. A boundary-semantic adaptive fusion module is embedded in each Transformer block to enhance the model's segmentation performance in complex remote sensing scenes. Specifically, while maintaining the basic image encoder architecture of the original SAM model, all original network parameters are frozen. Within each Transformer module, fine boundary features are introduced into the network, and multiple adapters are embedded in parallel to enhance the model's adaptability to remote sensing imagery. S2.

6. Construct an enhanced mask decoder. The enhanced mask decoder fully utilizes the important role of boundary information in semantic modeling by fusing the semantic features with precise boundary information extracted by the network encoder with the hint embedding information, significantly improving segmentation accuracy and detail recovery capabilities. Specifically: the output features with good boundary modeling generated by the encoder and the hint embedding features are jointly input to the prediction head for processing. First, all embedded features are efficiently modeled through two bidirectional Transformer modules to fully explore the rich information in the embedded features and significantly enhance the expressive power of the features. The self-attention mechanism is used to model the hint embedding features to effectively capture their contextual relevance. Secondly, the bidirectional cross-attention mechanism establishes an efficient interactive relationship between semantic features with precise boundary information and cue embeddings, effectively strengthening the long-range dependency between boundary details and semantic information, thereby significantly improving boundary clarity and the accuracy of semantic understanding. Subsequently, the class-level convolution module efficiently upsamples the fused features to further refine boundary information and generate fine semantic masks for multiple category channels, ensuring the accurate representation and coherence of boundary regions in the high-resolution mask. Finally, the prediction head efficiently combines the category information mask with the tokens generated by the multi-layer perceptron through a dot product operation, outputting a high-resolution multi-category semantic segmentation mask, achieving accurate classification of each pixel while maintaining the detailed structure and coherence of complex boundaries. S2.

7. Construct a loss function. For data with imbalanced categories, use the weighted cross entropy loss function, as follows: In the formula, N represents the total number of categories; i is the index of the category, from 1 to N; w i is the weight value of the i-th category, which is used to balance the importance between categories and is often used to deal with category imbalance problems; i is the true label of the i-th category; w i is the original prediction value of the model for the i-th type of output (that is, the output without activation function); σ(logit i ) represents the logit i Apply the Sigmoid function to calculate the probability that the prediction belongs to this category. The Sigmoid function is defined as Boundary loss is used to supervise the generation of boundary maps. The boundary pixels of an object are significantly less than the non-boundary pixels, so boundary prediction is a problem of imbalanced class distribution. Binary cross entropy loss is used for joint optimization of boundary learning. Where y i and p i Represents the true boundary and the predicted boundary; The total loss is calculated as follows: L total =L seg +α×L bound (14) Among them, L total is the total loss, L seg and L bound Represent segmentation loss and boundary loss respectively; L bound The hyperparameter α before represents its proportion in the total loss.

4. The method for constructing a large remote sensing model integrating edge semantic knowledge according to claim 3, characterized in that: The specific process of S2.2 is as follows: the image is first smoothed by a Gaussian filter to reduce noise; the two-dimensional convolution kernel of the Gaussian filter is: Represents the weight value at position (x, y) in a two-dimensional Gaussian filter. It is a distance-based weighting function used for image smoothing. In this formula, x and y represent the horizontal and vertical offsets of the pixel in the current convolution kernel relative to the center point, which are used to construct the discrete convolution kernel; σ is the standard deviation, which controls the blurriness of the filter. It is a normalization factor that ensures that the sum of the weights of the entire convolution kernel is 1 to maintain the overall brightness of the image unchanged; is the exponential part of the Gaussian function; The image I(x,y) is convolved with the Gaussian kernel G(x,y) to obtain a smoothed image I': I′(x,y)=I(x,y)*G(x,y) (2) The formula I′(x, y) = I(x, y) * G(x, y) represents a two-dimensional convolution operation between an image I(x, y) and a Gaussian kernel G(x, y) to obtain a smoothed image I(x, y). Here, I(x, y) is the pixel value of the input image at the coordinate (x, y), and G(x, y) is the convolution kernel defined by a Gaussian function. Each position is a weight value used to perform weighted averaging on the image. The convolution symbol "*" indicates that the Gaussian kernel is aligned with the image at a certain center position and the sum of the products of the pixels in that area of ​​the image and the corresponding weights in the kernel are calculated. Calculate the gradient strength and direction of the image using the Sobel operator; horizontal gradient G x and vertical gradient G y , obtained by convolution: G x =I′*S x ,G y =I′*S y (3) Among them, S x and S y They are the Sobel operators in the horizontal and vertical directions, respectively, and are initialized as: The gradient strength G and direction θ are: formula Represents the gradient intensity G and gradient direction θ of the image at a certain point; where G x It represents the gradient of the image on the horizontal x-axis, reflecting the speed of the grayscale value change of the image in the horizontal direction, usually calculated by the Sobel operator or other differential filters; G y It represents the gradient of the image on the vertical y-axis, reflecting the degree of grayscale change in the vertical direction; the gradient intensity G is the Euclidean norm of the gradients in these two directions, indicating the significance of the edge at that point. The larger the value, the more drastic the grayscale change at that point, and the more likely it is an edge; the gradient direction θ is the direction angle of the gradient at that point, indicating the direction of the edge normal, that is, the direction with the fastest grayscale change. It is calculated by the inverse tangent function and is generally expressed in degrees or radians. Finally, the gradient intensity values ​​of the image are divided into three categories through non-maximum suppression and double thresholding: strong edge, weak edge, and non-edge. Finally, edge connection is performed. After obtaining the preliminary binary boundary map B1, it is further enhanced using the boundary enhancement module. Specifically, the boundary enhancement module uses four sequentially connected convolutional layers to extract features, and each convolutional layer is followed by batch normalization and ReLU activation function. The convolution kernel size of the first three convolutional layers is 3x3, and the convolution kernel size of the last convolutional layer is 1x1 to compress the number of channels and extract local details. It is worth noting that no downsampling or pooling operations are used in the entire process to retain the spatial details of the boundary to the greatest extent. The final output boundary feature map can be recorded as: B f =ConvBNReLU 1×1 (ConvBNReLU 3×3 (...(ConvBNReLU 3×3 (B1))) (6) Among them ConvBNReLU k×k Represents a composite operation of applying k×k convolution, batch normalization, and ReLU activation in sequence; B1 is the initial binary boundary map obtained after detection using the Canny operator, which is used to represent the contour position of the object in the image; ConvBNReLU 3×3 Represents a composite operation module consisting of a 3×3 convolutional layer, batch normalization, and ReLU activation function. The composite operation module is used to extract local boundary features and enhance edge continuity and clarity. This composite module is stacked three times in sequence, enabling the network to gradually capture more complex boundary structure information. ConvBNReLU 1×1 It is a composite module that uses a 1×1 convolution kernel to compress the number of channels, integrate local features, and further highlight edge areas. It also includes BN and ReLU operations to ensure numerical stability and nonlinear expression capabilities.

5. The method for constructing a remote sensing large model integrating edge semantic knowledge according to claim 3, characterized in that: The specific process of S2.4 is as follows: First, the encoder intermediate feature F f The spatial resolution of F f It is input into a 1x1 convolutional layer and upsampled to generate features with higher resolution. Then, through three 1x1 convolutional layers, With boundary feature B f Together they are converted into new feature maps Q, K, and V; then, the attention mask of the intermediate features is calculated by matrix multiplication and Softmax function normalization, which is defined as follows: Among them, S(i, j) represents the influence of the i-th pixel on the j-th pixel in the semantic feature map. is the calculated affinity matrix, and The intermediate features The query feature map and key feature map generated, C′ is the number of feature channels, the size is 512, H and W are the height and width of the feature map respectively; After the affinity matrix S is generated, it is further processed by a 1×1 convolutional layer. The Softmax operation is applied to the boundary feature map V to generate a weighted matrix to balance the degree of fusion between boundary information and context information. Finally, the semantic features and the context attention map are summed element by element to obtain the output feature map F. o ; The specific aggregation operation is expressed as: Among them, F oj and F fj Represent the output feature map and input semantic feature map respectively, is a boundary feature map with a single channel, p s is the projection matrix, used to optimize S j Through the above process, boundary features and semantic features are deeply adaptively fused to improve the performance of the model.

6. The method for constructing a remote sensing large model integrating edge semantic knowledge according to claim 3, characterized in that: In S2.5, each Transformer Block module includes a multi-head self-attention layer, a boundary-semantic adaptive fusion module, and an MLP layer. The semantic features in the encoder backbone are input into the multi-head self-attention layer after layer normalization. The multi-head self-attention layer receives the query (Q), key (K), and value (V) matrices and calculates the attention weight, which is mathematically expressed as: Among them, d head Represents the dimension of each attention head; effectively captures the global dependencies of the input sequence through adaptive weight distribution; The semantic features processed by Transformer are passed to the feature fusion module together with the boundary features. In the boundary-semantic adaptive fusion module, the boundary features and intermediate features are weighted and jointly processed. Assuming that the output intermediate feature of the self-attention mechanism is f mid , first f mid Going into the first adapter and connecting it with the initial feature residual, it can be expressed as: F o =FFM(x i ,f Boundary ) (10) Among them, F o represents the fused output feature map, and FFM(·) is the joint weighted fusion of intermediate features and boundary features by the feature fusion module. Through this fusion method, the model combines boundary information and context information, thereby significantly improving the perception of local boundary areas and enhancing segmentation accuracy and robustness in complex scenarios.

7. The method for constructing a remote sensing large model integrating edge semantic knowledge according to claim 1, characterized in that: In S3, the multiple models or parameter configurations obtained in the training and validation stages are first screened through a unified evaluation on the test set. The data in the test set has not participated in model training or parameter adjustment. The diversity and complexity of the data in the test set reflect the generalization performance of the model in actual scenarios.

Citation Information

Cited By

  • Intelligent identification method based on oil field electric power operation safety behavior

    CN121305675A

  • Intelligent identification method based on safety behavior of oilfield electric operation

    CN121305675B

  • Structured image analysis method and system for fine-grained structure boundary and small target segmentation

    CN121482384A

  • Remote sensing image building boundary segmentation method based on U-net network model

    CN121582278A

  • Intelligent salt cavern form identification method based on edge optimization and evolution constraint

    CN122244689A