Remote sensing image target segmentation method and network based on expansion multi-scale fusion

By adopting the expansion multi-scale fusion method in remote sensing image processing, combined with the expansion convolution and attention mechanism, the problem of low segmentation accuracy of remote sensing images in the existing technology is solved, and more efficient multi-scale feature extraction and fusion is achieved, which significantly improves the segmentation accuracy.

CN119992100AActive Publication Date: 2025-05-13耕宇牧星(北京)空间科技有限公司

Patent Information

Application Number
CN202510175210.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The prior art uses low-contrast or small-scale remote sensing images in the prior art to process complex backgrounds, low-contrast or small targets, and it is difficult to effectively process features of different scales.

Method used

A method based on expansion multi-scale fusion is adopted, combined with expansion convolution and attention mechanism, and the receptive field is enhanced through expansion convolution operations, attention mechanism weighted calculation, and the attention matrix simulates expansion convolution to achieve multi-scale feature extraction and fusion.

Benefits of technology

The accuracy of remote sensing image target segmentation is significantly improved, especially in complex scenes and low contrast areas, and the model's perception of complex terrain and targets at different scales is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992100A_ABST
    Figure CN119992100A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image target segmentation method and network based on expansion multi-scale fusion, and belongs to the technical field of remote sensing image processing, and the method comprises the steps: carrying out the preprocessing and feature map extraction of an original remote sensing image; processing the extracted feature map by using expansion convolution, an attention mechanism and an attention matrix to obtain an enhanced feature map; and performing multi-scale feature extraction and fusion on the enhanced feature map, and outputting a target segmentation result with the same resolution as the original remote sensing image through a decoder. Through combination of expansion convolution and multi-scale feature fusion, the method can effectively improve the segmentation precision of the target region in the remote sensing image, especially in a complex scene and a low-contrast region. And by applying an efficient attention mechanism and feature enhancement in the feature extraction process, the performance of the segmentation model is further optimized, so that the segmentation model can be efficiently operated under different computing resources, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image target segmentation method and network based on expansion multi-scale fusion. Background Art

[0002] With the rapid development of remote sensing technology, the application of remote sensing images has been widely used in many fields such as geographic information systems, environmental monitoring, agriculture, urban planning, etc. Remote sensing images usually have complex scenes and high-dimensional features, which not only contain rich geographic information, but also involve target objects of different scales and types. In these images, the boundaries of the targets are often blurred and affected by factors such as noise, low contrast, and complex background, making the image segmentation task more difficult. Traditional image segmentation methods, such as those based on thresholds, edge detection, and region growing, are effective in some cases, but they often perform poorly when dealing with complex backgrounds, low contrast, or small targets, and have limited segmentation capabilities for multi-scale targets.

[0003] In recent years, deep learning methods, especially convolutional neural networks (CNNs), have made significant progress in the field of image segmentation. Deep learning can overcome the limitations of traditional methods to a certain extent by automatically learning high-level features of images, especially in complex image scenes. However, existing deep learning methods still have some problems, especially in remote sensing images. Due to the large differences in target size, complex background, and many low-contrast areas, it is difficult for the model to effectively process features of different scales. In addition, the current mainstream convolution operations mostly rely on fixed-size convolution kernels, which easily ignore long-distance dependencies with important information in the image, resulting in unsatisfactory segmentation effects in some small targets or low-contrast areas.

[0004] Therefore, how to improve the target segmentation accuracy of remote sensing images is a technical problem that technical personnel in this field need to solve urgently. Summary of the invention

[0005] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a remote sensing image target segmentation method and network based on dilated multi-scale fusion. The method enhances the perception ability of complex terrain and diversified targets by combining dilated convolution with attention mechanism, and can effectively improve the segmentation accuracy of target areas in remote sensing images, especially in complex scenes and low-contrast areas.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, an embodiment of the present invention provides a remote sensing image target segmentation method based on dilation multi-scale fusion, the method comprising the following steps:

[0008] Step 1: Preprocess the original remote sensing image and extract feature maps;

[0009] Step 2: Use dilated convolution, attention mechanism and attention matrix to process the extracted feature map to obtain enhanced feature map;

[0010] Step 3: Perform multi-scale feature extraction and fusion on the enhanced feature map, and output the target segmentation result with the same resolution as the original remote sensing image through the decoder.

[0011] Furthermore, in step 1, the original remote sensing image is preprocessed and feature maps are extracted, and the specific process includes:

[0012] Step 1.1: Preprocess the original remote sensing image, including normalization, image enhancement and denoising, contrast enhancement and detail enhancement;

[0013] Step 1.2: Combine the convolutional neural network and the transformer, and take advantage of the Diffusion model and the pre-trained model to extract features from the pre-processed remote sensing image.

[0014] Step 1.3: The extracted features are fused through the channel dimension, and then the spatial resolution is restored to generate a feature map F with spatial details and global context information suitable for image analysis tasks.

[0015] Furthermore, in step 2, the extracted feature map is processed using dilated convolution, attention mechanism and attention matrix to obtain an enhanced feature map; the specific process includes:

[0016] Step 2.1: The feature map F obtained in step 1 is passed through a normalization layer and then expanded through multiple dilated convolution layers to enhance the fusion of local and global information, and the feature map after the dilated convolution operation is obtained.

[0017] Step 2.2: Perform multi-path feature fusion. This process includes an upper branch and a lower branch, where:

[0018] Upper branch: First, the feature map Perform global average pooling to obtain global context information; then perform nonlinear transformation through the spike neuron layer, and enhance the expressiveness of the feature map through the activation function;

[0019] Lower branch: contains two layers of dilated convolutions, after which activation functions are applied to maintain the validity and nonlinear representation of the output;

[0020] Finally, the outputs of the upper branch and the lower branch are fused through an addition operation to obtain a fused feature map

[0021] Step 2.3: Fusion feature map After being processed by a normalization layer and three convolutional layers, the query, key, and value are obtained respectively;

[0022] Step 2.4: Construct a fixed attention matrix;

[0023] Step 2.5: Perform weighted calculation of the attention matrix and feature map to obtain the adjusted feature map

[0024] Step 2.6: Adjust the feature map And fusion feature map Perform residual connection to obtain the final enhanced feature map Used for subsequent remote sensing image target segmentation tasks.

[0025] Furthermore, in step 3, multi-scale feature extraction and fusion are performed on the enhanced feature map, and the target segmentation result with the same resolution as the original remote sensing image is output through the decoder. The specific process includes:

[0026] Step 3.1: Enhance the feature map obtained in step 2 Perform multi-scale convolution operations and use different convolution kernel sizes to obtain feature maps of multiple scales: and fuse these multi-scale feature maps by weighted averaging or splicing to obtain a feature map containing local features

[0027] Step 3.2: Feature map Perform global average pooling to obtain global features Then the global feature and feature maps containing local features Splice to get the final feature map

[0028] Step 3.3: The final feature map is passed through the decoder Converted into target segmentation results with the same resolution as the original remote sensing image.

[0029] Furthermore, in step 3, the total loss function is constructed in a weighted manner by combining the cross entropy loss and the Dice loss to perform remote sensing image target segmentation training.

[0030] Furthermore, the total loss function constructed is:

[0031]

[0032] in, represents the total loss; represents the cross entropy loss; represents Dice loss; α and β represent the weight coefficients of cross entropy loss and Dice loss, respectively.

[0033] In a second aspect, the present invention further provides a remote sensing image target segmentation network based on dilation multi-scale fusion, which applies the above-mentioned remote sensing image target segmentation method based on dilation multi-scale fusion to realize remote sensing image target segmentation, and the network includes:

[0034] Image preprocessing and feature map extraction module, used for preprocessing and feature map extraction of original remote sensing images;

[0035] The dilated multi-scale fusion module is used to process the extracted feature maps using dilated convolution, attention mechanism and attention matrix to obtain enhanced feature maps;

[0036] The target segmentation module is used to extract and fuse multi-scale features of the enhanced feature map, and output the target segmentation result with the same resolution as the original remote sensing image through the decoder.

[0037] In a third aspect, the present invention also provides an electronic device comprising a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above-mentioned remote sensing image target segmentation method based on dilation multi-scale fusion.

[0038] Compared with the prior art, the present invention has at least the following beneficial effects:

[0039] 1) Improved segmentation accuracy: The present invention introduces dilated convolution and multi-scale feature fusion to enhance the model’s ability to perceive local details and global context when processing complex remote sensing images, especially in the segmentation of low-contrast areas and small targets.

[0040] 2) Enhanced multi-scale feature extraction capability: The present invention combines dilated convolution with attention mechanism to effectively capture the features of remote sensing images at multiple scales, enhance the model’s perception of complex terrain and targets of different scales, thereby improving segmentation accuracy and robustness, especially in complex scenes.

[0041] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0042] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0045] Figure 1 A schematic flow chart of a remote sensing image target segmentation method based on dilation multi-scale fusion provided in an embodiment of the present invention.

[0046] Figure 2 A schematic diagram of the working principle of the expansion multi-scale fusion module provided in an embodiment of the present invention.

[0047] Figure 3 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0049] In the description of the present invention, it should be noted that: in some processes described in the specification and drawings of this application, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be performed in the order in which they appear in this document or may be performed in parallel. In addition, various serial numbers are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0050] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0051] See also Figure 1 As shown, the present invention provides a remote sensing image target segmentation method based on dilation multi-scale fusion, which aims to improve the accuracy of remote sensing image target segmentation in complex scenes. The specific implementation method and working principle of the method of the present invention are as follows:

[0052] Step 1: Preprocess the original remote sensing image and extract the feature map; the specific process is as follows:

[0053] Step 1.1: Preprocessing and image quality enhancement of input images: Remote sensing images are usually multi-channel images (such as RGB, depth or multispectral) with a size of H×W×C, where H is the height of the image, W is the width of the image, and C is the number of channels of the image. In order to ensure efficient processing and improve the performance of subsequent networks, the original remote sensing images are first preprocessed, including:

[0054] ① Normalization: Standardize the pixel values ​​and map them to a fixed range (such as [0, 1] or [-1, 1]) to ensure the numerical stability of the data and make it suitable for subsequent models.

[0055] ② Image enhancement and denoising: The Mamba model (adaptive feature enhancement module) is used to remove noise and enhance the input image. Mamba enhances the local features of the image and optimizes the visibility of long-distance and small targets, thereby improving the model's ability to recognize complex patterns and details in remote sensing images. This module uses an image enhancement strategy based on deep learning to suppress noise in low signal-to-noise ratio areas and retain effective features.

[0056] ③ Contrast enhancement and detail enhancement: Combining CLAHE (adaptive histogram equalization) and the enhancement method based on the Diffusion model, the local contrast of the image is enhanced, especially in the low-contrast area, thereby improving the ability to express details and providing richer input data for subsequent feature extraction.

[0057] Step 1.2: Feature extraction (combining convolution and transformer architecture). Feature extraction combines the convolutional neural network (CNN) and transformer architecture, and incorporates the advantages of the diffusion model and pre-trained model to make feature extraction more efficient and accurate; specifically, it includes:

[0058] ① Preliminary convolution feature extraction: The input image is subjected to a preliminary convolution layer (using a smaller convolution kernel, such as a 3×3 convolution kernel) for low-level feature extraction to capture local information such as edges and textures. These convolution layers gradually reduce the spatial size and increase the number of channels through multiple downsampling (such as convolution with a stride of 2), and provide effective low-level features for subsequent modules.

[0059] ② Diffusion enhanced feature extraction: Based on the traditional convolution layer, a Diffusion model is added. The Diffusion model enhances the local and global information of the feature map by simulating the diffusion process of the image. Specifically, Diffusion strengthens the feature representation by gradually transmitting and updating pixel information. Especially when processing complex images (such as the changing terrain and object structures in remote sensing images), it can effectively capture details and improve the recognition ability of textures and contours.

[0060] ④ Transformer: The features extracted by convolution are further processed using the visual transformer (ViT) architecture. First, the output feature map of the convolution layer is divided into image blocks of fixed size, and each image block is flattened into a one-dimensional vector. Then, the features of each image block are mapped to a fixed dimension through linear transformation, and position encoding is added to preserve the spatial position relationship of the blocks. Next, these features are input into the Transformer encoder for deep learning to capture the global and local feature dependencies of the image, and the self-attention mechanism is used to efficiently model long-range dependencies.

[0061] ⑤ Pre-training and fine-tuning: In order to improve the learning efficiency and robustness of the model, a pre-trained Transformer model based on large-scale remote sensing datasets is introduced. By fine-tuning the pre-trained model, the model can quickly adapt to the needs of specific tasks and effectively improve the quality and speed of feature extraction. The pre-trained model enables the network to learn rich visual representations from a large amount of image data, reducing training time and avoiding overfitting.

[0062] Step 1.3: Feature fusion and high-resolution restoration. In order to maximize the advantages of convolution and transformer, the extracted features are fused through the channel dimension and then the spatial resolution is restored for further processing; specifically:

[0063] ① Fusion of convolution and transformer features: The output feature maps of the convolution layer and the Transformer encoder are concatenated (Concat operation) in the channel dimension to obtain the fused feature map. This can simultaneously retain the local detail information captured by the convolution layer and the global context information learned by the transformer, making the feature map richer and more comprehensive.

[0064] ②Super-resolution restoration: restore the fused feature map to high resolution. Use a convolution-based super-resolution network (such as ESRGAN or deconvolution layer) to upsample the feature map and restore its spatial resolution, ensuring that finer local information and details can be captured when restoring the image. This step not only improves the visual quality of the image, but also increases the reliance of subsequent tasks (such as object detection, classification, etc.) on high-resolution images.

[0065] ③ Multi-scale feature fusion: After the final upsampling operation, the multi-scale feature information is combined to further enhance the image's expressiveness through cross-scale feature fusion. This process introduces multiple layers of feature maps and performs weighted fusion of features of different scales, so that the model can more comprehensively process image areas of different resolutions and enhance the ability to capture image details.

[0066] Step 1.4: Output the final feature map. Through the above steps, the feature map F is finally generated. The feature map F has strong spatial details and global context information, which is suitable for subsequent image analysis tasks.

[0067] Step 2: Construct an expansion multi-scale fusion module, use expansion convolution, attention mechanism and attention matrix to process the extracted feature map F to obtain the enhanced feature map like Figure 2 As shown, specifically including:

[0068] Step 2.1: Dilated convolution operation. The goal of this step is to expand the input feature map through dilated convolution, increase the receptive field, and enable the model to capture a wider range of contextual information. Especially for complex scenes in remote sensing images, dilated convolution can effectively enhance the fusion of local and global information. The feature map F obtained in step 1 is passed through the normalization layer and then through three dilated convolution layers to obtain

[0069] Then multiply Q and K to get Where g represents the grouping factor of the number of channels.

[0070] then Multiply it with V to get the feature map of dilated self-attention, which is then fused with F by skip link to get

[0071]

[0072] in, represents matrix multiplication, Represents matrix addition.

[0073] Step 2.2: Multi-path fusion of feature maps. After obtaining the feature map after the dilated convolution operation After that, multi-path feature fusion is performed. This process includes upper branches and lower branches, which are designed to extract information at different levels and fuse them to improve the final feature representation capability. Among them:

[0074] Upper branch: First A global average pooling (GAP) operation is performed to obtain global context information. Then, a nonlinear transformation is performed through a spiking neuron layer (SNN), and the expression ability of the feature map is further enhanced through an activation function (ReLU);

[0075] Lower branch: The lower branch contains two layers of dilated convolutions, the purpose of which is to further extract detail information in the image. Sigmoid activation function is applied after each layer of dilated convolution to maintain the validity and nonlinear representation of the output;

[0076] Fusion operation: The outputs of the upper branch and the lower branch are fused through the addition operation to obtain the final fused feature map

[0077]

[0078] Among them, SNN represents the spike neuron layer, BN represents the normalization layer, GAP represents the global average pooling layer, DConv represents the dilated convolution layer, and Sigmoid represents the Sigmoid activation function.

[0079] Step 2.3: Feature Map After the normalization layer, and then through three convolutional layers, we get M represents the size of the receptive field. The appropriate M value is selected based on the computational requirements and efficiency requirements of the model.

[0080] Step 2.4: Construct a fixed attention matrix. It is used to limit each query to focus on certain pixel positions in the key or value, specifically pixels at even coordinate positions. The core design idea of ​​this matrix is ​​to selectively mask certain pixels by controlling the relative position, thereby effectively simulating the effect of dilated convolution, while also meeting certain "blind spot" requirements. First, define a smaller binary matrix Used to determine whether attention is allowed. The value of this matrix depends on the relative position (x i -x j ,y i -y j ), and only allows queries to focus on pixel locations with even coordinate differences in the key value. The specific definition is as follows:

[0081]

[0082] Here, x and y are relative coordinate differences, which are assigned 0 when the conditions are met, and 1 otherwise. Next, using this binary matrix Can be further constructed A bigger M 2 ×M 2 Matrix. The specific construction rules are as follows:

[0083]

[0084] Among them, (x i ,y i ) and (x j ,y j ) are the positions of the query and key value respectively. If the relative position difference between the query and key value on some coordinate axis is an even number (i.e., it satisfies the condition), then indicates that the attention values ​​between these positions will not be affected; otherwise, When calculated by softmax, the corresponding attention value will be masked and become 0. The purpose of this design is to make the query only focus on the pixel positions with even coordinate differences in the key, thereby mimicking the behavior of the dilated convolution. The dilated convolution itself captures more contextual information by increasing the spacing between the convolution kernels, but at the same time retains a certain "blind spot" - that is, it does not pay attention to all positions. In this way, the network model can control local relationships while maintaining global perception capabilities, avoiding excessive interference and noise. This design is very useful in certain visual tasks (especially image processing tasks) because it can enhance the model's selective processing of local and global information. This M 2 ×M 2 The matrix is ​​obtained by This process ensures that the model's query only focuses on the pixels at the even-numbered coordinate positions of the key, thereby mimicking the behavior of the dilated convolution and maintaining a certain degree of localization and selective attention. This design helps improve the efficiency of the model and enhance its ability to learn specific spatial structures.

[0085] Step 2.5: Weighted calculation of attention matrix and feature map. After constructing the attention matrix, perform matrix multiplication with the query (Q') and key (K'), and multiply the result with the value (V') to obtain the adjusted feature map. This process simulates the effect of dilated convolution through the design of the attention matrix, so that the model can effectively learn the relationship between local and global features while avoiding excessive interference and noise.

[0086] Step 2.6: Residual connection and output feature map. Finally, the feature map obtained by the attention mechanism With the fused feature map Perform residual connection. The purpose of this step is to preserve the flow of information through skip connections while improving the expressiveness of the final features.

[0087]

[0088] Final feature map It is used in subsequent remote sensing image segmentation tasks. This feature map contains multi-scale information and local context features, which can effectively improve segmentation accuracy, especially when processing complex remote sensing image scenes.

[0089] Through the above steps, the present invention constructs an expanded multi-scale fusion module. The module uses expanded convolution and attention mechanism to combine local details and global context information, and designs an attention matrix to simulate the sparse perception characteristics of expanded convolution. These innovative designs enable the module to more effectively process complex features in remote sensing images, improve segmentation performance and computational efficiency, and are particularly suitable for efficient feature extraction and image structure modeling in remote sensing image segmentation tasks.

[0090] Step 3: From feature map to remote sensing image target segmentation result, as follows:

[0091] Step 3.1: Multi-scale feature extraction and fusion. The feature map obtained from step 2 The information of different scales is extracted from the image to enhance the network's resolution ability, especially when processing the diverse scales of the target image. Multi-scale features are extracted by using convolution kernels of different sizes. Multi-scale convolution operations are performed on the feature map, using different convolution kernel sizes {k1, k2, ..., k n}, and obtain feature maps of multiple scales:

[0092]

[0093] in, Indicates the application of convolution kernel size k i Next, in order to fuse these multi-scale feature maps, weighted averaging or concatenation is used:

[0094]

[0095] Among them, w i is the weight of each scale feature map, which can be learned through training or set based on experience.

[0096] Step 3.2: In order to further capture the contextual information of the target area in the remote sensing image, the global information modeling method is used to improve the model's ability to distinguish between the background and the target. Global average pooling is used to extract the global contextual information of the image. Perform global average pooling to obtain global features:

[0097]

[0098] in, is a global feature vector that contains the global context information of the image. Next, the global context information is fused with the local features to enhance the model’s ability to segment objects of different scales and types. After processing, it is compared with the local features To splice:

[0099]

[0100] Among them, Concat represents a concatenation operation.

[0101] Step 3.3: Predict the target segmentation result. The final feature map is decoded Convert to the target segmentation result with the same resolution as the input image. Use a 1x1 convolution layer to map the feature map to the number of channels corresponding to the number of target categories (for example, road and background segmentation, the number of categories is 2). Perform softmax or sigmoid activation on the output to obtain the probability of each pixel belonging to the target (for example, road) Where (i, j) represents the probability that the (i, j)th pixel in the image belongs to the target.

[0102] Step 3.4: Train the network. The present invention also combines cross entropy loss and Dice loss to optimize the model. Cross entropy loss is used to measure the difference between the predicted target segmentation probability and the true label. This loss function calculates the error between the predicted probability of each pixel category and the true label at the pixel level, and is particularly suitable for class imbalance problems (such as background accounts for a large part and the target area is small). Dice loss is used to measure the overlap between the predicted segmentation result and the true label. Dice loss is calculated based on the Dice coefficient, which is an indicator commonly used to evaluate image segmentation accuracy. By optimizing the Dice loss, the network can pay more attention to the accuracy of the segmented area, especially when the target area is small or complex in shape, which can effectively improve the segmentation quality. Finally, weighted loss is used. Considering the role of different loss functions, cross entropy loss and Dice loss are usually combined in a weighted manner. Cross entropy loss is responsible for adjusting the overall classification ability of the model, while Dice loss ensures the accurate segmentation of the target area. The weighting coefficient can be adjusted through experiments. Usually, we will give a higher weight to Dice loss to ensure the accuracy of the segmented area. The final loss function is a weighted combination of cross entropy loss and Dice loss:

[0103]

[0104] in, represents the total loss; represents the cross entropy loss; represents Dice loss; α and β represent the weight coefficients of cross entropy loss and Dice loss respectively

[0105] During the training process, the segmentation results are gradually optimized by minimizing the total loss. By optimizing the total loss function, the model can learn more accurate target segmentation results, especially in complex backgrounds and target edge areas. Through this step, the network's decoding process and the total loss function work together, allowing the network to accurately segment targets in remote sensing images and improve the accuracy and robustness of target segmentation.

[0106] From the description of the above embodiments, those skilled in the art can know that: the present invention provides a remote sensing image target segmentation method based on dilated multi-scale fusion, and the focus is on combining dilated convolution with the attention mechanism through step 2 to further improve the segmentation accuracy and multi-scale feature extraction ability of the network model. Through the dilated convolution operation, the network can capture long-distance features in remote sensing images within a larger receptive field, thereby improving the perception of complex backgrounds and low-contrast areas. And through the introduction of multi-scale feature fusion, image features of different scales can be effectively and efficiently fused, enhancing the perception of complex terrain and diverse targets, thereby improving segmentation accuracy and robustness.

[0107] In addition, the present invention has also made optimizations in reducing the demand for computing resources. Through efficient fine-tuning of parameters, the computing overhead in the training and deployment process is reduced, and it can run efficiently in different computing environments. Compared with traditional deep learning methods, the present invention not only improves the segmentation accuracy, but also shows higher efficiency in computing resource utilization, and can better adapt to the computing needs in practical applications.

[0108] In summary, the present invention can effectively improve the performance of remote sensing image target segmentation in complex scenes, especially in the segmentation of low-contrast areas and small targets. At the same time, through efficient computing strategies, the present invention can reduce the demand for computing resources while ensuring segmentation accuracy, and has broad application prospects.

[0109] Furthermore, the present invention also provides a remote sensing image target segmentation network based on dilation multi-scale fusion, which applies a remote sensing image target segmentation method based on dilation multi-scale fusion in the above embodiment to achieve remote sensing image target segmentation. The network includes:

[0110] Image preprocessing and feature map extraction module, used for preprocessing and feature map extraction of original remote sensing images;

[0111] The dilated multi-scale fusion module is used to process the extracted feature maps using dilated convolution, attention mechanism and attention matrix to obtain enhanced feature maps;

[0112] The target segmentation module is used to extract and fuse multi-scale features of the enhanced feature map, and output the target segmentation result with the same resolution as the original remote sensing image through the decoder.

[0113] The network provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the system embodiment, reference can be made to the corresponding contents in the aforementioned method embodiment, which will not be repeated here.

[0114] In addition, refer to Figure 3 As shown, an embodiment of the present invention further provides an electronic device, which may include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, and the processor executes the computer program to implement a remote sensing image target segmentation method based on dilation multi-scale fusion in the above method embodiment.

[0115] The processor 10 may be composed of an integrated circuit in some embodiments, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 10 is the control core of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing programs or modules stored in the memory 11, and calling data stored in the memory 11.

[0116] The memory 11 may be, for example, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples of storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (RAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, and any suitable combination thereof.

[0117] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, network systems, electronic devices or computer program products, etc. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0118] It should be noted that the word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several distinct components, and by means of a suitably programmed computer.

[0119] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0120] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image target segmentation method based on dilation multi-scale fusion, characterized in that: The method comprises the following steps: Step 1: Preprocess the original remote sensing image and extract feature maps; Step 2: Use dilated convolution, attention mechanism and attention matrix to process the extracted feature map to obtain enhanced feature map; Step 3: Perform multi-scale feature extraction and fusion on the enhanced feature map, and output the target segmentation result with the same resolution as the original remote sensing image through the decoder.

2. The remote sensing image target segmentation method based on dilation multi-scale fusion according to claim 1, characterized in that: In step 1, the original remote sensing image is preprocessed and feature maps are extracted. The specific process includes: Step 1.1: Preprocess the original remote sensing image, including normalization, image enhancement and denoising, contrast enhancement and detail enhancement; Step 1.2: Combine the convolutional neural network and the transformer, and take advantage of the Diffusion model and the pre-trained model to extract features from the pre-processed remote sensing image. Step 1.3: The extracted features are fused through the channel dimension, and then the spatial resolution is restored to generate a feature map F with spatial details and global context information suitable for image analysis tasks.

3. The remote sensing image target segmentation method based on dilation multi-scale fusion according to claim 2, characterized in that: In the step 2, the extracted feature map is processed using dilated convolution, attention mechanism and attention matrix to obtain an enhanced feature map; The specific process includes: Step 2.1: The feature map F obtained in step 1 is passed through a normalization layer and then expanded through multiple dilated convolution layers to enhance the fusion of local and global information, and the feature map after the dilated convolution operation is obtained. Step 2.2: Perform multi-path feature fusion. This process includes an upper branch and a lower branch, where: Upper branch: First, the feature map Perform global average pooling to obtain global context information; then perform nonlinear transformation through the spike neuron layer, and enhance the expressiveness of the feature map through the activation function; Lower branch: contains two layers of dilated convolutions, after which activation functions are applied to maintain the validity and nonlinear representation of the output; Finally, the outputs of the upper branch and the lower branch are fused through an addition operation to obtain a fused feature map Step 2.3: Fusion feature map After being processed by a normalization layer and three convolutional layers, the query, key, and value are obtained respectively; Step 2.4: Construct a fixed attention matrix; Step 2.5: Perform weighted calculation of the attention matrix and feature map to obtain the adjusted feature map Step 2.6: Adjust the feature map And fusion feature map Perform residual connection to obtain the final enhanced feature map Used for subsequent remote sensing image target segmentation tasks.

4. The remote sensing image target segmentation method based on dilation multi-scale fusion according to claim 3 is characterized in that: In step 3, multi-scale feature extraction and fusion are performed on the enhanced feature map, and the target segmentation result with the same resolution as the original remote sensing image is output through the decoder. The specific process includes: Step 3.1: Enhance the feature map obtained in step 2 Perform multi-scale convolution operations and use different convolution kernel sizes to obtain feature maps of multiple scales: and fuse these multi-scale feature maps by weighted averaging or splicing to obtain a feature map containing local features Step 3.2: Feature map Perform global average pooling to obtain global features Then the global feature and feature maps containing local features Splice to get the final feature map Step 3.3: The final feature map is passed through the decoder Converted into target segmentation results with the same resolution as the original remote sensing image.

5. The remote sensing image target segmentation method based on dilation multi-scale fusion according to claim 1, characterized in that: In step 3, the total loss function is constructed in a weighted manner by combining the cross entropy loss and the Dice loss to perform remote sensing image target segmentation training.

6. The method for remote sensing image target segmentation based on dilation multi-scale fusion according to claim 5, characterized in that: The total loss function constructed is: in, represents the total loss; represents the cross entropy loss; represents Dice loss; α and β represent the weight coefficients of cross entropy loss and Dice loss, respectively.

7. A remote sensing image target segmentation network based on dilation multi-scale fusion, characterized in that: A remote sensing image target segmentation method based on dilation multi-scale fusion as claimed in any one of claims 1 to 6 is applied to achieve remote sensing image target segmentation, wherein the network comprises: Image preprocessing and feature map extraction module, used for preprocessing and feature map extraction of original remote sensing images; The dilated multi-scale fusion module is used to process the extracted feature maps using dilated convolution, attention mechanism and attention matrix to obtain enhanced feature maps; The target segmentation module is used to extract and fuse multi-scale features of the enhanced feature map, and output the target segmentation result with the same resolution as the original remote sensing image through the decoder.

8. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement a remote sensing image target segmentation method based on expansion multi-scale fusion as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Retinal fundus vessel segmentation method based on deep multi-scale attention convolutional neural network

    CN112102283A

  • Remote sensing image target tracking method based on multi-scale feature aggregation enhancement

    CN116777953A

  • Logistics park safety helmet wearing detection and segmentation method and system and medium

    CN118609054A

  • Image segmentation method and image segmentation system

    CN119229107A

Cited By

  • Intelligent high-precision measuring instrument small target segmentation method and system

    CN120580254A

  • Image segmentation method and device, equipment, storage medium and computer program product

    CN120876848A

  • Remote sensing burned area semantic segmentation method based on multi-scale fusion and attention mechanism, terminal and storage medium

    CN121353681A