Remote sensing image semantic segmentation method and system fusing feature redundancy removal operation, and readable storage medium
Through the improved segmentation network, combined with feature redundancy removal module and global-local perception network, high-quality multi-scale features are extracted, and the problem of poor segmentation effect and efficiency in the existing technology is solved, and higher quality segmentation results and stronger generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510051870.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The existing remote sensing image segmentation technology has bottlenecks in complex feature modeling, resulting in poor segmentation effect and efficiency. Especially when dealing with complex and spectral differences, there are problems such as unreliable context and large intra-class variance, and the generalization ability is weak.
The semantic segmentation method of remote sensing image with fusion feature redundancy removal operation is adopted. Through an improved segmentation network, including encoder and decoder, U-Net or ResNet is used as the backbone network, combined with spatial redundancy removal module and channel redundancy removal module, high-quality multi-scale features are extracted, and fusion is carried out through a global-local perceptual network to finally generate high-quality segmentation results.
The performance of semantic segmentation of remote sensing images is improved, the problem of insufficient feature extraction quality is solved, the consumption of computing resources is reduced, the training speed and task performance is improved, and the generalization ability of the method is enhanced, which is suitable for different types of remote sensing images.
Smart Images

Figure CN119964161A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and more specifically to a remote sensing image semantic segmentation method, system and readable storage medium integrating feature redundancy removal operations. Background Art
[0002] At present, remote sensing images show more complex features than traditional images due to their complex and changeable backgrounds, significant intra-class differences and wide scale changes. These characteristics greatly increase the difficulty and challenge of remote sensing image semantic segmentation tasks.
[0003] However, the current remote sensing image segmentation technology encounters bottlenecks in complex feature modeling, and the segmentation effect and efficiency are both poor. How to efficiently capture and process the intricate features in remote sensing images is a difficult problem that needs to be overcome in this field. For example, the remote sensing image semantic segmentation method based on the Convolutional Neural Network (CNN) fusion feature redundancy removal operation mainly relies on context modeling, which specifically includes two methods: spatial context modeling and relational context modeling. Among them, the spatial context modeling method often ignores the differences between different classes. When processing remote sensing images with complex targets and large spectral differences, the context will be unreliable; and the relational context modeling method produces a large intra-class variance, resulting in a large gap between the pixel and the global class representation.
[0004] At the same time, some methods show weak generalization ability in the face of different types of remote sensing images, and it is difficult to accurately parse and segment image content, thus limiting its breadth and effectiveness in practical applications. The encoder-decoder architecture for remote sensing image segmentation significantly optimizes traditional methods. The encoder effectively extracts image features, while the decoder accurately reconstructs based on the encoded features to generate high-quality segmentation results. Although this method has made significant progress compared to previous technologies, it still has limitations in the precision of feature extraction, especially when faced with complex and changeable remote sensing images, its segmentation ability needs to be further enhanced.
[0005] Therefore, how to provide a remote sensing image semantic segmentation method that can solve the above-mentioned problems by removing redundant fusion features is an urgent problem that those skilled in the art need to solve. Summary of the invention
[0006] In view of this, the present invention provides a remote sensing image semantic segmentation method, system and readable storage medium that integrate feature redundancy removal operations, which can extract higher quality image features and further improve the performance of remote sensing image semantic segmentation.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] A remote sensing image semantic segmentation method integrating feature redundancy removal operation comprises the following steps:
[0009] Acquire a remote sensing image to be identified, and preprocess the remote sensing image to be identified;
[0010] Constructing and improving a segmentation network, wherein the improved segmentation network includes an encoder and a decoder connected in sequence, wherein the encoder selects any one of a U-Net network or a ResNet network as a backbone network for extracting and processing features of the remote sensing image to be identified to obtain corresponding multi-scale features; and the decoder is used to detect and fuse the multi-scale features;
[0011] The pre-processed remote sensing image to be identified is input into the improved segmentation network for processing to obtain the corresponding remote sensing image semantic segmentation result.
[0012] Preferably, the improved segmentation network specifically further includes:
[0013] The data smoothing module is used to smooth the result output by the decoder to obtain the final remote sensing image semantic segmentation result.
[0014] Preferably, the encoder includes a residual network and a spatial redundancy removal module which are connected in sequence, and the spatial redundancy removal module implements a spatial redundancy removal operation on the feature extraction result of the remote sensing image to be identified through a depth-separable convolutional layer.
[0015] Preferably, the spatial redundancy removal module is implemented by a spatial feature separation operation and a feature reconstruction operation, and the specific expression is:
[0016]
[0017]
[0018] Where μ and σ are the mean and standard deviation of the input feature X, ε represents the preset parameter, C represents the number of channels, and w i Represents information weight.
[0019] Preferably, the network structure of the encoder further includes: a channel redundancy removal module, which is used to perform a channel redundancy removal operation on the feature extraction result processed by the spatial redundancy removal module.
[0020] Preferably, the channel redundancy removal module is implemented through feature separation, conversion and fusion operations.
[0021] Preferably, the decoder includes a global perception network and a local perception network;
[0022] The global perception network processes the multi-scale features through a fully connected layer and an attention mechanism to obtain corresponding global perception features;
[0023] The local perception network processes the global perception features by introducing an attention mechanism to obtain corresponding local perception features;
[0024] The global perception features and the local perception features are fused to obtain a remote sensing image semantic segmentation result.
[0025] Preferably, the residual network, the spatial redundancy removal module and the channel redundancy removal module are operated multiple times in a cycle.
[0026] The present invention also provides a remote sensing image semantic segmentation system integrating feature redundancy removal operation, comprising:
[0027] An acquisition module, used for acquiring remote sensing images to be identified;
[0028] Model building and improvement module, used to build and improve segmentation networks;
[0029] The segmentation module is used to input the remote sensing image to be identified into the improved segmentation network for processing to obtain the corresponding remote sensing image semantic segmentation result.
[0030] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the remote sensing image semantic segmentation method for removing redundant fusion features as described in any one of the above items is implemented.
[0031] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a remote sensing image semantic segmentation method, system and readable storage medium integrating feature redundancy removal operation, which have the following beneficial effects:
[0032] (1) Designed a simple and practical encoder-decoder structure for semantic segmentation of remote sensing images;
[0033] (2) The feature redundancy removal method and the global-local perception network are integrated, so that the designed encoder-decoder structure can extract high-quality image features, further improving the performance of remote sensing image semantic segmentation;
[0034] (3) It solves the problem of insufficient feature extraction quality in remote sensing images due to complex background and large intra-class variance, which can effectively reduce the consumption of computing resources during model training, so that the designed network model can be effectively improved in terms of training speed. At the same time, the extraction of high-quality features can also effectively improve the task performance of the model.
[0035] (4) The feature redundancy removal method and the global-local perception network are effectively integrated to promote the generalization ability of the remote sensing image semantic segmentation method that integrates the feature redundancy removal operation, including semantic segmentation of satellite remote sensing images, visible light remote sensing images, multispectral remote sensing images, and infrared remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0037] Figure 1 An overall flow chart of a remote sensing image semantic segmentation method that incorporates feature redundancy removal operations provided by the present invention;
[0038] Figure 2 The overall implementation flow chart of the encoder provided by the present invention, wherein (a) represents the spatial feature de-redundancy process, and (b) represents the channel feature de-redundancy process;
[0039] Figure 3 The overall implementation flow chart of the decoder provided by the present invention;
[0040] Figure 4 A structural principle block diagram of a remote sensing image semantic segmentation system provided by the present invention. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] See also Figure 1 As shown, the embodiment of the present invention discloses a remote sensing image semantic segmentation method integrating feature redundancy removal operation, comprising the following steps:
[0043] Acquire a remote sensing image to be identified, and preprocess the remote sensing image to be identified;
[0044] Constructing and improving a segmentation network, wherein the improved segmentation network includes an encoder and a decoder connected in sequence, wherein the encoder selects any one of a U-Net network or a ResNet network as a backbone network for extracting and processing features of the remote sensing image to be identified to obtain corresponding multi-scale features; and the decoder is used to detect and fuse the multi-scale features;
[0045] The pre-processed remote sensing image to be identified is input into the improved segmentation network for processing to obtain the corresponding remote sensing image semantic segmentation result.
[0046] In a specific embodiment, the improved segmentation network further includes:
[0047] The data smoothing module is used to smooth the result output by the decoder to obtain the final remote sensing image semantic segmentation result.
[0048] In a specific embodiment, the encoder includes a residual network and a spatial redundancy removal module which are connected in sequence, and the spatial redundancy removal module implements a spatial redundancy removal operation on the feature extraction result of the remote sensing image to be identified through a depth-separable convolutional layer.
[0049] In a specific embodiment, the spatial redundancy removal module is implemented by a spatial feature separation operation and a feature reconstruction operation, and the specific expression is:
[0050]
[0051] Where μ and σ are the mean and standard deviation of the input feature X, ε represents the preset parameter, C represents the number of channels, and w i Represents information weight.
[0052] In a specific embodiment, the network structure of the encoder further includes: a channel redundancy removal module, which is used to perform a channel redundancy removal operation on the feature extraction result processed by the spatial redundancy removal module.
[0053] In a specific embodiment, the channel redundancy removal module is implemented through feature separation, conversion and fusion operations.
[0054] In a specific embodiment, the decoder includes a global perception network and a local perception network;
[0055] The global perception network processes the multi-scale features through a fully connected layer and an attention mechanism to obtain corresponding global perception features;
[0056] The local perception network processes the global perception features by introducing an attention mechanism to obtain corresponding local perception features;
[0057] The global perception features and the local perception features are fused to obtain a remote sensing image semantic segmentation result.
[0058] In a specific embodiment, the residual network, the spatial redundancy removal module and the channel redundancy removal module are operated multiple times in a cycle.
[0059] Specifically, the backbone network of the encoder can be ResNet50 or ResNet101 as the cornerstone of feature extraction, which can extract deep image features from the original image. Subsequently, in order to improve the effectiveness of feature representation, the embodiment of the present invention introduces spatial redundancy removal operations and channel redundancy removal operations after completing the initial feature extraction, and will again undergo deep refinement of the residual network and feature redundancy removal method. This cycle is repeated four times, deepening the learning depth and richness of the features layer by layer, which can effectively strip off the redundant information in the features, so that subsequent processing can focus on more useful feature information.
[0060] For details, see Figure 2 As shown in Figure 1, (a) shows the process of spatial redundancy removal, which mainly adopts spatial feature separation and feature reconstruction operations. For the feature separation operation, given the input feature map X, first standardize the input feature X by subtracting the mean μ and dividing it by the standard deviation σ, see formula (1):
[0061]
[0062] Where μ and σ are the mean and standard deviation of X, and ε is a small positive constant (e -10 ), α and β are trainable affine transformations. The standardized features will be normalized to obtain the importance of different feature maps. The specific formula is shown in formula (2):
[0063]
[0064] In the formula, C represents the number of channels. Then, W is transformed into α Mapped to the range of (0,1) and filtered by the threshold. Specifically, the threshold is set to 1 / 2, and the information above the threshold is merged into the useful information weight W1, and vice versa. The information below the threshold is merged into the useless information weight W2. After that, the input feature X is multiplied by W1 and W2 respectively to obtain the corresponding information-rich and less information
[0065] In the feature reconstruction operation, the feature with rich information is added to the feature with less information through cross reconstruction operation to fully combine the two different information features. The specific processing process is shown in formula (3):
[0066]
[0067] In the formula, The matrix addition operation is the ∪ matrix concatenation operation.
[0068] (b) shows the process of channel redundancy removal, which mainly includes three operations: feature separation, conversion and fusion. In terms of feature separation, two 1*1 convolution operations are used to separate the features obtained by spatial redundancy removal into sum. In the feature conversion stage, the acquired features are processed by group convolution (Gate-Wise Convolution, GWC) and point-wise convolution (Point-Wise Convolution, PWC) to obtain the corresponding rich features. For the corresponding conversion features obtained by point-wise convolution, the specific process is shown in formula (4):
[0069]
[0070] In the feature fusion stage, for the acquired and Pooling and normalization operations are used to obtain the corresponding importance vector, and then feature fusion is performed through the importance vector to finally obtain the channel redundancy removal feature. The specific process is shown in formula (5):
[0071]
[0072] In a specific embodiment, the specific processing process of the decoder includes:
[0073] The multi-scale features are processed through the global perception network to obtain the corresponding global perception features;
[0074] The global perception features are processed through the local perception network to obtain the corresponding local perception features;
[0075] The global perception features and local perception features are fused to obtain the semantic segmentation results of the remote sensing image.
[0076] Specifically, in the decoding stage, the multi-scale features output by the encoder first pass through the global perception network, which aims to capture the classification information at the overall level of the image and generate global perception features. At the same time, each de-redundant feature generated cyclically during the encoding process will be combined with the global perception feature and input into the local perception network. The local perception network focuses on parsing more detailed and specific local area information from multi-scale and multi-level features to generate local perception features. Finally, the acquired global perception features and local perception features will be effectively fused to generate the corresponding segmentation results.
[0077] For details, see Figure 3 As shown in FIG, for the obtained de-redundant features, the global perception network first obtains the corresponding category probability distribution through a convolution operation, and then obtains the corresponding global perception features by fusing the de-redundant features and the category probability distribution. For the specific process, see formula (6):
[0078]
[0079] Where, X G and Class represent the redundant features and category probability distribution after separation, X C is the obtained local category representation.
[0080] See also Figure 4 As shown, the present invention also provides a remote sensing image semantic segmentation system integrating feature redundancy removal operation, comprising:
[0081] An acquisition module, used for acquiring remote sensing images to be identified;
[0082] Model building and improvement module, used to build and improve segmentation networks;
[0083] The segmentation module is used to input the remote sensing image to be identified into the improved segmentation network for processing to obtain the corresponding remote sensing image semantic segmentation result.
[0084] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the remote sensing image semantic segmentation method for removing redundant fusion features as described in any one of the above items is implemented.
[0085] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0086] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image semantic segmentation method integrating feature redundancy removal operation, characterized in that: The following steps are involved: Acquire a remote sensing image to be identified, and preprocess the remote sensing image to be identified; Constructing and improving a segmentation network, wherein the improved segmentation network includes an encoder and a decoder connected in sequence, wherein the encoder selects any one of a U-Net network or a ResNet network as a backbone network for extracting and processing features of the remote sensing image to be identified to obtain corresponding multi-scale features; and the decoder is used to detect and fuse the multi-scale features; The pre-processed remote sensing image to be identified is input into the improved segmentation network for processing to obtain the corresponding remote sensing image semantic segmentation result.
2. A remote sensing image semantic segmentation method integrating feature redundancy removal operation according to claim 1, characterized in that: The improved segmentation network also includes: The data smoothing module is used to smooth the result output by the decoder to obtain the final remote sensing image semantic segmentation result.
3. The remote sensing image semantic segmentation method according to claim 2, characterized in that: The encoder includes a residual network and a spatial redundancy removal module connected in sequence. The spatial redundancy removal module implements a spatial redundancy removal operation on the feature extraction result of the remote sensing image to be identified through a depth-separable convolutional layer.
4. The remote sensing image semantic segmentation method according to claim 3, characterized in that: The spatial redundancy removal module is implemented by spatial feature separation operation and feature reconstruction operation, and the specific expression is: Where μ and σ are the mean and standard deviation of the input feature X, ε represents the preset parameter, C represents the number of channels, and w i Represents information weight.
5. The remote sensing image semantic segmentation method integrating feature redundancy removal operation according to claim 3, characterized in that: The network structure of the encoder further includes: a channel redundancy removal module, which is used to perform a channel redundancy removal operation on the feature extraction result processed by the spatial redundancy removal module.
6. The remote sensing image semantic segmentation method integrating feature redundancy removal operation according to claim 5, characterized in that: The channel redundancy removal module is implemented through feature separation, conversion and fusion operations.
7. The remote sensing image semantic segmentation method integrating feature redundancy removal operation according to claim 2, characterized in that: The decoder includes a global perception network and a local perception network; The global perception network processes the multi-scale features through a fully connected layer and an attention mechanism to obtain corresponding global perception features; The local perception network processes the global perception features by introducing an attention mechanism to obtain corresponding local perception features; The global perception features and the local perception features are fused to obtain a remote sensing image semantic segmentation result.
8. The remote sensing image semantic segmentation method integrating feature redundancy removal operation according to claim 6, characterized in that: The residual network, the spatial redundancy removal module and the channel redundancy removal module are operated multiple times in a cycle.
9. A segmentation system using a remote sensing image semantic segmentation method with a fusion feature redundancy removal operation as described in any one of claims 1 to 8, characterized in that: include: An acquisition module, used for acquiring remote sensing images to be identified; Model building and improvement module, used to build and improve segmentation networks; The segmentation module is used to input the remote sensing image to be identified into the improved segmentation network for processing to obtain the corresponding remote sensing image semantic segmentation result.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the remote sensing image semantic segmentation method for removing redundant fusion features according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Breast ultrasonic image analysis method and system based on deep convolutional neural network
CN117392125A
Multi-scale building and construction waste extraction method based on high-resolution remote sensing image
CN118736433A
Gearbox fault diagnosis method and device based on digital twinborn and composite adversarial domain discrimination network
CN119104295A
Semantic segmentation-based sea wave height detection method, medium and system
CN119151974A