Remote sensing image fine-grained feature extraction method based on feature conversion and knowledge enhancement
By employing feature transformation and knowledge enhancement methods, a feature knowledge enhancement network was constructed, which solved the problem of difficulty in distinguishing ground features with high inter-class similarity in remote sensing images, and achieved higher-precision fine-grained ground feature extraction.
Patent Information
- Application Number
- CN202411326298.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-09-23
AI Technical Summary
In existing technologies, fine-grained land cover extraction methods for remote sensing images suffer from problems such as complex network structures, difficulty in distinguishing land cover with high similarity between classes, and low extraction accuracy.
A feature transformation module is used to convert pixel information of remote sensing images into feature-level information, and a knowledge enhancement module is used to fuse feature space information at different stages and scales to construct a feature knowledge enhancement network. The segmentation results are optimized by an iterative method controlled by time step.
It improves the network's ability to automatically identify fine-grained land features, enhances the recognition accuracy of land features with high inter-class similarity, and achieves higher extraction accuracy and better land feature differentiation capabilities.
Smart Images

Figure CN119418189B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing image processing technology, and more specifically, relates to a method for fine-grained ground feature extraction from remote sensing images based on feature transformation and knowledge enhancement. Background Technology
[0002] Fine-grained ground feature extraction is a key component of remote sensing image analysis. Its goal is to accurately identify and extract highly similar and finely distinguishable ground features from high-resolution satellite or UAV images, such as irrigated land and water-fed fields, lakes and rivers, residential buildings and factories. This technology has significant applications in urban planning, environmental protection, and agricultural monitoring. For example, in urban planning, extracting information on the distribution of residential areas and factory buildings can help assess the urban heat island effect; in environmental protection, precise identification of tree species and health conditions contributes to biodiversity conservation; and in agricultural monitoring, accurate assessment of crop growth status is crucial for improving crop yields.
[0003] To extract these fine-grained features, traditional methods rely on manual visual interpretation or simple thresholding. However, these methods are often time-consuming, labor-intensive, and inconsistent. Therefore, in recent years, researchers have focused on developing more advanced automated processing techniques to address this challenge. The development of machine learning, especially deep learning, has provided new insights into solving this problem. By constructing complex neural network models, not only can efficient processing of remote sensing images be achieved, but the knowledge and experience of human experts can also be simulated to some extent, thus reaching or even surpassing the accuracy of traditional methods. The application of these technologies has greatly promoted the automation of remote sensing image analysis and is expected to further enhance its intelligence level in the future.
[0004] In recent years, with advancements in deep learning, computer vision technology has achieved remarkable progress. In everyday image classification tasks, deep learning models have reached performance levels comparable to humans. Furthermore, for multi-class ground feature extraction from remote sensing images—i.e., semantic segmentation—deep learning technology has also enabled a significant leap in accuracy. Nevertheless, a gap still exists compared to the level of human experts. Therefore, one of the current research hotspots is how to leverage deep learning technology to achieve higher accuracy and more efficient fine-grained ground feature information extraction.
[0005] In the existing technology, the patent title is: "A Method for Extracting Typical Land Features from Remote Sensing Images Based on Spectral Enhancement and Dual-Path Encoding." This method utilizes a spectral enhancement module and a dual-path encoding module to extract features from both spatial and spectral information, fusing multiple features to increase the model's ability to extract and fuse different types of information. This addresses the problems of excessive manual intervention and low accuracy in previous methods. It also employs four attention mechanisms to construct different types of decoders, improving the network's ability to reconstruct information from different types of features. Furthermore, it utilizes spectral attention to enhance the network's feature extraction and enhancement capabilities for spectral information, thereby improving the network's robustness. However, this method for extracting typical land features from remote sensing images based on spectral enhancement and dual-path encoding has the following problems:
[0006] 1. Complex network structure. The spectral enhancement dual-channel coding network designs a decoder for each type of land cover, making the network structure bloated and unable to handle tasks that require extracting multiple types of land covers.
[0007] 2. Poor ability to extract and identify land features with high inter-class similarity. Although the spectral-enhanced dual-path coding network can distinguish land features with high similarity to a certain extent from a spectral perspective, its feature extraction and feature enhancement capabilities are poor, making it difficult to distinguish highly similar land features such as irrigated land and irrigated land. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a fine-grained ground feature extraction method for remote sensing images based on feature transformation and knowledge enhancement. By utilizing feature transformation, the pixel spatial information of remote sensing images is fully explored and utilized and encoded into the feature space. Through knowledge enhancement, feature space information at different scales and stages is repeatedly explored and fused, thereby realizing the intelligent extraction of fine-grained ground feature elements. This solves the problem of poor resolution and low extraction accuracy of ground features with high inter-class similarity in spectral enhancement dual-path encoding.
[0009] This invention provides a method for fine-grained ground feature extraction from remote sensing images based on feature transformation and knowledge enhancement, characterized by comprising the following steps:
[0010] (1) Construct a training dataset for training the feature knowledge enhancement network;
[0011] (2) Construct and train a feature knowledge enhancement network;
[0012] (3) Use the trained feature knowledge to enhance the network for fine-grained visualization extraction of ground features in remote sensing images.
[0013] The objective of this invention is achieved as follows:
[0014] This invention discloses a method for extracting fine-grained ground features from remote sensing images based on feature transformation and knowledge enhancement. The feature transformation module converts pixel information in remote sensing images into feature-level information, enhancing the network's ability to extract and utilize information from different domains, thereby improving its ability to extract and utilize feature information at different scales and aspects. The knowledge enhancement module integrates information from different stages and scales, further enhancing the network's ability to extract and utilize feature knowledge while simultaneously improving its ability to distinguish and extract fine-grained ground features. This achieves automated and intelligent identification of fine-grained ground features, with high accuracy in identifying features with high inter-class similarity.
[0015] Meanwhile, the fine-grained ground feature extraction method for remote sensing images based on feature transformation and knowledge enhancement of the present invention also has the following beneficial effects:
[0016] (1) Based on the traditional deep learning network structure, this invention introduces a feature conversion module to improve the network model's ability to extract feature space information. The feature conversion module can also improve the network model's accuracy in extracting fine-grained ground features.
[0017] (2) To address the issue of high similarity among some land cover types in remote sensing images, a knowledge enhancement module was designed. The knowledge enhancement module can improve the network's ability to extract fine-grained information and solve the problem of low extraction accuracy of similar land cover types.
[0018] (3) To address the problem of small differences between similar land cover types in remote sensing images and the difficulty in fusing and utilizing different feature information, this invention introduces a time-step controlled iterative method to optimize the segmentation results. This method involves multiple iterations of network processing, each controlled by a time step. The results of each iteration are integrated into the next step, providing additional refinement information to generate more accurate and complete results. This iterative process ensures that the segmentation accuracy can be gradually improved. Attached Figure Description
[0019] Figure 1 This is an overall structural diagram of the feature knowledge enhancement network of this invention;
[0020] Figure 2 Feature transformation module structure diagram;
[0021] Figure 3 Knowledge Enhancement Module Diagram;
[0022] Figure 4 ST module structure diagram;
[0023] Figure 5 RBT structure diagram;
[0024] Figure 6The figures show the experimental results: (a) the original image, (b) the labeled image, (c) the UNet experimental results, (d) the spectral enhancement dual-path coding network experimental results, and (e) the feature knowledge enhancement network experimental results. Detailed Implementation
[0025] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.
[0026] Example
[0027] For ease of description, the relevant technical terms appearing in the specific implementation method will be explained first:
[0028] ST (Spatial Transformer): Spatial Transformation Module
[0029] RBT (Resnet Block with Timestep): Time Residual Module
[0030] In this embodiment, the present invention provides a method for extracting typical land features from remote sensing images based on spectral enhancement and dual-path coding, comprising the following steps:
[0031] (1) Construct the training dataset;
[0032] (1.1) Download multiple remote sensing images and crop each remote sensing image into a tile of size m*n;
[0033] (1.2) Use semantic segmentation annotation tools to label fine-grained features in remote sensing images with different shapes. Fine-grained features include industrial areas, urban areas, rural areas, sports fields, squares, roads, overpasses, railway stations, airports, irrigated land, water-irrigated land, dry land, gardens, arbor forests, shrub forests, parks, natural meadows, artificial meadows, rivers, lakes, ponds, fishponds, snowfields, and wastelands.
[0034] (1.3) Set the pixel values corresponding to each typical land feature to 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 to generate a label image. The pixel value of industrial areas is set to 0, the pixel value of urban areas is set to 1, and so on.
[0035] (1.4) Each remote sensing image and its corresponding label image are used as a set of training data to form a training dataset;
[0036] (2) Building Figure 1 The feature knowledge enhancement network shown is trained;
[0037] A set of training data is used as input to the feature knowledge enhancement network;
[0038] In this embodiment, the feature conversion module structure is as follows: Figure 2 As shown, the feature knowledge enhancement network starts with encoder E of the feature transformation module. Encoder E contains three residual blocks and three 2x2 average pooling layers. Each residual block includes two convolutional modules, each containing a 3x3 convolutional layer, a batch normalization layer, and a ReLU activation function. The two convolutional modules form a residual connection. After the original image passes through the first residual block and the 2x2 average pooling layer, an intermediate feature map is obtained. intermediate feature map The intermediate feature map is obtained after passing through the second residual block and a 2x2 average pooling layer. intermediate feature map After the third residual block and a 2x2 average pooling layer, the initial knowledge feature map Z is obtained. C×H×W Where C, H, and W represent the number of channels, height, and width of the initial knowledge feature map, respectively; then the initial knowledge feature map Z... C×H×W Input is fed into the knowledge enhancement module for knowledge enhancement operations;
[0039] like Figure 3 As shown, the knowledge enhancement module includes a loop counter t, a knowledge fusion module, and a loop predictor; the loop counter t controls the knowledge fusion module and the loop predictor to loop three times to obtain the output of the knowledge enhancement module;
[0040] Among them, the initial knowledge feature map Z C×H×W After entering the knowledge enhancement module, the first step is to enter the knowledge fusion module. At this point, the loop counter t first checks the number of loops. If t > 1, then Z is... C×H×W and By superimposing along the channel direction, a feature map is obtained. If t = 1, then Z C×H×W The feature map is obtained by superimposing a zero-bound tensor of size C×W×H along the channel direction. Then feature map After passing through two convolutional modules and a 1×1 convolutional layer, the fused feature map P is obtained. t C×H×W ;
[0041] Fusion Feature Map P t C×H×W Then it enters the recurrent predictor; the feature map P is fused. tC×H×W First, we enter an RBT module to obtain the feature map. The feature map is then obtained by sequentially passing it through a 2x2 average pooling downsampling layer, an ST module, and an RBT module. The feature map is then obtained by sequentially passing it through a 2x2 average pooling downsampling layer, an ST module, and an RBT module. After passing through a 2x2 average pooling downsampling layer and an ST module, the feature map is obtained. Then it enters an RBT module and an upsampling layer to obtain the feature map. and Together, they enter the ST module to obtain the feature map. The feature map is obtained after passing through an RBT module and an upsampling layer. and Together they enter an ST module to obtain feature maps The feature map Y is obtained after passing through an RBT module and an upsampling layer. t C×H×W ;
[0042] In this embodiment, the ST module structure is as follows: Figure 4 As shown, assume the input feature map of the ST module is F1. C×H×W and F1 C×H×W This is achieved through three parallel fully connected layers. Will and After performing a dot product, a softmax operation is performed to obtain the result. and Perform matrix multiplication, then combine with F1 C ×H×W Adding them together gives Obtained through two parallel fully connected layers and Obtained through a fully connected layer Will and After performing a dot product, a softmax operation is performed to obtain the result. and Perform matrix multiplication and then F1 C×H×W The output of the ST module is obtained by adding them together. If the ST module has only one input, then
[0043] In this embodiment, the RBT structure is as follows: Figure 5 As shown, the RBT module includes two input branches, one of which is the time step T. 1×C T 1×C One input is a tensor vector of size 1×C, with all values being t; the other input branch is a feature map.
[0044] Suppose the feature map of an input branch is F1 C×H×W Time step T 1×C First, it goes through two fully connected layers to obtain T1. C ×H×W F1 C×H×W First, it passes through a convolutional module to obtain... F1 C×H×W T1 C×H×W Adding the bits together, we get F3. C×H×W F3 C×H×W The output of the RBT module is obtained after passing through a convolutional module.
[0045]
[0046] Y t C×H×W The decoder D and Y are input to the feature transformation module. t C×H×W First, with Z C×H×W By superimposing along the channel direction, a feature map is obtained. After passing through an upsampling layer and a residual block, it is then compared with the feature map. By superimposing along the channel direction, a feature map is obtained. Feature map After passing through an upsampling layer and a residual block, and then compared with the feature map... By superimposing along the channel direction, a feature map is obtained. Final feature map After an upsampling layer and a residual block, the staged classification probability weight map pre is obtained. t ;
[0047] At this point, the time counter t is evaluated; if the time counter t = 3, the loop ends, and pre3 is used as the final classification probability weight map; otherwise, t = t + 1 is set, and then the feature map Y is evaluated. t C×H×W With the initial knowledge feature map Z C×H×W Let's enter the knowledge enhancement module together to start the next cycle;
[0048] For example: if the time counter t = 1, then increment t by 1, Y1 C×H×W With the initial knowledge feature map Z C×H×WLet's enter the knowledge enhancement module together and proceed to the next loop; if the time counter t = 2, then increment t by 1. With the initial knowledge feature map Z C ×H×W Enter the knowledge enhancement module together to proceed to the next loop; if the time counter t equals 3, then end the loop;
[0049] Calculate the loss function value (Loss) of the spectral enhancement dual-path coding model after this round of training:
[0050]
[0051] Among them, pre t Let p represent the classification probability weight map obtained in the t-th iteration, where label represents the input label map. ti Pre represents the classification probability weight graph. t The value of the i-th pixel, l i This represents the value of the i-th pixel in the label image, and N represents the total number of pixels, N = m * n.
[0052] Finally, the feature knowledge enhancement network is trained using each set of training data until the loss function converges, at which point training stops, thus obtaining the trained feature knowledge enhancement network.
[0053] (3) Fine-grained visualization extraction of ground features from remote sensing images;
[0054] The remote sensing image to be extracted is cropped into m*n tiles, and then input into the trained feature knowledge enhancement network to output the label values 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 corresponding to each fine-grained land feature in the remote sensing image. The label values are then mapped to a color range to form a visualization image.
[0055] Figure 6 The images show the experimental results, with (a) showing the original image, (b) showing the labeled image, (c) showing the UNet experimental results, (d) showing the experimental results of the spectral-enhanced dual-path coding network, and (e) showing the experimental results of the feature-knowledge-enhanced network. A comparison reveals that (e) has the following advantages over (d): 1. It has better ability to distinguish and extract features with high similarity; 2. It misclassifies features into other categories less frequently. Overall, the feature-knowledge-enhanced network demonstrates significantly higher accuracy in fine-grained feature extraction from remote sensing images than the spectral-enhanced dual-path coding network.
[0056] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.
Claims
1. A method for fine-grained ground feature extraction from remote sensing images based on feature transformation and knowledge enhancement, characterized in that, Includes the following steps: (1) Construct a training dataset for training the feature knowledge enhancement network; (2) Construct and train a feature knowledge enhancement network; A set of training data is used as input to the feature knowledge enhancement network; The feature knowledge enhancement network begins with encoder E in the feature transformation module. Encoder E contains three residual blocks and three 2x2 average pooling layers. Each residual block includes two convolutional modules, each containing a 3x3 convolutional layer, a batch normalization layer, and a ReLU activation function. The two convolutional modules form a residual connection. After the original image passes through the first residual block and the 2x2 average pooling layer, an intermediate feature map is obtained. intermediate feature map The intermediate feature map is obtained after passing through the second residual block and a 2x2 average pooling layer. intermediate feature map After the third residual block and a 2x2 average pooling layer, the initial knowledge feature map Z is obtained. C×H×W Where C, H, and W represent the number of channels, height, and width of the initial knowledge feature map, respectively; then the initial knowledge feature map Z... C×H×W Input is fed into the knowledge enhancement module for knowledge enhancement operations; The knowledge enhancement module includes a loop counter t, a knowledge fusion module, and a loop predictor; the loop counter t controls the knowledge fusion module and the loop predictor to loop three times to obtain the output of the knowledge enhancement module; Among them, the initial knowledge feature map Z C×H×W After entering the knowledge enhancement module, the first step is to enter the knowledge fusion module. At this point, the loop counter t first checks the number of loops. If t > 1, then Z is... C×H×W and By superimposing along the channel direction, a feature map is obtained. If t = 1, then Z C×H×W The feature map is obtained by superimposing a zero-bound tensor of size C×W×H along the channel direction. Then feature map After passing through two convolutional modules and a 1×1 convolutional layer, the fused feature map P is obtained. t C×H×W ; Fusion Feature Map P t C×H×W Then it enters the recurrent predictor; the feature map P is fused. t C×H×W First, we enter an RBT module to obtain the feature map. The feature map is then obtained by sequentially passing it through a 2x2 average pooling downsampling layer, an ST module, and an RBT module. The feature map is then obtained by sequentially passing it through a 2x2 average pooling downsampling layer, an ST module, and an RBT module. After passing through a 2x2 average pooling downsampling layer and an ST module, the feature map is obtained. Then it enters an RBT module and an upsampling layer to obtain the feature map. and Together, they enter the ST module to obtain the feature map. The feature map is obtained after passing through an RBT module and an upsampling layer. and Together they enter an ST module to obtain feature maps The feature map Y is obtained after passing through an RBT module and an upsampling layer. t C×H×W ; Y t C×H×W The decoder D and Y are input to the feature transformation module. t C×H×W First, with Z C×H×W By superimposing along the channel direction, a feature map is obtained. After passing through an upsampling layer and a residual block, it is then compared with the feature map. By superimposing along the channel direction, a feature map is obtained. Feature map After passing through an upsampling layer and a residual block, and then compared with the feature map... By superimposing along the channel direction, a feature map is obtained. Final feature map After an upsampling layer and a residual block, the staged classification probability weight map pre is obtained. t ; At this point, the time counter t is evaluated; if the time counter t = 3, the loop ends, and pre3 is used as the final classification probability weight map; otherwise, t = t + 1 is set, and then the feature map Y is evaluated. t C×H×W With the initial knowledge feature map Z C×H×W Let's enter the knowledge enhancement module together to start the next cycle; Calculate the loss function value (Loss) of the spectral enhancement dual-path coding model after this round of training: Among them, pre t Let p represent the classification probability weight map obtained in the t-th iteration, where label represents the input label map. ti Pre represents the classification probability weight graph. t The value of the i-th pixel, l i This represents the value of the i-th pixel in the label image, and N represents the total number of pixels, N = m * n; Finally, the feature knowledge enhancement network is trained using each set of training data until the loss function converges, at which point training stops, thus obtaining the trained feature knowledge enhancement network. The structure of the RBT module is as follows: The RBT module includes two input branches, one of which is the time step T. 1×C T 1×C One input is a tensor vector of size 1×C, with all values being t; the other input branch is a feature map. Suppose the feature map of an input branch is F1 C×H×W Time step T 1×C First, it goes through two fully connected layers to obtain T1. C×H×W F1 C×H×W First, it passes through a convolutional module to obtain... F1 C×H×W T1 C×H×W Adding digits together, we get The output of the RBT module is obtained after passing through a convolutional module. The structure of the ST module is as follows: Assume the input feature map of the ST module is F1. C×H×W and F1 C×H×W This is achieved through three parallel fully connected layers. Will and After performing a dot product, a softmax operation is performed to obtain the result. and Perform matrix multiplication, then combine with F1 C×H×W Adding them together gives Obtained through two parallel fully connected layers and Obtained through a fully connected layer Will and After performing a dot product, a softmax operation is performed to obtain the result. and Perform matrix multiplication and then F1 C×H×W The output of the ST module is obtained by adding them together. If the ST module has only one input, then (3) Use the trained feature knowledge to enhance the network for fine-grained visualization extraction of ground features in remote sensing images.
2. The method for fine-grained ground feature extraction from remote sensing images based on feature transformation and knowledge enhancement according to claim 1, characterized in that, The training dataset is constructed as follows: (2.1) Download multiple remote sensing images and crop each remote sensing image into a patch of size m*n; (2.2) Use semantic segmentation annotation tools to label fine-grained features in remote sensing images with different shapes. Fine-grained features include industrial areas, urban areas, rural areas, sports fields, squares, roads, overpasses, railway stations, airports, irrigated land, water-irrigated land, dry land, gardens, arbor forests, shrub forests, parks, natural meadows, artificial meadows, rivers, lakes, ponds, fishponds, snowfields, and wastelands. (2.3) Set the pixel values corresponding to each fine-grained feature to 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 to generate label images corresponding to each fine-grained feature. The pixel values of industrial areas are set to 0, the pixel values of urban areas are set to 1, and so on. (2.4) Each remote sensing image and its corresponding label image are used as a set of training data to form a training dataset.
3. The method for fine-grained ground feature extraction from remote sensing images based on feature transformation and knowledge enhancement according to claim 1, characterized in that, The process of extracting fine-grained ground features from remote sensing images is as follows: The remote sensing image to be extracted is cropped into m*n tiles, and then input into the trained feature knowledge enhancement network to output the label values 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 corresponding to each fine-grained land feature in the remote sensing image. The label values are then mapped to a color range to form a visualization image.
Citation Information
Patent Citations
Method for extracting multi-ground-feature change information of high-resolution remote sensing image
CN114898212A
Wavelet-space double-attention image rain removal method and system guided by priori knowledge
CN118014890A
Cited By
Low-altitude remote sensing image-oriented ground feature fine-grained attribute extraction method
CN121661531A