Land utilization classification method based on high-frequency feature guidance
By introducing frequency domain feature encoding blocks into the segmentation network model of land use classification, combined with spatial domain characteristics, the efficient fusion of spatial and frequency domain characteristics is achieved, the problem of insufficient utilization of frequency domain information in the existing technology is solved, and the classification accuracy is improved.
Patent Information
- Application Number
- CN202510051357.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The prior art is difficult to effectively utilize frequency domain information in land use classification, which leads to difficulties in distinguishing foreground characteristics from complex background noise, which in turn affects classification accuracy.
Using an improved segmentation network model, this model combines spatial domain feature encoding blocks and frequency domain feature encoding blocks guided by high-frequency features. By extracting deep-level features in the frequency domain and adding them to the spatial domain features, it realizes efficient fusion of spatial and frequency domain features.
By introducing frequency domain information, the improved segmentation network model effectively avoids the loss of spatial information and improves the accuracy and adaptability of land use classification.
Smart Images

Figure CN119992321A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of remote sensing science and technology, and in particular relates to a land use classification method based on high-frequency feature guidance. Background Art
[0002] Land is the material basis for human survival and development, and it is also an important carrier of natural resources and social and economic activities. Land use is the process by which humans meet their own needs through specific activities based on the characteristics of the land. Scientifically classifying land use types and clarifying the meaning of each type not only has far-reaching significance for ecological and environmental protection, but also plays a key role in economic development and social progress. As the core manifestation of the transformation from natural ecosystems to artificial ecosystems, land use not only reflects the changes in natural conditions on the surface, but also deeply reflects the impact of human activities on surface cover. This characteristic makes land use a hot topic in current earth observation research. Therefore, carrying out accurate classification of land use has important guiding significance for disaster management, urban planning, environmental protection, agricultural production and other fields.
[0003] With the continuous development of earth observation technology, the temporal resolution and spatial resolution of remote sensing images have been significantly improved, and their rich spectral information and spatial characteristics provide strong data support for accurate land use classification. Based on this technological progress, land use classification methods have also developed rapidly, mainly including manual visual interpretation, pixel-based methods, object-oriented methods, and deep learning-based methods. Initially, land use classification mainly relied on manual visual interpretation, which relied on the professional knowledge of interpreters to make judgments and had high classification accuracy. However, the visual interpretation method has the disadvantages of poor timeliness and low reusability, and cannot meet the current needs of large-scale and rapid classification. Therefore, with the development of computer technology, other automated classification methods have gradually emerged and gradually replaced traditional manual interpretation. Pixel-based methods mainly include maximum likelihood classification, support vector machine, K-means clustering, etc. These methods obtain classification results by performing feature analysis on each pixel in the image. Although these methods have good classification accuracy, they do not fully consider the spatial relationship between pixels and it is difficult to distinguish complex land object types. Object-oriented classification methods (such as multi-scale segmentation and rule-based classification methods) segment images, merge adjacent pixels into objects with similar features, and then classify them according to the texture, shape and other properties of the objects. Compared with pixel-based methods, object-oriented methods can better consider spatial relationships and contextual information, thereby improving classification accuracy. However, this method also has certain limitations. For example, the selection of segmentation scale is affected by human subjective factors, and the influence of complex surface morphology on segmentation results may also lead to low segmentation accuracy, especially in areas with blurred boundaries of objects, the accuracy of target extraction is poor.
[0004] With the rapid development of artificial intelligence technology, deep learning-based methods have become the mainstream means of accurate land use classification, mainly including convolutional neural networks, recurrent neural networks and generative adversarial networks. These methods extract high-level features from the spatial domain of remote sensing images through feature extraction modules such as convolutional layers, long short-term memory, generators and discriminators, and have shown excellent performance, especially when processing large-scale data and complex landform types. These methods can not only make full use of the spectral, spatial and texture information of images, but also capture complex spatial and spatiotemporal change patterns through multi-level learning, thereby significantly improving classification accuracy and adaptability. However, existing methods mainly rely on feature extraction in the spatial domain and ignore the role of frequency domain information, which makes it difficult to distinguish foreground features from complex background noise. Some methods extract frequency domain features by wavelet transform to replace spatial domain features, thereby improving the performance and generalization ability of the model to a certain extent. However, relying solely on frequency domain features will lead to the loss of spatial information, resulting in insufficient land use classification accuracy. Summary of the invention
[0005] The embodiment of the present application provides a land use classification method based on high-frequency feature guidance, which can solve the problem of insufficient land use classification accuracy.
[0006] The application embodiment provides a land use classification method based on high-frequency feature guidance, including:
[0007] Obtain remote sensing image data of the study area;
[0008] The remote sensing image data is input into the improved segmentation network model for processing to obtain the land use classification segmentation probability map;
[0009] Based on the land use classification segmentation probability map, the land use classification results of the study area are obtained;
[0010] The feature extraction part of the improved segmentation network model includes a spatial domain feature coding block, a frequency domain feature coding block based on high-frequency feature guidance, and a feature addition module. The spatial domain feature coding block is used to extract spatial features of input data from the spatial domain, the frequency domain feature coding block is used to extract deep features of input data from the frequency domain, and the feature addition module is used to add spatial features and deep features and output feature extraction results. The spatial domain feature coding block and the frequency domain feature coding block are both input ends of the feature extraction part, and the feature addition module is the output end of the feature extraction part.
[0011] Optionally, the spatial domain feature encoding block includes a first convolution block and a second convolution block connected in sequence, the first convolution block and the second convolution block both include a convolution layer, a batch normalization layer and a Relu activation function layer connected in sequence, the input end of the convolution layer of the first convolution block is the input end of the spatial domain feature encoding block, the output end of the Relu activation function layer of the first convolution block is connected to the input end of the convolution layer of the second convolution block, and the Relu activation function layer of the second convolution block is the output end of the spatial domain feature encoding block.
[0012] Optionally, the frequency domain feature encoding block includes: a first linear layer, a discrete wavelet transform layer, a low-frequency feature enhancement module based on high-frequency features, an inverse wavelet transform layer, a first multiplication module, a layer normalization layer and a Dropout layer, a second linear layer and a Sigmoid activation function layer connected in sequence;
[0013] Among them, the input end of the first linear layer and the input end of the second linear layer are both the input ends of the frequency domain feature coding block, the output end of the Sigmoid activation function layer is connected to the input end of the first multiplication module, and the output end of the Dropout layer is the output end of the frequency domain feature coding block.
[0014] Optionally, a discrete wavelet transform layer is used to decompose the spatial feature map output by the first linear layer into low-frequency features, vertical high-frequency features, horizontal high-frequency features, and diagonal high-frequency features.
[0015] Optionally, the low-frequency feature enhancement module based on the high-frequency feature includes a splicing module, a first high-frequency processing module, a second high-frequency processing module, a low-frequency processing module, a second multiplication module, an addition module, and a three-division module;
[0016] The splicing module is used to splice vertical high-frequency features, horizontal high-frequency features and diagonal high-frequency features to obtain high-frequency features;
[0017] The first high-frequency processing module is used to process high-frequency features and output a high-frequency feature probability map;
[0018] The second high-frequency processing module is used to perform convolution processing on the high-frequency features and output the high-frequency features after the convolution processing;
[0019] The low-frequency processing module is used to process the low-frequency features and output a low-frequency feature probability map;
[0020] The second multiplication module is used to multiply the low-frequency features, the high-frequency feature probability map, and the low-frequency feature probability map, and output the low-frequency enhanced features;
[0021] The addition module is used to perform addition processing on the low-frequency features and the low-frequency enhancement features to obtain the low-frequency enhancement features after addition processing;
[0022] The three-equal division module is used to divide the high-frequency features after convolution processing into three equal parts according to the channel dimension, and output the equally divided vertical high-frequency features, horizontal high-frequency features, and diagonal high-frequency features;
[0023] The output end of the trisection module and the output end of the addition module are both connected to the input end of the inverse wavelet transform layer.
[0024] Optionally, the first high-frequency processing module includes a third convolution block and a Sigmoid activation function layer connected in sequence, the input end of the third convolution block is the input end of the first high-frequency processing module, and the output end of the Sigmoid activation function layer is the output end of the first high-frequency processing module.
[0025] Optionally, the second high-frequency processing module is a fourth convolution block.
[0026] Optionally, the low-frequency processing module includes a fifth convolution block and a Sigmoid activation function layer connected in sequence, the input end of the fifth convolution block is the input end of the low-frequency processing module, and the output end of the Sigmoid activation function layer is the output end of the low-frequency processing module.
[0027] Optionally, the third convolution block, the fourth convolution block and the fifth convolution block each include a first convolution layer, a Relu activation function layer and a second convolution layer connected in sequence;
[0028] The input end of the first convolution layer of the third convolution block is the input end of the third convolution block, and the output end of the second convolution layer of the third convolution block is the output end of the third convolution block;
[0029] The input end of the first convolution layer of the fourth convolution block is the input end of the fourth convolution block, and the output end of the second convolution layer of the fourth convolution block is the output end of the fourth convolution block;
[0030] The input end of the first convolution layer of the fifth convolution block is the input end of the fifth convolution block, and the output end of the second convolution layer of the fifth convolution block is the output end of the fifth convolution block.
[0031] Optionally, the land use classification results of the study area are obtained based on the land use classification segmentation probability map, including:
[0032] The Argmax function is used to calculate the land use classification segmentation probability map to obtain the land use classification results of the study area.
[0033] The above solution of the present application has the following beneficial effects:
[0034] In the embodiment of the present application, the remote sensing image data of the study area is processed by using an improved segmentation network model, and then the land use classification result of the study area is calculated based on the land use classification segmentation probability map obtained by processing. Among them, since the improved segmentation network model can introduce frequency domain information on the basis of retaining spatial domain features, and realize the efficient fusion of spatial and frequency domain features by extracting deep features in the frequency domain, it can effectively avoid the loss of spatial information and improve the accuracy of land use classification.
[0035] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 A flow chart of a land use classification method based on high-frequency feature guidance provided in one embodiment of the present application;
[0038] Figure 2 A schematic diagram of the structure of an improved segmentation network model provided in an embodiment of the present application;
[0039] Figure 3 A schematic diagram of the structure of a feature extraction part provided in one embodiment of the present application;
[0040] Figure 4 A schematic diagram of the structure of a frequency domain feature coding block provided in an embodiment of the present application;
[0041] Figure 5 A schematic diagram of a low-frequency feature enhancement module based on high-frequency features provided in an embodiment of the present application Figure 1 ;
[0042] Figure 6 A schematic diagram of a low-frequency feature enhancement module based on high-frequency features provided in an embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0043] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0044] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0045] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0046] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0047] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0048] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0049] In order to solve the problem of insufficient land use classification accuracy, the embodiment of the present application provides a land use classification method based on high-frequency feature guidance, which processes the remote sensing image data of the study area by using an improved segmentation network model, and then calculates the land use classification result of the study area based on the processed land use classification segmentation probability map. Among them, since the improved segmentation network model can introduce frequency domain information on the basis of retaining spatial domain features, and realizes efficient fusion of spatial and frequency domain features by extracting deep features in the frequency domain, it can effectively avoid the loss of spatial information and improve the accuracy of land use classification.
[0050] The land use classification method based on high-frequency feature guidance provided by the present application is exemplified below with reference to specific embodiments.
[0051] like Figure 1 As shown, the land use classification method based on high-frequency feature guidance provided in the embodiment of the present application includes the following steps:
[0052] Step 11, obtain remote sensing image data of the study area.
[0053] The above study area is an area that requires land use classification, and the above remote sensing image data can be obtained through satellite collection.
[0054] Step 12, input the remote sensing image data into the improved segmentation network model for processing to obtain a land use classification segmentation probability map.
[0055] It is understandable that, in order to facilitate data processing, the remote sensing image data can be normalized first, and then the normalized remote sensing image data can be input into the improved segmentation network model for processing to obtain a land use classification segmentation probability map.
[0056] The improved segmentation network model is obtained by improving the feature extraction part on the basis of the traditional segmentation network. The improved segmentation network model can introduce frequency domain information on the basis of retaining spatial domain features, and realize the efficient fusion of spatial and frequency domain features by extracting deep features in the frequency domain.
[0057] Step 13, obtaining the land use classification result of the study area based on the land use classification segmentation probability map.
[0058] In some embodiments of the present application, the Argmax function can be used to calculate the land use classification segmentation probability map to obtain the land use classification result of the study area. The land use classification result can be that some areas of the study area (such as areas with a pixel value of 1) are gardens, some areas (such as areas with a pixel value of 2) are urban residences, and some areas (such as areas with a pixel value of 3) are irrigated land.
[0059] It is worth mentioning that the improved segmentation network model can introduce frequency domain information while retaining spatial domain features, and realize the efficient fusion of spatial and frequency domain features by extracting deep features in the frequency domain, thereby effectively avoiding the loss of spatial information and improving the accuracy of land use classification.
[0060] The above improved segmentation network model is exemplarily described below in conjunction with specific embodiments.
[0061] like Figure 2 As shown in Figure 1, the improved segmentation network model is the same as the traditional segmentation network, including 4 feature extraction parts, 4 downsampling layers, intermediate feature sampling layers, 4 upsampling layers and feature classification layers. The difference is that, Figure 3 As shown, the feature extraction part of the present application includes a spatial domain feature coding block, a frequency domain feature coding block guided by high-frequency features, and a feature addition module. The spatial domain feature coding block is used to extract spatial features of input data from the spatial domain, and the frequency domain feature coding block is used to extract deep features of input data from the frequency domain. The feature addition module is used to add the spatial features and the deep features and output the feature extraction results. The spatial domain feature coding block and the frequency domain feature coding block are both input ends of the feature extraction part, and the feature addition module is the output end of the feature extraction part.
[0062] Among them, the spatial domain feature encoding block includes a first convolution block and a second convolution block connected in sequence, the first convolution block and the second convolution block both include a convolution layer, a batch normalization layer and a Relu activation function layer connected in sequence, the input end of the convolution layer of the first convolution block is the input end of the spatial domain feature encoding block, the output end of the Relu activation function layer of the first convolution block is connected to the input end of the convolution layer of the second convolution block, and the Relu activation function layer of the second convolution block is the output end of the spatial domain feature encoding block.
[0063] Specifically, remote sensing image data Input into the first convolution block to obtain the spatial feature map The spatial feature map Input into the second convolution block to obtain the spatial feature map in The dimensions are 1×3×w×h; and The size of is 1×64×w×h. The size of w and h is 512. The convolution kernel size in the above convolution layer is 3, the stride is 1, and the padding is 1.
[0064] like Figure 4 As shown, the frequency domain feature encoding block includes: a first linear layer, a discrete wavelet transform layer, a low-frequency feature enhancement module based on high-frequency features, an inverse wavelet transform layer, a first multiplication module (i.e. Figure 4 The circle with an × in the middle), the layer normalization layer and the Dropout layer, the second linear layer and the Sigmoid activation function layer connected in sequence (i.e. Figure 4 The input end of the first linear layer and the input end of the second linear layer are both input ends of the frequency domain feature coding block, the output end of the Sigmoid activation function layer is connected to the input end of the first multiplication module, and the output end of the Dropout layer is the output end of the frequency domain feature coding block. The discrete wavelet transform layer is used to decompose the spatial feature map output by the first linear layer into low-frequency features, vertical high-frequency features, horizontal high-frequency features, and diagonal high-frequency features.
[0065] Specifically, remote sensing image data Input to the first linear layer to obtain the spatial feature map Remote sensing image data Input to the second linear layer to get the feature map The feature map Input to the Sigmoid activation function to get the feature probability map in and The dimensions are 1×64×w×h.
[0066] The discrete wavelet transform layer uses the Haar basis to decompose the spatial features into low-frequency and high-frequency parts. Its size is 1×64×w×h, and it is decomposed in two directions: horizontal and vertical. First, the spatial feature map is sorted by row Decompose.
[0067] For any location Low frequency part:
[0068]
[0069] High frequency part:
[0070]
[0071] where F L (a, b) and F H The dimensions of (a,b) are Then decompose each row again, the low-frequency part:
[0072]
[0073] Vertical high frequency section:
[0074]
[0075] Horizontal high frequency section:
[0076]
[0077] Diagonal high frequency part:
[0078]
[0079] After decomposition, the low-frequency feature F is obtained. LL (a, b) and high-frequency features F HL (a,b),F LH (a, b) and F HH (a,b), both of which are
[0080] like Figure 5 As shown, the above-mentioned low-frequency feature enhancement module based on high-frequency features includes a splicing module (i.e. Figure 5 The circle with C in the middle), the first high-frequency processing module, the second high-frequency processing module, the low-frequency processing module, the second multiplication module (i.e. Figure 5 The circle with × in the middle), the addition module (i.e. Figure 5 The module is divided into three equal parts.
[0081] The splicing module is used to splice the vertical high-frequency features, the horizontal high-frequency features and the diagonal high-frequency features to obtain the high-frequency features; the first high-frequency processing module is used to process the high-frequency features and output the high-frequency feature probability map; the second high-frequency processing module is used to convolve the high-frequency features and output the high-frequency features after the convolution processing; the low-frequency processing module is used to process the low-frequency features and output the low-frequency feature probability map; the second multiplication module is used to multiply the low-frequency features, the high-frequency feature probability map and the low-frequency feature probability map, and output the low-frequency enhancement feature; the addition module is used to add the low-frequency features and the low-frequency enhancement features to obtain the low-frequency enhancement features after the addition processing; the three-equal division module is used to divide the high-frequency features after the convolution processing into three equal parts according to the channel dimension, and output the equally divided vertical high-frequency features, horizontal high-frequency features and diagonal high-frequency features; the output end of the three-equal division module and the output end of the addition module are both connected to the input end of the inverse wavelet transform layer.
[0082] Among them, Figure 5 and Figure 6 As shown, the first high frequency processing module includes a third convolution block and a Sigmoid activation function layer (i.e. Figure 6 The input end of the third convolution block is the input end of the first high-frequency processing module, and the output end of the Sigmoid activation function layer is the output end of the first high-frequency processing module. The second high-frequency processing module is the fourth convolution block. The low-frequency processing module includes the fifth convolution block and the Sigmoid activation function layer (i.e. Figure 6 The input end of the fifth convolution block is the input end of the low-frequency processing module, and the output end of the Sigmoid activation function layer is the output end of the low-frequency processing module.
[0083] The third convolution block, the fourth convolution block and the fifth convolution block all include a first convolution layer, a Relu activation function layer and a second convolution layer connected in sequence. The input end of the first convolution layer of the third convolution block is the input end of the third convolution block, and the output end of the second convolution layer of the third convolution block is the output end of the third convolution block; the input end of the first convolution layer of the fourth convolution block is the input end of the fourth convolution block, and the output end of the second convolution layer of the fourth convolution block is the output end of the fourth convolution block; the input end of the first convolution layer of the fifth convolution block is the input end of the fifth convolution block, and the output end of the second convolution layer of the fifth convolution block is the output end of the fifth convolution block. The convolution kernel size in the first convolution layer is 1, the step size is 1, and the padding is 0. The convolution kernel size in the second convolution layer is 3, the step size is 1, and the padding is 1.
[0084] Specifically, the high-frequency feature F HL (a,b),F LH (a, b) and F HH (a, b) High-frequency features are obtained by channel splicing Its size is The high frequency features Input into the third convolution block to obtain high-frequency features The high frequency features Input Sigmoid activation function to get high-frequency feature probability map The high frequency features Input into the fourth convolution block to obtain high-frequency features in and The size is The size is
[0085] The low frequency feature F LL (a, b) are input into the fifth convolutional block to obtain low-frequency features. The low frequency features Input Sigmoid activation function to get low-frequency feature probability map Then, the low-frequency features High frequency feature probability map And low frequency feature probability map Element-by-element multiplication to obtain low-frequency enhancement features Next, the low-frequency feature F LL (a, b) and low-frequency enhancement features Add element by element to get low-frequency enhancement features The high frequency features Divide into three equal parts according to the channel dimension and obtain the high-frequency features F' HL (a,b),F' LH (a,b) and F' HH (a,b). Then the low frequency enhancement feature and high frequency features F' HL (a,b),F' LH (a,b) and F' HH (a,b) is mapped back to the spatial domain through the inverse wavelet transform layer. F' HL (a,b),F' LH (a,b) and F' HH The dimensions of (a,b) are
[0086] The inverse wavelet transform first restores the low-frequency and high-frequency parts of each column, and then restores the low-frequency and high-frequency parts of each row. The column direction is restored to satisfy the following formula:
[0087]
[0088] Restore the row direction to satisfy the following formula:
[0089]
[0090] The spatial feature map is obtained by the above formula The spatial feature map With feature probability plot Multiply element by element to get the spatial feature map Through the spatial feature map After the layer normalization layer, we get the spatial feature map The spatial feature map After the Dropout layer, we get the spatial feature map The spatial feature map With spatial feature map Add element by element to get the fused feature map extracted in spatial domain and frequency domain in, and The dimensions are 1×64×w×h.
[0091] The parameters used in the above explanation of the feature extraction part are the parameters of the first feature extraction part in the model (the first one from left to right).
[0092] Figure 2The first sampling layer (from left to right) in includes a maximum pooling layer with a step size of 2. The fused feature map Input the maximum pooling layer to obtain the spatial feature map Its size becomes
[0093] Figure 2 The convolution kernel size in the convolution layer of the spatial domain feature encoding block in the second feature extraction part is 3, the stride is 1, and the padding is 1. Input into the first convolution block to obtain the spatial feature map The spatial feature map Input into the second convolution block to obtain the spatial feature map and The size is
[0094] against Figure 2 The second frequency domain feature encoding block of the second feature extraction part (the second one from left to right) in Input to the first linear layer to obtain the spatial feature map The spatial feature map Input to the second linear layer to get the feature map The feature map Input to the Sigmoid activation function to get the feature probability map in and The dimensions are
[0095] It should be noted that the four feature extraction parts in the model have the same structure but slightly different sizes, the two downsampling layers have the same structure but slightly different sizes, and the four upsampling layers have the same structure but slightly different sizes. Therefore, the functions of the same structures in the model are not described in detail here, and only the sizes are exemplified.
[0096] Specifically, the second downsampling layer includes a maximum pooling layer with a step size of 2. The fused feature map Input the maximum pooling layer to obtain the spatial feature map Its size becomes The size of the spatial domain feature encoding code output of the third feature extraction part is The size of the frequency domain feature encoding output of the third feature extraction part is The output of the third downsampling layer is The size of the spatial domain feature encoding code output of the fourth feature extraction part is The size of the frequency domain feature encoding output of the fourth feature extraction part is The size of the output of the fourth downsampling layer is The intermediate feature sampling layer consists of the first convolution block and the second convolution block. The first convolution block consists of a convolution layer, a batch normalization layer, and a Relu activation function layer in sequence; the second convolution block consists of a convolution layer, a batch normalization layer, and a Relu activation function layer in sequence. Input into the first convolution block to obtain the spatial feature map The spatial feature map Input into the second convolution block to obtain the spatial feature map and The size is In the above convolutional layer, the convolution kernel size is 3, the stride is 1, and the padding is 1.
[0097] The above is an introduction to the encoder of the model. The following is an introduction to the decoder of the model.
[0098] like Figure 2 As shown in the figure, the decoder consists of 4 upsampling layers and a feature classification layer. Specifically, the decoder is used to perform the following process:
[0099] Step 1: The first upsampling layer (from right to left) consists of a bilinear interpolation layer, a first convolutional block, and a second convolutional block. The first convolutional block consists of a convolutional layer, a batch normalization layer, and a Relu activation function layer; the second convolutional block consists of a convolutional layer, a batch normalization layer, and a Relu activation function layer. Input the bilinear interpolation layer to get the spatial feature map Its size is The spatial feature map With spatial feature map (Output of the fourth feature extraction part) is obtained by splicing according to the channel dimension, and the spliced feature map Its size is 1× The concatenated feature map Input the first convolution block to get the feature map Its size is The feature map Input the second convolution block to get the feature map Its size is In the above convolutional layer, the convolution kernel size is 3, the stride is 1, and the padding is 1.
[0100] Step 2: The second upsampling layer is composed of a bilinear interpolation layer, a first convolutional block, and a second convolutional block. The first convolutional block is composed of a convolutional layer, a batch normalization layer, and a Relu activation function layer; the second convolutional block is composed of a convolutional layer, a batch normalization layer, and a Relu activation function layer. Input the bilinear interpolation layer to get the spatial feature map Its size is The spatial feature map With spatial feature map (Output of the third feature extraction part) is obtained by splicing according to the channel dimension, splicing feature map Its size is The concatenated feature map Input the first convolution block to get the feature map Its size is The feature map Input the second convolution block to get the feature map Its size is In the above convolutional layer, the convolution kernel size is 3, the stride is 1, and the padding is 1.
[0101] Step 3: The third upsampling layer is composed of a bilinear interpolation layer, a first convolutional block, and a second convolutional block. The first convolutional block is composed of a convolutional layer, a batch normalization layer, and a Relu activation function layer; the second convolutional block is composed of a convolutional layer, a batch normalization layer, and a Relu activation function layer. Input the bilinear interpolation layer to get the spatial feature map Its size is 1×128× The spatial feature map With spatial feature map (Output of the second feature extraction part) is obtained by splicing according to the channel dimension, splicing feature map Its size is The concatenated feature map Input the first convolution block to get the feature map Its size is The feature map Input the second convolution block to get the feature map Its size is In the above convolutional layer, the convolution kernel size is 3, the stride is 1, and the padding is 1.
[0102] Step 4: The fourth upsampling layer is composed of a bilinear interpolation layer, a first convolutional block, and a second convolutional block. The first convolutional block is composed of a convolutional layer, a batch normalization layer, and a Relu activation function layer; the second convolutional block is composed of a convolutional layer, a batch normalization layer, and a Relu activation function layer. Input the bilinear interpolation layer to get the spatial feature map Its size is 1×64×w×h. With spatial feature map (Output of the first feature extraction part) is obtained by splicing according to the channel dimension, splicing feature map Its size is 1×128×w×h. Input the first convolution block to get the feature map Its size is 1×32×w×h. Input the second convolution block to get the feature map Its size is 1×32×w×h. In the above convolutional layer, the convolution kernel size is 3, the stride is 1, and the padding is 1.
[0103] Step 5: The feature classification layer is composed of a convolution layer and a Softmax activation function, with a convolution kernel and a step size of 1. Input convolution layer to get feature map The feature map Input the Softmax activation function to obtain the land use classification segmentation probability map F P , whose size is 1×NC×w×h, where NC is the number of categories of land use classification.
[0104] It should be noted that before using the improved segmentation network model to process remote sensing image data, the segmentation network model needs to be trained, and the training method can adopt the training method of the traditional segmentation network.
[0105] In the actual training process, the training data is processed as follows:
[0106] Land use category label preprocessing: Land use category label files contain two data formats: vector and raster. For the vector format label file S, S = {S1, S2, S3, ..., S i ,…,S N},S i Represents the i-th vector format label file. Use the gdal.RasterizeLayer function in the GDAL third-party library in Python to convert S iAccording to the category attributes and vectors contained in the Class field, the high-resolution remote sensing image SI, SI = {SI1, SI2, SI3, ..., SI i ,…,SI N}, converted to single-band raster data SL, SL = {SL1, SL2, SL3, …, SL i ,…,SL N For raster format tag file T, T = {T1, T2, T3, ..., T i ,…,T M}, T i Represents the i-th raster format label file. Use the Numpy third-party library in Python to correspond to the high-resolution remote sensing image TI according to the color map and raster in the metadata, TI = {TI1, TI2, TI3, ..., TI i ,…,TI M}, converted to single-band raster data TL, TL = {TL1, TL2, TL3, …, TL i ,…,TL M}.
[0107] The obtained single-band raster data {SL, TL} and the corresponding high-resolution remote sensing images {SI, TI} are divided into training set and validation set in a ratio of 7:3. Considering the performance of the computer hardware system, the single-band raster data {SL, TL} in the training set are divided into Train Cropping is performed with a 50% overlap rate and a size of 512 × 512 pixels to obtain X training images, and the xth image is The high-resolution remote sensing images {SI,TI} in the training set Train Cropping with 50% overlap and 512×512 pixels to get X training labels, the xth image is The single-band raster data {SL, TL} in the validation set Val and the corresponding high-resolution remote sensing image {SI,TI} Val According to the non-overlapping rate and 512×512 pixel size, we get Y verification images and verification labels respectively. The yth verification image is Verification label of the yth The above training set and validation set are stored to obtain the land use classification dataset.
[0108] During the training process, the multi-classification cross entropy loss function Loss is used to calculate the land use classification segmentation probability map F P and the xth label L x TThe segmentation network model is trained using the AdamW optimizer, cosine annealing learning rate and loss function Loss to obtain the optimized segmentation network model. The initial learning rate is set to 0.001, the batch size is set to 8, and the epoch is set to 128. The T_max of the cosine annealing learning rate is set to 16.
[0109] In order to verify the accuracy of the land use classification method of the present application, the land use classification method of the present application is used to process the remote sensing image land cover dataset GID. GID is a large-scale high-resolution remote sensing image land cover dataset constructed based on Gaofen-2 satellite data. The GID dataset is divided into two parts: a large-scale classification set (GID-5) and a fine land cover set (GID-15). The fine land cover set (GID-15) contains 15 categories including rice fields, irrigated land, dry land, gardens, arbor forests, shrub forests, natural grasslands, artificial grasslands, industrial land, urban residential buildings, rural residential buildings, transportation land, rivers, lakes, ponds, etc., with a total of 30,000 image blocks.
[0110] RS3Mamba is used as the baseline model. The present application is embedded into the spatial domain feature encoding block of its auxiliary encoder: the Mamba module, and the frequency domain features are introduced while retaining the spatial domain features, and the fusion of spatial and frequency domain features is realized. In the above-mentioned GID-15 dataset, the average F1 Score of the present application reached 84.51%, and the average intersection over union (mIoU) reached 74.57%. Compared with the baseline model, the F1 Score and IoU of the garden category increased by 3.22% and 3.65%, and the F1 Score and IoU of the pond category increased by 1.84% and 1.99% respectively. Quantitative data show that the method of the present application can improve the accuracy of land use classification.
[0111] In summary, the land use classification method of this application has the following advantages:
[0112] (1) Frequency domain coding module: This module is flexibly designed and can be seamlessly integrated with the existing spatial domain feature coding module. It can map spatial domain information to the frequency domain for feature extraction, and then map it back to the spatial domain after the extraction is completed, thereby achieving efficient fusion of frequency domain and spatial domain features; (2) Low-frequency feature enhancement module based on high-frequency features: This module converts high-frequency information into a probabilistic form based on the extraction of high-frequency and low-frequency features, thereby promoting the extraction and optimization of low-frequency features and enhancing the ability of low-frequency features to distinguish between different types of land objects. This application provides a new theoretical basis and technical means for the precise classification of land use in complex scenes and large areas, which helps to solve the current problem of insufficient classification accuracy.
[0113] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A land use classification method based on high-frequency feature guidance, characterized in that: include: Obtain remote sensing image data of the study area; Inputting the remote sensing image data into the improved segmentation network model for processing to obtain a land use classification segmentation probability map; Obtaining the land use classification result of the study area based on the land use classification segmentation probability map; The feature extraction part of the improved segmentation network model includes a spatial domain feature coding block, a frequency domain feature coding block based on high-frequency feature guidance, and a feature addition module. The spatial domain feature coding block is used to extract the spatial features of the input data from the spatial domain, the frequency domain feature coding block is used to extract the deep features of the input data from the frequency domain, and the feature addition module is used to add the spatial features and the deep features and output the feature extraction results. The spatial domain feature coding block and the frequency domain feature coding block are both input ends of the feature extraction part, and the feature addition module is the output end of the feature extraction part.
2. The land use classification method according to claim 1, characterized in that: The spatial domain feature encoding block includes a first convolution block and a second convolution block connected in sequence, the first convolution block and the second convolution block both include a convolution layer, a batch normalization layer and a Relu activation function layer connected in sequence, the input end of the convolution layer of the first convolution block is the input end of the spatial domain feature encoding block, the output end of the Relu activation function layer of the first convolution block is connected to the input end of the convolution layer of the second convolution block, and the Relu activation function layer of the second convolution block is the output end of the spatial domain feature encoding block.
3. The land use classification method according to claim 1, characterized in that: The frequency domain feature encoding block includes: a first linear layer, a discrete wavelet transform layer, a low-frequency feature enhancement module based on high-frequency features, an inverse wavelet transform layer, a first multiplication module, a layer normalization layer and a Dropout layer, a second linear layer and a Sigmoid activation function layer connected in sequence; Among them, the input end of the first linear layer and the input end of the second linear layer are both the input end of the frequency domain feature coding block, the output end of the Sigmoid activation function layer is connected to the input end of the first multiplication module, and the output end of the Dropout layer is the output end of the frequency domain feature coding block.
4. The land use classification method according to claim 3, characterized in that: The discrete wavelet transform layer is used to decompose the spatial feature map output by the first linear layer into low-frequency features, vertical high-frequency features, horizontal high-frequency features and diagonal high-frequency features.
5. The land use classification method according to claim 4, characterized in that: The low-frequency feature enhancement module based on high-frequency features includes a splicing module, a first high-frequency processing module, a second high-frequency processing module, a low-frequency processing module, a second multiplication module, an addition module, and a three-division module; The splicing module is used to splice the vertical high-frequency feature, the horizontal high-frequency feature and the diagonal high-frequency feature to obtain a high-frequency feature; The first high-frequency processing module is used to process the high-frequency features and output a high-frequency feature probability map; The second high-frequency processing module is used to perform convolution processing on the high-frequency features and output the high-frequency features after the convolution processing; The low-frequency processing module is used to process the low-frequency features and output a low-frequency feature probability map; The second multiplication module is used to multiply the low-frequency feature, the high-frequency feature probability map, and the low-frequency feature probability map to output a low-frequency enhancement feature; The adding module is used to perform an adding process on the low-frequency feature and the low-frequency enhancement feature to obtain the low-frequency enhancement feature after the adding process; The three-equal division module is used to divide the high-frequency features after convolution processing into three equal parts according to the channel dimension, and output the equally divided vertical high-frequency features, horizontal high-frequency features and diagonal high-frequency features; The output end of the three-division module and the output end of the addition module are both connected to the input end of the inverse wavelet transform layer.
6. The land use classification method according to claim 5, characterized in that: The first high-frequency processing module includes a third convolution block and a Sigmoid activation function layer connected in sequence, the input end of the third convolution block is the input end of the first high-frequency processing module, and the output end of the Sigmoid activation function layer is the output end of the first high-frequency processing module.
7. The land use classification method according to claim 6, characterized in that: The second high frequency processing module is a fourth convolution block.
8. The land use classification method according to claim 7, characterized in that: The low-frequency processing module includes a fifth convolution block and a Sigmoid activation function layer connected in sequence, the input end of the fifth convolution block is the input end of the low-frequency processing module, and the output end of the Sigmoid activation function layer is the output end of the low-frequency processing module.
9. The land use classification method according to claim 8, characterized in that: The third convolution block, the fourth convolution block and the fifth convolution block each include a first convolution layer, a Relu activation function layer and a second convolution layer connected in sequence; The input end of the first convolution layer of the third convolution block is the input end of the third convolution block, and the output end of the second convolution layer of the third convolution block is the output end of the third convolution block; The input end of the first convolution layer of the fourth convolution block is the input end of the fourth convolution block, and the output end of the second convolution layer of the fourth convolution block is the output end of the fourth convolution block; The input end of the first convolution layer of the fifth convolution block is the input end of the fifth convolution block, and the output end of the second convolution layer of the fifth convolution block is the output end of the fifth convolution block.
10. The land use classification method according to claim 1, characterized in that: The obtaining of the land use classification result of the study area based on the land use classification segmentation probability map includes: The Argmax function is used to calculate the land use classification segmentation probability map to obtain the land use classification result of the study area.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method and device, computer equipment and storage medium
CN113034506A
Land utilization classification method based on deep space-time mode interaction network
CN114863266A
Progressive double-decoupling SAR (Synthetic Aperture Radar)-assisted remote sensing image thick cloud removal method
CN117689579A
Remote sensing image prediction method and system based on multi-scale encoder and decoder structure
CN117877033A
Land utilization classification and image reconstruction method and device, medium and product
CN118212467A