High-resolution city green space classification method and device based on double decoders and fine segmentation head, equipment and medium

The dual decoder and fine segmentation head method addresses the limitations of existing deep learning networks by integrating local and global features in high-resolution satellite imagery, resulting in improved accuracy and efficiency for urban green space classification.

CN120318704APending Publication Date: 2025-07-15TSINGHUA UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510376086.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art has problems such as local feature limitations, loss of spatial details and difficulty in handling complex landscapes in urban green space recognition in high-resolution remote sensing images, resulting in low classification accuracy and efficiency.

Method used

A high-resolution urban green space classification method based on dual decoder and fine segmentation head is adopted. Local and global information are captured through the combination of internal encoder and external encoder, and feature enhancement is used for fine segmentation head to achieve high-precision classification.

Benefits of technology

It significantly improves the accuracy and efficiency of green space classification, especially in complex scenarios, and is suitable for the automated classification of large-scale high-resolution satellite city image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318704A_ABST
    Figure CN120318704A_ABST
Patent Text Reader

Abstract

The invention provides a high-resolution city green space classification method and device based on double decoders and a fine segmentation head, equipment and a medium, and relates to the technical field of image classification. The method comprises the following steps: acquiring a high-resolution satellite remote sensing image; preprocessing the satellite remote sensing image to obtain an image training sample and an image test sample; inputting an image test sample into the trained green space classification neural network for processing, and determining a green space classification result; wherein the green space classification neural network is obtained by training the image training sample. According to the embodiment of the invention, the defect of low accuracy of the traditional image classification technology and the deep learning network image classification at the present stage in the prior art is overcome, and large-scale high-resolution satellite city image data can be quickly processed through a classification algorithm for capturing local and global information at the same time; urban green space images are automatically classified in a high-precision mode, and the precision and efficiency of green space classification are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and in particular, to a high-resolution urban green space classification method, device, equipment and medium based on a dual decoder and a fine segmentation head. Background Art

[0002] With the acceleration of the global urbanization process, the role of urban green spaces in improving the living environment of residents and enhancing the ecological benefits of cities has become increasingly prominent. Urban green spaces provide various ecological benefits to residents, including reducing noise, regulating temperature, purifying air, and improving physical and mental health. Therefore, how to effectively utilize urban green spaces and maximize their ecological benefits has become a key focus for achieving high-quality sustainable development.

[0003] However, due to the wide and complex distribution of urban green spaces, accurately identifying and assessing their conditions has always faced major challenges. Although traditional field survey methods can directly obtain green space information, they often require a large amount of manpower and material resources, are time-consuming and laborious in the survey process, and have limited spatial coverage, which is insufficient when conducting large-scale green space analysis.

[0004] In recent years, with the rapid development of satellite remote sensing technology, the acquisition of high-resolution images has become more convenient. Compared with traditional field surveys, remote sensing technology provides a more efficient and scalable solution for urban green space mapping. Through remote sensing technology, researchers can achieve comprehensive, extensive, and large-scale analysis of green spaces and extract a large amount of key information. At the same time, with the rapid development of deep learning technology, semantic segmentation technology based on remote sensing images has been further applied and optimized. These technologies use a convolutional neural network (CNN) as an encoder and introduce a Transformer module in the decoding stage, which can improve the semantic segmentation accuracy while reducing the model complexity.

[0005] Although existing deep learning technologies provide effective solutions for semantic segmentation of remote sensing images, there are still some deficiencies that cannot be ignored: (1) Limitations of local features: Existing CNN encoders mainly capture local features of images and are difficult to fully reflect the large-scale long-range dependencies required in urban green space classification. This makes it impossible for the model to comprehensively grasp global information when processing complex urban green space landscapes, thus affecting the classification accuracy. (2) Loss of spatial details: The current mainstream decoder design is relatively simple, and it is difficult to effectively retrieve the lost spatial details in the image during the processing of high-resolution remote sensing images. This directly leads to insufficient ability to identify subtle features in urban green spaces. (3) Difficulty in processing complex landscapes: Urban green spaces contain a large number of complex landscape features, and traditional semantic segmentation methods require relatively complex data processing techniques to extract accurate green space information. However, this complexity increases the difficulty of data processing and restricts the efficiency and accuracy of the model in practical applications.

[0006] In summary, currently, in the aspect of identifying complex urban green spaces in high-resolution remote sensing images, the accuracy and efficiency of classifying the green space are relatively low. Summary of the Invention

[0007] The present invention provides a high-resolution urban green space classification method, device, equipment and medium based on a dual decoder and a fine segmentation head, to solve the defect of low accuracy of traditional image classification technology and current deep learning network image classification in the prior art. Through a classification algorithm that captures both local and global information simultaneously, it realizes rapid processing of large-scale high-resolution satellite urban image data, automatically classifies urban green space images with high precision, and significantly improves the accuracy and efficiency of green space classification.

[0008] In the first aspect, the present invention provides a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head, including the following steps: Obtain high-resolution satellite remote sensing images; Preprocess the satellite remote sensing images to obtain image training samples and image test samples; Input the image test samples into a trained green space classification neural network for processing to determine the green space classification result; wherein, the green space classification neural network is obtained by training the image training samples.

[0009] Preferably, according to the high-resolution urban green space classification method provided by the present invention, the preprocessing of the satellite remote sensing images to obtain image training samples and image test samples includes: Perform pixel-by-pixel annotation on the green space features of the satellite remote sensing images to obtain an annotated image; Perform the same-scale segmentation processing on the labeled image and the satellite remote sensing image to obtain a labeled image sample and an original image sample, and use the labeled image sample as the image test sample; Perform cropping processing on the original image sample to obtain a first original sub-image training sample and a second original sub-image training sample, and use the first original sub-image training sample and the second original sub-image training sample as the image training sample.

[0010] Preferably, according to a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention, the green space classification neural network includes: an internal encoder, an external encoder, a decoder, and a fine segmentation head; The training method of the green space classification neural network includes: Input the first original sub-image training sample into the internal encoder for downsampling processing to obtain local image features; Input the second original sub-image training sample into the external encoder for downsampling processing to obtain global image features; Input the local image features and the global image features into the decoder for feature fusion processing to obtain fused image features; Input the fused image features into the fine segmentation head for feature enhancement processing to obtain target image features; Input the target image features and the image test sample into a preset training evaluation model for evaluation processing, and use the target image features with an evaluation score greater than or equal to a preset score threshold as the training output result of the trained green space classification neural network; Use the training parameters corresponding to the training output result as the model parameters of the trained green space classification neural network.

[0011] Preferably, according to a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention, the step of inputting the local image features and the global image features into the decoder for feature fusion processing to obtain fused image features includes: Generate a local feature vector by passing the local image features through a multi-layer perceptron, and generate a first global feature vector and a second global feature vector by passing the global image features through a multi-layer perceptron; Perform calculation processing on the first global feature vector, the second global feature vector, and the local feature vector based on an activation function to generate an attention weight matrix; Perform a product operation on the attention weight matrix and the decoder features of the decoder to generate the fused image features.

[0012] Preferably, for a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention, the step of inputting the fused image features into the fine segmentation head for feature enhancement processing to obtain target image features includes: Inputting the fused image features into the fine segmentation head for channel feature enhancement processing to obtain channel-enhanced features; Performing spatial feature enhancement processing on the channel-enhanced features to obtain the target image features.

[0013] Preferably, for a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention, the step of inputting the fused image features into the fine segmentation head for channel feature enhancement processing to obtain channel-enhanced features includes: Using the fine segmentation head to perform global max pooling processing on the fused image features to obtain a global max feature vector, and performing global average pooling processing on the fused image features to obtain a global average feature vector; Processing the global max feature vector and the global average feature vector based on the multi-layer perceptron of the fine segmentation head to generate a channel attention map; Applying the channel attention map to the fused image features to generate the channel-enhanced features.

[0014] Preferably, for a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention, the step of performing spatial feature enhancement processing on the channel-enhanced features to obtain the target image features includes: Performing global max pooling processing on the channel-enhanced features to obtain a global max feature map, and performing global average pooling processing on the channel-enhanced features to obtain a global average feature map; Performing convolution processing on the global max feature map and the global average feature map to generate a spatial attention map; Applying the spatial attention map to the channel-enhanced features to generate spatially channel-enhanced features; Performing prediction processing on the spatially channel-enhanced features to obtain the target image features; wherein the prediction processing includes at least a single ReLu activation function, a single convolutional layer, and upsampling operation processing.

[0015] In a second aspect, the present invention also provides a high-resolution urban green space classification device based on a dual decoder and a fine segmentation head, including: An acquisition module, configured to acquire high-resolution satellite remote sensing images; A preprocessing module, configured to preprocess the satellite remote sensing images to obtain image training samples and image test samples; A green space classification result determination module, which is configured to input the image test samples into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is obtained by training the image training samples.

[0016] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the high-resolution urban green space classification method based on a dual decoder and a fine segmentation head as described in any one of the above.

[0017] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the high-resolution urban green space classification method based on a dual decoder and a fine segmentation head as described in any one of the above.

[0018] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the high-resolution urban green space classification method based on a dual decoder and a fine segmentation head as described in any one of the above.

[0019] A high-resolution urban green space classification method, device, equipment, and medium based on a dual decoder and a fine segmentation head provided by the present invention, by obtaining high-resolution satellite remote sensing images; preprocessing the satellite remote sensing images to obtain image training samples and image test samples; inputting the image test samples into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is obtained by training the image training samples. It is used to solve the defect of low accuracy of traditional image classification technology and current deep learning network image classification in the prior art. Through a classification algorithm that captures local and global information simultaneously, it realizes the rapid processing of large-scale high-resolution satellite urban image data, automatically classifies urban green space images with high precision, and significantly improves the accuracy and efficiency of green space classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is one of the flow diagrams of a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention.

[0022] Figure 2 It is the second schematic diagram of a high - resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention.

[0023] Figure 3 It is the structural schematic diagram of the fine segmentation head provided by the present invention.

[0024] Figure 4 It is the structural schematic diagram of a high - resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention.

[0025] Figure 5 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0027] First, several nouns involved in the present invention are analyzed: CNN: It is a feed - forward neural network with a deep structure of convolutional layers, mainly used for processing image data, and can also be used in fields such as speech recognition and natural language processing. It automatically extracts features of the input data through convolutional layers, reduces the complexity of the model, and reduces the risk of overfitting through weight sharing.

[0028] The key components of CNN include convolutional layers, pooling layers and fully - connected layers. The convolutional layer is responsible for feature extraction, the pooling layer is used to reduce the spatial dimension of the feature map, and the fully - connected layer performs classification or regression tasks.

[0029] Transformer: It is a brand - new neural network architecture. The core idea is to use the self - attention mechanism to model long - distance dependencies in sequence data, so as to better capture semantic information. It completely abandons the traditional recurrent neural network (RNN) structure and only relies on the attention mechanism to process sequence data. The Transformer model mainly consists of two parts: an encoder and a decoder, and each part is composed of multiple stacked identical modules. These modules usually contain multi - head self - attention mechanisms and feed - forward neural networks.

[0030] ReLu (Rectified Linear Unit): That is, the Rectified Linear Unit, which is a commonly used non-linear activation function. Its mathematical expression is f(x) = max(0, x), meaning that when the input x is positive, the output is equal to the input; when the input x is negative, the output is 0.

[0031] The ReLu activation function plays a role in introducing non-linear factors in the neural network, enabling the neural network to learn and represent more complex functional relationships. It is very efficient in calculation and does not involve any exponential operations, so it is widely used in practice.

[0032] In related technologies, there are at least the following technical problems: (1) Limitations of local features: Existing CNN encoders mainly capture local features of images and are difficult to fully reflect the large-scale long-range dependencies required in urban green space classification. This makes it impossible for the model to comprehensively grasp global information when processing complex urban green space landscapes, thus affecting the accuracy of classification.

[0033] (2) Loss of spatial details: The current mainstream decoder design is relatively simple. During the processing of high-resolution remote sensing images, it is difficult to effectively retrieve the lost spatial details in the images. This directly leads to insufficient ability to identify subtle features in urban green spaces.

[0034] (3) Difficulty in processing complex landscapes: Urban green spaces contain a large number of complex landscape features. Traditional semantic segmentation methods require relatively complex data processing techniques to extract accurate green space information. However, this complexity increases the difficulty of data processing and restricts the efficiency and accuracy of the model in practical applications.

[0035] In summary, in the aspect of identifying complex urban green spaces in high-resolution remote sensing images, there is an urgent need for a classification algorithm that can simultaneously capture local and global information to overcome the deficiencies of existing technologies and improve the accuracy and efficiency of green space classification.

[0036] The following combines Figures 1 - 5 Describe a high-resolution urban green space classification method, device, equipment and medium based on a dual decoder and a fine segmentation head of the present invention, to solve the defect of low accuracy of traditional image classification technology and current deep learning network image classification in the prior art. Through a classification algorithm that simultaneously captures local and global information, it realizes rapid processing of large-scale high-resolution satellite urban image data, automatically classifies urban green space images with high precision, and significantly improves the accuracy and efficiency of green space classification.

[0037] Figure 1 It is one of the flow schematic diagrams of a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention, asFigure 1 As shown, the method may include but is not limited to steps S100 to S300: S100, obtaining high-resolution satellite remote sensing images; S200, preprocessing the satellite remote sensing images to obtain image training samples and image test samples; S300, inputting the image test samples into a trained green space classification neural network for processing to determine the green space classification result; wherein, the green space classification neural network is obtained by training the image training samples.

[0038] In step S100 of some embodiments, high-resolution satellite remote sensing images are obtained.

[0039] It can be understood that high-resolution satellite remote sensing images refer to high-precision and high-clarity images of the urban ground obtained by high-resolution sensors carried by satellites.

[0040] In step S200 of some embodiments, the satellite remote sensing images are preprocessed to obtain image training samples and image test samples.

[0041] It can be understood that after step S100 is executed, the specific execution steps may be: performing per-pixel annotation on the green space features of the satellite remote sensing images to obtain an annotated image; performing same-scale segmentation processing on the annotated image and the satellite remote sensing images to obtain annotated image samples and original image samples, and using the annotated image samples as the image test samples; performing cropping processing on the original image samples to obtain first original sub-image training samples and second original sub-image training samples, and using the first original sub-image training samples and the second original sub-image training samples as the image training samples.

[0042] Further, using satellite image processing software such as Arcgis, the number of categories of green space features is determined in advance by means such as visual interpretation or on-site survey. For example, in an area covered by satellite remote sensing images, through on-site investigation and visual interpretation, it is determined that the green space features in this area include different types such as forests, grasslands, and wetlands.

[0043] According to the determined categories, the green space features in the remote sensing images are divided into different colors. For example, the forest area is represented by green, the grassland area is represented by yellow, the wetland area is represented by blue, etc.

[0044] Save the annotated image in png format, which can clearly store the color information of each pixel and is convenient for subsequent processing and analysis.

[0045] Pixel-by-pixel annotation can provide accurate information on land cover classes, enabling the computer to accurately identify and distinguish different green space features. This provides the basic data for subsequent classification and analysis, helping to improve the accuracy of green space extraction and monitoring. The original image and the annotated image are segmented into sub-images according to the same length and width (such as 2048 2048 pixels), that is, the annotated image samples and the original image samples are obtained, ensuring the corresponding relationship between the two in terms of spatial position and size.

[0046] Furthermore, the segmented sub-images are respectively cropped into sub-image a (256 256 pixels) and sub-image b (768 768 pixels), where sub-image b is centered on sub-image a and the cropping size is three times that of sub-image a. Such a cropping method can obtain image information at different scales. Sub-image a can provide local detailed information, while sub-image b contains larger context information.

[0047] It should be noted that sub-image a is the first original sub-image training sample, and sub-image b is the second original sub-image training sample.

[0048] Through the same-scale segmentation process, the consistency of the original image and the annotated image in space can be ensured, avoiding analysis errors caused by inconsistent image sizes or misalignment. By using the annotated image as a test sample, the classification accuracy of the model for green space features can be tested in actual application scenarios, providing a basis for the improvement and optimization of the model.

[0049] The original image samples are cropped to obtain the first original sub-image training sample and the second original sub-image training sample. During the cropping process, appropriate cropping positions and sizes can be selected according to specific research requirements and data characteristics.

[0050] The first original sub-image training sample and the second original sub-image training sample are used as image training samples to train the green space classification neural network model. During the training process, the green space classification neural network model will learn the features and patterns of different features, thereby improving the ability to identify green space features.

[0051] Through cropping, representative sub-image samples can be obtained from the original image, reducing the data volume while retaining key information and improving the efficiency of model training. By using these sub-images as training samples, the model can learn the features of green space features at different scales and in different environments, improving the generalization ability and robustness of the model, enabling it to better adapt to various changes and challenges in actual application scenarios.

[0052] In step S300 of some embodiments, the image test sample is input into the trained green space classification neural network for processing to determine the green space classification result; wherein, the green space classification neural network is obtained by training the image training samples.

[0053] It should be noted that inputting the image test sample into the trained green space classification neural network for processing means starting the forward propagation process of the green space classification neural network, and the image test sample is sequentially calculated through each layer of the network. In this process, the network will perform feature extraction and classification decision on each pixel or data block according to the learned weight and bias parameters, that is, the model parameters of the trained green space classification neural network. For each pixel or data block, the network will output a probability vector, and the length of the vector is equal to the number of green space categories. Each element in the vector represents the probability that the pixel belongs to the corresponding category. The category corresponding to the highest probability in each pixel probability vector is selected as the prediction result of the network, so as to obtain the prediction classification map of the entire test set.

[0054] The prediction result is compared with the corresponding labeled image of the test set pixel by pixel. The labeled image is reference data with known correct categories. By comparing with the labeled image, the correctness and error conditions of the network prediction can be determined.

[0055] A confusion matrix is constructed to summarize the relationship between the prediction result and the true label. The confusion matrix is a square matrix, the rows and columns of which represent the true category and the predicted category respectively, and the elements in the matrix represent the number of pixels corresponding to the true category and the predicted category.

[0056] Evaluation indicators such as accuracy and IoU (Intersection over Union) are calculated according to the confusion matrix. Accuracy refers to the ratio of the number of correctly predicted pixels to the total number of pixels, and IoU comprehensively considers the cases of true positives and false positives, and can more accurately reflect the segmentation performance of the network.

[0057] Furthermore, it should be noted that as shown in Figure 2 the green space classification neural network includes: an internal encoder, an external encoder, a decoder, and a fine segmentation head.

[0058] The internal encoder is mainly responsible for feature extraction and downsampling operations on the input high-resolution green space image. It usually consists of a series of convolutional layers, pooling layers, and possibly non-linear activation functions (such as ReLu). The convolutional layers are used to learn local features from the image, such as edges, textures, etc. These local features are crucial for identifying different land cover types (such as vegetation, water bodies, bare land, etc.). The pooling layers gradually reduce the spatial dimension of the feature map, reducing the data volume while retaining important feature information. This process also helps the model focus on the intrinsic structure and patterns of the image rather than just paying attention to details. For example, when processing green space images, the internal encoder can capture features such as the leaf edges and textures of vegetation through convolutional layers, and gradually generalize these features through pooling layers to provide a compact and representative low-resolution feature representation for subsequent processing.

[0059] The functions of the internal encoder at least include: Downsampling and feature simplification: The downsampling operation of the internal encoder gradually reduces the spatial size of the input image, for example, from the original high-resolution image (such as 256×256 pixels) to a lower resolution (such as 32×32 pixels or lower). This downsampling process not only reduces the computational amount but also helps the model remove some unnecessary detailed information and highlight the main structure and features of the image. This is very important for the green space classification task because it can focus on the features of larger areas, such as the overall features of different types of green space plots (large parks, small gardens, etc.), rather than the details of individual pixels, thus better identifying and classifying different green space categories.

[0060] Feature vectorization: Through a series of convolutional and pooling operations, the internal encoder converts the high-resolution image features into a low-resolution vector representation. This vector contains the important feature information after the input image has undergone multiple layers of processing and is a compact and abstract representation of the original image. For example, in green space images, the internal encoder can map regions with similar vegetation cover types, texture features, etc. to similar vectors so that subsequent classifiers can more easily distinguish different categories.

[0061] The external encoder: It is mainly used to process additional data sources that contain semantic information related to green spaces but may not have high-resolution spatial information. For example, it can process data such as land use type data, vegetation index data (such as the normalized difference vegetation index NDVI), and terrain data (elevation, slope, etc.). The external encoder usually processes and encodes these non-image data separately so that they can be effectively fused with the image features processed by the internal encoder. For example, for vegetation index data, the external encoder may convert it into a representation form that matches the image features through specific mathematical transformations or neural network layers.

[0062] Resblock is short for residual block, which is a commonly used module structure in deep learning. Its basic structure is as follows: The input features are subjected to feature extraction through a convolutional layer. After passing through an activation function, they are then subjected to feature extraction through another convolutional layer. The output of the second convolutional layer is added to the input features. Finally, after passing through another activation function, the output of the ResBlock is obtained.

[0063] By stacking multiple ResBlocks, a deep residual network such as ResNet can be constructed. ResNet has achieved good results in tasks such as image classification and object detection.

[0064] The role of the external encoder: It provides supplementary semantic information for the entire neural network and realizes the fusion of multi-modal data. In green space classification, relying solely on the spatial features of images may not be able to accurately distinguish some categories with similar appearances but different semantics. For example, a grassland and a crop planting area may look similar in an image, but by incorporating external data such as vegetation indices, the model can better understand the differences between them. The external encoder combines this additional information with the image features of the internal encoder, enriching the model's understanding of green spaces and improving the accuracy and robustness of classification.

[0065] Enhance the model's understanding of complex scenes: By processing various types of data, the external encoder can help the model better understand the complex environment in which the green space is located. For example, combining terrain data can help the model distinguish between a forest in a valley and farmland on a plain, because terrain differences are often closely related to land use patterns and vegetation distributions. This fusion of multi-source data enables the model to evaluate green spaces from multiple perspectives, thus completing the classification task more comprehensively.

[0066] Decoder: Feature Restoration and Upsampling: The main task of the decoder is to gradually restore the low-resolution features output by the internal encoder to high resolution to generate an output with the same or similar size as the input image. It typically consists of a series of transposed convolutional layers (also known as deconvolutional layers or transposed convolutional layers), skip connections, and possibly non-linear activation functions. The transposed convolutional layers are used to upsample the low-resolution features, increasing the spatial dimensions of the feature map by inserting zero values, etc. The skip connections fuse the features at the corresponding levels in the internal encoder with the features in the decoder, and this fusion can provide more detailed information to help restore the spatial structure of the image. For example, during the processing, the decoder first magnifies the small-sized feature map output by the internal encoder by a certain multiple through the transposed convolutional layer, and then concatenates it with the output of the corresponding layer in the internal encoder (the larger-sized feature map before the pooling operation in this layer), thereby gradually constructing a higher-resolution feature representation.

[0067] Transformer Fusion Incorporating the self-attention mechanism of Transformer, features are fused at different levels to improve the model's ability to represent features and generalization ability.

[0068] The functions of the decoder at least include: In the case of green space classification for semantic segmentation tasks, the decoder plays a crucial role. It restores the low-resolution classification features to high resolution, enabling each pixel to be assigned the corresponding class label. For example, in the semantic segmentation of urban green space landscapes, the decoder can restore the features processed by the internal encoder and the external encoder to the size of the original image, thereby accurately dividing different classes such as vegetation areas, water areas, and road areas. This process of restoring features from low resolution to high resolution is like a "refinement" process, enabling the model to more accurately locate and identify the class to which each pixel in the image belongs.

[0069] In classification tasks, the decoder helps the model better focus on local details in the image by restoring spatial information, thereby improving the accuracy of classification. For example, when differentiating different types of vegetation, the high-resolution features restored by the decoder can clearly show information such as the morphology and boundaries of the vegetation, enabling the model to more accurately distinguish between broad-leaved forests and coniferous forests, single-species vegetation and mixed vegetation, etc. At the same time, the decoder can also combine the semantic information provided by the external encoder to further optimize the classification results, because the features at this time have fused information from multiple aspects and can more comprehensively reflect the real situation of the green space.

[0070] The fine segmentation head is a structure specifically designed to process high-resolution image features and is usually located at the end of the decoder. It is mainly used to further refine the features after upsampling by the decoder to generate the final fine segmentation result. The fine segmentation head generally includes convolutional layers with small convolutional kernels (such as 3×3 or 1×1) and possibly batch normalization layers, ReLu activation functions, etc. These small convolutional kernels can better capture the tiny details and local features in the image and classify each pixel and its neighborhood more precisely. For example, when processing the vegetation boundary in a green space image, the fine segmentation head can accurately identify the position and shape of the boundary through small convolutional kernels.

[0071] The main function of the fine segmentation head is to generate high-quality fine segmentation images. In green space classification, this means being able to accurately divide each pixel into different categories, such as different types of vegetation, water bodies, bare land, etc. It makes the segmentation result more accurate and smooth through the fine processing of high-resolution features. For example, when identifying flower beds in an urban park, the fine segmentation head can accurately outline the contours of the flower beds and distinguish the planting areas of different types of flowers, providing detailed information for applications such as urban planning and ecological environment monitoring.

[0072] In complex green space scenes, there are various types of ground objects and complex boundary situations. The fine segmentation head can adapt to this complexity and effectively handle the interlacing and transition between different ground objects through the fine extraction and classification of local features. For example, at the junction of a river and riverside vegetation, the fine segmentation head can use high-resolution features to accurately determine whether each pixel belongs to the water body or the vegetation, avoiding misclassification, thereby improving the reliability and accuracy of the entire green space classification.

[0073] The training method of the green space classification neural network includes: Inputting the first original subgraph training sample into the internal encoder for downsampling to obtain local image features; Inputting the second original subgraph training sample into the external encoder for downsampling to obtain global image features; Inputting the local image features and the global image features into the decoder for feature fusion processing to obtain fused image features; Inputting the fused image features into the fine segmentation head for feature enhancement processing to obtain target image features; Inputting the target image features and the image test sample into a preset training evaluation model for evaluation processing, and taking the target image features with an evaluation score greater than or equal to the preset score threshold as the training output result of the trained green space classification neural network; Take the training parameters corresponding to the training output result as the model parameters of the trained green space classification neural network.

[0074] It can be understood that sub - figure a (local image): Crop a small - sized sub - figure a from the original high - resolution remote - sensing image, that is, the first original sub - figure training sample, for extracting local features.

[0075] Sub - figure b (global image): Centered on sub - figure a, crop a large sub - figure b on the original image with a size three times that of sub - figure a, that is, the second original sub - figure training sample, for capturing a wider range of global context information.

[0076] Internal encoder processing: Input sub - figure a (the first original sub - figure training sample) into the internal encoder.

[0077] The internal encoder can use the ResNet18 network structure. Through four down - sampling operations (sampling rates are 1 / 2, 1 / 4, 1 / 8, 1 / 16 respectively), the feature dimensions are 64, 128, 256, 512 in sequence after each down - sampling. Stack four down - sampling operations and four ResNet18 modules. The purpose is to extract multi - scale local image internal feature maps.

[0078] External encoder processing: Input sub - figure b (the second original sub - figure training sample) into the external encoder.

[0079] The external encoder uses the CNN network structure and also performs four down - sampling operations (sampling rates are 1 / 2, 1 / 4, 1 / 8, 1 / 16 respectively), and the feature dimensions are also 64, 128, 256, 512.

[0080] The purpose is to capture a wider range of global context information and long - distance dependencies between images.

[0081] By inputting the second original sub - figure training sample into the external encoder for down - sampling processing, global image features are obtained.

[0082] In some embodiments of the present invention, the global image features obtained from the output of the external encoder are reshaped and up - sampled to generate additional external prediction information. This prediction is used for the calculation of the loss function to supervise network training.

[0083] The overall loss function for training the green space classification neural network in the present invention , consists of the Dice loss function and the cross - entropy loss function and is shown as follows: Among them, and represent the number of samples and the number of categories respectively; and are the true labels of sub - figure a and sub - figure b; and are the internal prediction information and the external additional prediction information. The loss of the external prediction is multiplied by a coefficient (default is 0.5) to better integrate the loss function and balance the importance of internal and external predictions. The internal prediction information is obtained by reshaping and upsampling the local image features output by the internal encoder. The Dice loss measures the similarity between the prediction result and the true label.

[0084] Through the above double - encoder design, the internal encoder focuses on capturing fine - grained local features, while the external encoder focuses on obtaining global information. The two are highly complementary and provide a comprehensive feature representation for subsequent classification reasoning. The steps of inputting the local image features and the global image features into the decoder for feature fusion processing specifically include: Generating a local feature vector from the local image features through a multi - layer perceptron, and generating a first global feature vector and a second global feature vector from the global image features through a multi - layer perceptron; Performing a calculation process on the first global feature vector, the second global feature vector and the local feature vector based on an activation function to generate an attention weight matrix; Performing a product operation on the attention weight matrix and the decoder features of the decoder to generate the fused image features.

[0085] It should be noted that the present invention uses a Transformer feature fusion module to layer - by - layer integrate the features from the internal encoder and the external encoder. The local image features from the internal encoder pass through a multi - layer perceptron to generate a q - vector (local feature vector), and the global image features from the external encoder pass through a multi - layer perceptron to generate k and v vectors (the first global feature vector and the second global feature vector), then generate an attention weight matrix through the softmax function, and then perform a product operation on the attention weight matrix and the decoder features of the decoder, thereby completing the operation of fusing the internal and external decoder features. As shown in the following formula: Among them, is the number of channels of the current feature map, is the number of heads in the multi - head attention mechanism, the k and v vectors correspond to the first global feature vector and the second global feature vector respectively, and the q - vector is the local feature vector, To fuse image features.

[0086] In this way, the green space classification neural network model can effectively model the context dependence between the in-graph feature map mm and the cross-image feature map momo, so as to more effectively project the cross-image features into the in-graph context relationship. Then, the image resolution is gradually restored through upsampling.

[0087] By inputting local image features into a multi-layer perceptron, the non-linear representation of local image features can be automatically learned. This non-linear representation can better capture the complex structures and patterns in local images, providing more representative local feature information for subsequent feature fusion.

[0088] For global image features, two different first global feature vectors and second global feature vectors are generated through two different multi-layer perceptrons respectively, which can abstract and represent global image features from different perspectives. In this way, the global image features can be modeled from multiple different subspaces by using the multi-head attention mechanism, so as to more comprehensively capture the context information in the global image.

[0089] Calculating and processing the three feature variables based on the activation function is the core part of the attention mechanism. The activation function (such as Softmax) can convert the feature values after linear combination into an attention weight matrix in the form of a probability distribution. This attention weight matrix can reflect the correlation between local features and global features, and determines the importance degree of each feature during information fusion.

[0090] For example, when using the Softmax activation function, the element values in the attention weight matrix range from 0 to 1, and the sum of all elements is 1. A larger weight value indicates that the corresponding feature has a higher priority during the fusion process, thus highlighting important feature information and suppressing unimportant feature information.

[0091] Furthermore, the product operation is performed on the attention weight matrix and the decoder features of the decoder, realizing the effective fusion of local image features and global image features at the decoder level. In this way, the model can integrate the context information of the global image into the local image features.

[0092] Specifically, each weight value in the attention weight matrix corresponds to a correlation relationship between local features and global features. During the product operation, the global features with higher weights can have a greater impact on local features, so that the fused image features not only contain local detailed information but also integrate global context information. This fused feature is of great significance for subsequent tasks such as image classification and object detection, and can improve the model's ability to understand and recognize images.

[0093] In some embodiments of the present invention, the step of inputting the fused image features into the fine segmentation head for feature enhancement processing to obtain target image features includes: Inputting the fused image features into the fine segmentation head for channel feature enhancement processing to obtain channel-enhanced features; Performing spatial feature enhancement processing on the channel-enhanced features to obtain the target image features.

[0094] It can be understood that as Figure 3 shown, the step of inputting the fused image features into the fine segmentation head for channel feature enhancement processing to obtain channel-enhanced features specifically includes: Using the fine segmentation head to perform global max pooling processing on the fused image features to obtain a global max feature vector, and performing global average pooling processing on the fused image features to obtain a global average feature vector; Based on the multi-layer perceptron of the fine segmentation head, processing the global max feature vector and the global average feature vector to generate a channel attention map; Applying the channel attention map to the fused image features to generate the channel-enhanced features.

[0095] The fine segmentation head is used to provide more classification bases for the network, better enhance the channel and spatial representations of the fused features, ensure that the key green space information is amplified, and at the same time suppress redundant or irrelevant features. In the embodiments of the present invention, the fine segmentation head is used to generate the prediction information of the green space classification neural network.

[0096] Specifically as follows: First, the features from the decoder (i.e., the fused image features) are subjected to global max pooling and global average pooling processing, respectively generating two global feature vectors (i.e., the global max feature vector) and (i.e., the global average feature vector), as shown in the following formula: (1) (2) where, represents the global max pooling operation, represents the global average pooling operation.

[0097] Passing the two vectors, i.e., the global max feature vector and the global average feature vector, respectively through a shared multi-layer perceptron (MLP) to generate a channel attention map .

[0098] (3) Among them, is the Sigmoid function, which is used to generate attention scores. is the global maximum feature vector. is the global average feature vector.

[0099] Apply the attention map to the channel dimension of the fused image features to generate channel-enhanced features , which can be expressed by the following formula: (4) Among them, is the channel-enhanced feature, is the channel attention map, is the fused image feature.

[0100] Global max pooling is a commonly used pooling operation method. It can extract the most important feature information from the fused image features. By sliding a fixed-size window on the feature map and taking the maximum value in each window as the output, the global maximum feature vector is obtained.

[0101] This operation helps to highlight the most critical local features in the fused image features. For example, in an image classification task, it may highlight the most prominent part of the target object in the image and ignore irrelevant information such as the background. At the same time, max pooling can effectively reduce the spatial dimension of the feature map, reduce the computational amount, and improve the operation efficiency of the model.

[0102] Global average pooling is another important pooling method. It calculates the average value of each channel of the fused image features to obtain the global average feature vector. Different from max pooling, average pooling pays more attention to retaining the overall feature information.

[0103] It can avoid the influence of individual extreme values (such as noise points) in the feature map on the entire feature extraction process, making the model pay more attention to the overall distribution of features. For example, when dealing with image texture features, average pooling can better capture the average pattern of the texture and provide a more stable basis for subsequent feature processing. The role of generating the channel attention map: Generating the channel attention map by processing the global maximum feature vector and the global average feature vector based on the multi-layer perceptron of the fine segmentation head is to dynamically adjust the importance of different channel features.

[0104] The multi-layer perceptron can learn the complex non-linear relationship between the global maximum feature vector and the global average feature vector. By comprehensively processing these two feature vectors, the model can automatically judge which channel features are more important for the current task (such as classification, segmentation, etc.) according to the image content.

[0105] The generated channel attention map is actually a weight assignment mechanism that can enhance the expression of important channel features while suppressing unimportant channel features. In this way, when the channel attention map is subsequently applied to fuse image features, it can highlight key information, improve the model's sensitivity to this key information, and thus enhance the model's performance.

[0106] Function of generating channel enhanced features: The purpose of applying the channel attention map to fuse image features to generate channel enhanced features is to further optimize and refine the fused image features.

[0107] As a weight information, multiplying the channel attention map with the fused image features can adjust each channel in the fused image features according to the corresponding weights in the attention map. Those channel features assigned higher weights will be enhanced in the final channel enhanced features, while those with lower weights will be relatively weakened.

[0108] This operation helps the model to focus more on the key information related to the task and improve the model's representation ability of image features. In practical applications, such as in the high-resolution green space image classification task, the channel enhanced features can better highlight the features related to the green space, enabling the model to more accurately distinguish different green space categories.

[0109] In some embodiments of the present invention, the step of performing spatial feature enhancement processing on the channel enhanced features to obtain the target image features specifically includes: Performing global max pooling processing on the channel enhanced features to obtain a global max feature map, and performing global average pooling processing on the channel enhanced features to obtain a global average feature map; Performing convolution processing on the global max feature map and the global average feature map to generate a spatial attention map; Applying the spatial attention map to the channel enhanced features to generate a spatial channel enhanced feature; Performing prediction processing on the spatial channel enhanced features to obtain the target image features; wherein the prediction processing includes at least a single ReLu activation function, a single convolutional layer, and an upsampling operation.

[0110] It can be understood that, as shown in Figure 3 performing global max pooling and global average pooling operations on the channel dimension of the channel enhanced features to generate two spatial feature maps (global max feature map) and (global average feature map).

[0111] ​Perform convolution processing on the global maximum feature map and the global average feature map to generate a spatial attention map, that is, stack the global maximum feature map and the global average feature map, these two spatial feature maps, and generate a spatial attention map through a convolution operation , as shown in the following formula: (5) Among them, is the spatial attention map, [·] represents the concatenation of feature maps, is the global maximum feature map, is the global average feature map.

[0112] Apply the spatial attention map to the spatial dimension of the channel-enhanced feature to generate a spatially and channel-optimized spatial-channel enhanced feature .

[0113] (6) Among them, is the spatial-channel enhanced feature, is the spatial attention map, is the channel-enhanced feature.

[0114] Finally, after passing the spatial-channel enhanced feature through a ReLu activation function, a convolutional layer, and an upsampling operation, the final prediction result of the green spatial classification neural network is obtained, that is, the target image feature.

[0115] Finally, input the target image feature and the image test sample into a preset training and evaluation model for evaluation processing, and use the target image feature with an evaluation score greater than or equal to the preset score threshold as the training output result of the trained green spatial classification neural network. Use the training parameters corresponding to the training output result as the model parameters of the trained green spatial classification neural network.

[0116] Train a high-resolution urban green space classification network based on a dual decoder and a fine segmentation head through a training set, obtain a weight parameter file and save it.

[0117] Function of generating spatial attention map: Convolving the global maximum feature map and the global average feature map to generate the spatial attention map is to integrate the information of the two pooling features and determine the importance of each position in the image. The convolution operation can learn the complex non-linear relationship between these two feature maps. Through convolution processing, the model can automatically judge which features at which spatial positions are more important for the current task (such as classification, segmentation, etc.) according to the image content. For example, in a green space image, the spatial attention map can highlight key positions such as the edges of vegetation, the boundaries between water bodies and vegetation, making the model pay more attention to these areas that help distinguish different land cover classes.

[0118] Function of generating spatially-channel enhanced features: Applying the spatial attention map to the channel enhanced features to generate spatially-channel enhanced features aims to optimize the channel enhanced features based on spatial positions. As a weight information, multiplying the spatial attention map with the channel enhanced features can adjust each channel in the channel enhanced features according to the corresponding weights in the attention map at different spatial positions. The features at the spatial positions assigned higher weights will be enhanced in the final spatially-channel enhanced features, while the features at the positions with lower weights will be relatively weakened. This operation helps the model to focus more on the key spatial position information related to the task in the image and improves the model's ability to represent image features. In practical applications, such as in the classification task of high-resolution green space images, the spatially-channel enhanced features can better highlight the spatial features related to the category, enabling the model to more accurately distinguish different green space categories.

[0119] Function of the single ReLu activation function: The ReLu (Rectified Linear Unit) activation function is a commonly used non-linear activation function. It is computationally simple and can quickly judge whether the input is greater than 0. If it is greater than 0, the output is equal to the input; otherwise, the output is 0. When performing prediction processing on the spatially-channel enhanced features, the ReLu activation function introduces non-linearity and enhances the model's expressive ability. It can convert the linear relationship in the spatially-channel enhanced features into a non-linear relationship, enabling the model to learn more complex patterns. For example, when processing green space images, the ReLu activation function can help the model better capture the complex relationships between factors such as vegetation growth status, lighting conditions and image features, thus more accurately predicting the target image features.

[0120] Function of a single convolutional layer: A single convolutional layer performs a convolution operation on the features after ReLu activation to further extract high-level features in the spatially channel-enhanced features. The convolutional layer slides a convolutional kernel over the feature map to detect local patterns such as edges and textures. During this process, the convolutional layer can learn complex relationships between different channels and spatial positions, and combine these relationships into more abstract feature representations. For example, in a green space image, the convolutional layer can identify high-level features such as the morphological structures and distribution patterns of different types of vegetation, providing richer feature information for subsequent classification or segmentation tasks.

[0121] Function of the upsampling operation: The upsampling operation is to restore the spatial resolution, map the low-resolution features back to high resolution, and obtain an output image with the same or similar size as the input image. When predicting the features of the target image, due to a series of previous operations such as pooling and convolution, the size of the feature map decreases, and upsampling can restore the processed features to the size of the original image. This can ensure that the prediction results have sufficient spatial details and can accurately correspond to each pixel position of the input image. For example, in the green space image classification task, the upsampled prediction results can accurately label the category to which each pixel belongs, such as vegetation-covered areas and bare areas, providing detailed information for subsequent analysis and evaluation.

[0122] The high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the present invention has at least the following technical effects: (1) The dual encoder structure adopted by the present invention combines local feature extraction and global context modeling. The internal encoder extracts fine-grained boundary and texture information, and the external encoder captures long-range dependencies and global features across the image. Through the collaborative action of the dual encoders, the problem of incomplete green space classification caused by limited receptive fields in the prior art is effectively solved, and the classification accuracy is significantly improved, especially in complex scenarios (such as fragmented green spaces and mixed-type areas).

[0123] (2) Compared with the high computational complexity of traditional global attention mechanisms (such as pure Transformer structures), the decoder of the present invention that embeds a feature fusion module based on a transformer significantly reduces the computational overhead, making the method more efficient in processing high-resolution remote sensing images, while maintaining the global feature modeling ability, and is suitable for the application requirements of large-scale urban green space scenarios.

[0124] (3) The present invention provides an efficient and reliable technical means for urban green space classification, which can be widely applied in fields such as urban ecological planning, environmental management, and land use assessment. Through more accurate classification results, the present invention helps to promote the rational allocation of urban green space resources and the maximization of ecological benefits.

[0125] A high-resolution urban green space classification method, device, equipment and medium based on a dual decoder and a fine segmentation head provided by the present invention obtains a high-resolution satellite remote sensing image; preprocesses the satellite remote sensing image to obtain an image training sample and an image test sample; inputs the image test sample into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is obtained by training the image training sample. It is used to solve the defect of low accuracy of traditional image classification technology and current deep learning network image classification in the prior art. Through a classification algorithm that captures local and global information simultaneously, it can quickly process large-scale high-resolution satellite urban image data, automatically classify urban green space images with high precision, and significantly improve the accuracy and efficiency of green space classification.

[0126] The high-resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention is described below. The high-resolution urban green space classification device based on a dual decoder and a fine segmentation head described below can be mutually referred to corresponding to the high-resolution urban green space classification method described above.

[0127] As Figure 4 shown is a schematic structural diagram of a high-resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention. A high-resolution urban green space classification device based on a dual decoder and a fine segmentation head includes the following modules: An acquisition module 410, configured to acquire a high-resolution satellite remote sensing image; A preprocessing module 420, configured to preprocess the satellite remote sensing image to obtain an image training sample and an image test sample; A module 430 for determining a green space classification result, configured to input the image test sample into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is obtained by training the image training sample.

[0128] Preferably, for the high-resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention, the preprocessing module 420 is specifically configured to perform per-pixel annotation on the green space features of the satellite remote sensing image to obtain an annotated image; perform proportional segmentation processing on the annotated image and the satellite remote sensing image to obtain an annotated image sample and an original image sample, and use the annotated image sample as the image test sample; Crop the original image sample to obtain a first original sub - image training sample and a second original sub - image training sample, and use the first original sub - image training sample and the second original sub - image training sample as the image training sample.

[0129] Preferably, for the high - resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention, the module 430 for determining the green space classification result is specifically configured such that the green space classification neural network includes: an internal encoder, an external encoder, a decoder, and a fine segmentation head; The training method of the green space classification neural network includes: Input the first original sub - image training sample into the internal encoder for down - sampling processing to obtain local image features; Input the second original sub - image training sample into the external encoder for down - sampling processing to obtain global image features; Input the local image features and the global image features into the decoder for feature fusion processing to obtain fused image features; Input the fused image features into the fine segmentation head for feature enhancement processing to obtain target image features; Input the target image features and the image test sample into a preset training evaluation model for evaluation processing, and use the target image features with an evaluation score greater than or equal to a preset score threshold as the training output result of the trained green space classification neural network; Use the training parameters corresponding to the training output result as the model parameters of the trained green space classification neural network.

[0130] Preferably, for the high - resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention, the module 430 for determining the green space classification result is specifically configured to Generate local feature vectors from the local image features through a multi - layer perceptron, and generate a first global feature vector and a second global feature vector from the global image features through a multi - layer perceptron; Perform calculation processing on the first global feature vector, the second global feature vector, and the local feature vector based on an activation function to generate an attention weight matrix; Perform a product operation on the attention weight matrix and the decoder features of the decoder to generate the fused image features.

[0131] Preferably, for the high - resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention, the module 430 for determining the green space classification result is specifically configured to input the fused image features into the fine segmentation head for channel feature enhancement processing to obtain channel - enhanced features; Perform spatial feature enhancement processing on the channel enhancement feature to obtain the target image feature.

[0132] Preferably, for the high-resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention, the module 430 for determining the green space classification result is specifically configured to perform global max pooling processing on the fused image feature by using the fine segmentation head to obtain a global max feature vector, and perform global average pooling processing on the fused image feature to obtain a global average feature vector; Process the global max feature vector and the global average feature vector based on the multi-layer perceptron of the fine segmentation head to generate a channel attention map; Apply the channel attention map to the fused image feature to generate the channel enhancement feature.

[0133] Preferably, for the high-resolution urban green space classification device based on a dual decoder and a fine segmentation head provided by the present invention, the module 430 for determining the green space classification result is specifically configured to perform global max pooling processing on the channel enhancement feature to obtain a global max feature map, and perform global average pooling processing on the channel enhancement feature to obtain a global average feature map; Perform convolution processing on the global max feature map and the global average feature map to generate a spatial attention map; Apply the spatial attention map to the channel enhancement feature to generate a spatial channel enhancement feature; Perform prediction processing on the spatial channel enhancement feature to obtain the target image feature; wherein the prediction processing at least includes a single ReLu activation function, a single convolutional layer, and upsampling operation processing.

[0134] A high-resolution urban green space classification method, device, equipment, and medium based on a dual decoder and a fine segmentation head provided by the present invention obtain a high-resolution satellite remote sensing image; perform preprocessing on the satellite remote sensing image to obtain an image training sample and an image test sample; input the image test sample into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is trained on the image training sample. It is used to solve the defect of low accuracy in traditional image classification technology and current deep learning network image classification in the prior art. Through a classification algorithm that captures both local and global information simultaneously, it realizes fast processing of large-scale high-resolution satellite urban image data, automatically classifies urban green space images with high accuracy, and significantly improves the accuracy and efficiency of green space classification.

[0135] Figure 5 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head. The method includes: obtaining a high-resolution satellite remote sensing image; preprocessing the satellite remote sensing image to obtain image training samples and image test samples; inputting the image test samples into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is obtained by training the image training samples.

[0136] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0137] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a high-resolution urban green space classification method provided by the above-mentioned various methods. The method includes: obtaining a high-resolution satellite remote sensing image; preprocessing the satellite remote sensing image to obtain image training samples and image test samples; inputting the image test samples into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is obtained by training the image training samples.

[0138] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a high-resolution urban green space classification method based on a dual decoder and a fine segmentation head provided by the above-mentioned various methods. The method includes: obtaining a high-resolution satellite remote sensing image; preprocessing the satellite remote sensing image to obtain image training samples and image test samples; inputting the image test samples into a trained green space classification neural network for processing to determine a green space classification result; wherein, the green space classification neural network is obtained by training the image training samples.

[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A high-resolution urban green space classification method based on a dual decoder and a fine segmentation head, characterized in that Including: Obtain high-resolution satellite remote sensing images; Preprocess the satellite remote sensing images to obtain image training samples and image test samples; Input the image test samples into a trained green space classification neural network for processing to determine the green space classification results; wherein, the green space classification neural network is trained with the image training samples.

2. The high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to claim 1, wherein The preprocessing of the satellite remote sensing images to obtain image training samples and image test samples includes: Perform pixel-by-pixel annotation on the green space features of the satellite remote sensing images to obtain an annotated image; Perform proportional segmentation processing on the annotated image and the satellite remote sensing images to obtain annotated image samples and original image samples, and use the annotated image samples as the image test samples; Perform cropping processing on the original image samples to obtain first original sub-image training samples and second original sub-image training samples, and use the first original sub-image training samples and the second original sub-image training samples as the image training samples.

3. The high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to claim 2, characterized in that, The green space classification neural network includes: an internal encoder, an external encoder, a decoder, and a fine segmentation head; The training method of the green space classification neural network includes: Input the first original sub-image training samples into the internal encoder for downsampling processing to obtain local image features; Input the second original sub-image training samples into the external encoder for downsampling processing to obtain global image features; Input the local image features and the global image features into the decoder for feature fusion processing to obtain fused image features; Input the fused image features into the fine segmentation head for feature enhancement processing to obtain target image features; Input the target image features and the image test samples into a preset training evaluation model for evaluation processing, and use the target image features with evaluation scores greater than or equal to a preset score threshold as the training output results of the trained green space classification neural network; Use the training parameters corresponding to the training output results as the model parameters of the trained green space classification neural network.

4. The high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to claim 3, wherein The inputting the local image features and the global image features into the decoder for feature fusion processing to obtain fused image features includes: Generate local feature vectors from the local image features through a multi-layer perceptron, and generate first global feature vectors and second global feature vectors from the global image features through a multi-layer perceptron; Perform calculation processing on the first global feature vectors, the second global feature vectors, and the local feature vectors based on an activation function to generate an attention weight matrix; Perform product operation processing on the attention weight matrix and the decoder features of the decoder to generate the fused image features.

5. The high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to claim 3, characterized in that The inputting the fused image features into the fine segmentation head for feature enhancement processing to obtain target image features includes: Input the fused image features into the fine segmentation head for channel feature enhancement processing to obtain channel-enhanced features; Perform spatial feature enhancement processing on the channel-enhanced features to obtain the target image features.

6. The high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to claim 5, wherein, Inputting the fused image features into the fine segmentation head for channel feature enhancement processing to obtain channel-enhanced features includes: Using the fine segmentation head to perform global max pooling on the fused image features to obtain a global max feature vector, and performing global average pooling on the fused image features to obtain a global average feature vector; Processing the global max feature vector and the global average feature vector based on the multi-layer perceptron of the fine segmentation head to generate a channel attention map; Applying the channel attention map to the fused image features to generate the channel-enhanced features.

7. The high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to claim 5, characterized in that, Performing spatial feature enhancement processing on the channel-enhanced features to obtain the target image features includes: Performing global max pooling on the channel-enhanced features to obtain a global max feature map, and performing global average pooling on the channel-enhanced features to obtain a global average feature map; Performing convolution processing on the global max feature map and the global average feature map to generate a spatial attention map; Applying the spatial attention map to the channel-enhanced features to generate spatially channel-enhanced features; Performing prediction processing on the spatially channel-enhanced features to obtain the target image features; wherein the prediction processing at least includes a single ReLu activation function, a single convolutional layer, and upsampling operation processing.

8. A high-resolution urban green space classification device based on a dual decoder and a fine segmentation head, characterized in that, Including: An acquisition module for acquiring high-resolution satellite remote sensing images; A preprocessing module for preprocessing the satellite remote sensing images to obtain image training samples and image test samples; A green space classification result determination module for inputting the image test samples into a trained green space classification neural network for processing to determine the green space classification result; wherein the green space classification neural network is trained with the image training samples.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the program, it implements the high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the high-resolution urban green space classification method based on a dual decoder and a fine segmentation head according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • High-resolution remote sensing image land coverage classification method and device and storage medium

    CN117036936A

  • Hyperspectral and laser radar data fusion classification method based on AGLT network

    CN117475216A

  • Sea-land port segmentation method based on space and semantic alignment fusion

    CN118691827A

  • High-resolution remote sensing image building function type classification method based on multi-hop graph neural network

    CN119274079A

  • Multi-mode unmanned aerial vehicle image segmentation method

    CN119516187A