Urban green plant identification method and device, medium and equipment

By using the MSSA-Encoder, a multi-shape and spectral feature fusion encoder, and a feature fusion guidance module, the problem of accurate extraction of urban green spaces in complex environments was solved, achieving high-precision classification and identification of urban green spaces.

CN120913089AActive Publication Date: 2025-11-07CHINA UNIV OF PETROLEUM (EAST CHINA)

Patent Information

Application Number
CN202511041394.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-07
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately extract different types of urban green spaces in complex urban environments, especially due to small spectral differences and limitations in high-resolution image bands, resulting in insufficient identification accuracy.

Method used

The MSSA-Encoder, which integrates multi-shape and spectral features, enhances the identification of vegetation types and land cover boundaries by extracting and fusing deep and strip features, combined with pyramidal convolution and dynamic directional convolution kernels. Furthermore, the feature fusion guidance module weightedly fuses multispectral and multi-shape feature maps to achieve accurate classification of urban green spaces.

Benefits of technology

It significantly improves the accuracy of urban green space classification and the extraction precision of complex-shaped features, enhances the ability to identify subtle spectral differences among different vegetation types, and improves the completeness and accuracy of urban green space identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913089A_ABST
    Figure CN120913089A_ABST
Patent Text Reader

Abstract

The invention discloses an urban green plant identification method and device, a medium and equipment, and relates to the technical field of image identification, and the method comprises the steps: obtaining a to-be-identified remote sensing image; deep features in the remote sensing image are extracted, multi-scale spectral features in the deep features are obtained, the deep features and the multi-scale spectral features are fused, and a multi-spectral feature map is generated; wherein the deep features comprise color features and texture features; extracting strip features and multi-scale shape features in the remote sensing image, and fusing the strip features and the multi-scale shape features to generate a multi-shape feature map; performing weighted fusion on weights corresponding to the multi-spectral feature map and the multi-shape feature map to determine a fused feature map; and recovering the resolution of the fused feature map to the size of the original image, and classifying the feature map with the recovered resolution to obtain a final urban green land classification result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and in particular relates to a city green plant recognition method, device, medium and equipment. BACKGROUND

[0002] City green land is an important part of the city ecological system and has an important influence on the city environment. City green land can play an important role in many aspects, such as reducing air pollution, saving water resources, and serving as a carbon sink; reducing stress and anxiety, increasing the happiness of citizens; and providing entertainment venues for residents. More importantly, different types and structural patterns of city green land have different effects on relieving the urban heat environment, for example: the distribution structure of trees, shrubs and grasslands is different, and the degree of relief of the urban heat environment is also very different. In addition, in some areas, due to the lack of reliable data on the distribution of different types of underground geological exploration teams, this poses a major obstacle to effective policy-making and resource allocation. Therefore, fine-grained and accurate extraction of trees, shrubs and grasslands in city green land can more effectively improve the quality of the urban environment and provide a reference for city managers to more reasonably allocate city trees, shrubs and grasslands.

[0003] In recent years, in order to improve the integrity and accuracy of city green land recognition, high-resolution and ultra-high-resolution satellite images have been gradually applied to city green land classification. These high-resolution images provide complex details that the naked eye cannot detect, and the improvement of spatial resolution enables more fragmented and dispersed vegetation in the city to be identified, enhancing the integrity of city green land recognition. However, in narrow urban spaces, the spectral differences between different types of city green land are small, and the limitations of high-resolution image bands make it impossible to achieve accurate extraction of city green land solely relying on high-resolution images.

[0004] At the same time, in order to solve the problem of complex vegetation morphology in the city, researchers have gradually paid attention to the design of multi-scale network models. Multi-scale networks can improve the recognition effect of different size objects. Although multi-scale networks can greatly improve the recognition accuracy of city green land, in complex urban environments, green land is not always regular in shape, such as linear street trees and strip-shaped shrubs. Therefore, simply extracting regular information of different scales is far from enough, and it is also difficult to achieve accurate extraction of city green land. SUMMARY

[0005] The present application provides a city green plant recognition method, device, medium and equipment to solve the above-mentioned problems existing in the prior art, i.e. how to achieve accurate extraction of city green land in the prior art. The present application provides a city green plant recognition method, which comprises: obtaining a remote sensing image to be recognized; extract deep features and multi-scale spectral features from the remote sensing image, fuse the deep features and the multi-scale spectral features to generate a multi-spectral feature map, wherein the deep features include color features and texture features; extract strip features and multi-scale shape features from the remote sensing image, and fuse the strip features and the multi-scale shape features to generate a multi-shape feature map; determine a fused feature map by weighted fusion of weights corresponding to the multi-spectral feature map and the multi-shape feature map; restore the resolution of the fused feature map to the original image size, and then classify the feature map with restored resolution to obtain a final urban green space classification result.

[0006] Optionally, the deep features, the strip features, the multi-spectral feature map and the multi-shape feature map are obtained through a multi-shape and spectral feature fusion encoder MSSA-Encoder, and the multi-shape and spectral feature fusion encoder MSSA-Encoder specifically includes: a first branch, a second branch, a third branch and a fourth branch; wherein the first branch includes two 3x3 DOConvs connected in sequence, for extracting deep features from the remote sensing image; the second branch and the third branch are two parallel 1x3 and 3x1 DOConvs, for extracting strip features from the remote sensing image; and the fourth branch is a dilated spatial convolution pyramid module including a 1x1 convolution, three 3x3 convolutions with different dilated rates, for extracting multi-scale spectral features and multi-scale shape features from the deep features and the remote sensing image respectively.

[0007] Optionally, the weights corresponding to the multi-spectral feature map and the multi-shape feature map are weighted fused through a feature fusion guide module FFG, and the feature fusion guide module FFG includes a first branch and a second branch in parallel; the first branch is used to select optimal shape features from multi-shape and multi-spectral information, and the second branch is used to obtain fused features of shallow spatial and spectral information; The shape features and the fused features of shallow spatial and spectral information are dynamically aggregated, and the specific calculation formula is as follows: wherein, is the total weight, and are initial weights of the feature maps, and are dynamic weights assigned to the two feature maps by the network during the training process.

[0008] Optionally, the remote sensing image is a pre-processed Gaofen-2 image and a Sentinel-2 image, and the pre-processing specifically includes radiation correction, atmospheric correction and orthorectification processing.

[0009] The present application provides a kind of urban green plant identification device, comprising: The acquisition module is configured to acquire a remote sensing image to be identified. The multispectral feature map acquisition module is configured to extract deep features in the remote sensing image, acquire multiscale spectral features in the deep features, and fuse the deep features and the multiscale spectral features to generate a multispectral feature map. The multi-shape feature map acquisition module is configured to extract strip features and multiscale shape features in the remote sensing image, and fuse the strip features and the multiscale shape features to generate a multi-shape feature map. The fusion module is configured to fuse the multispectral feature map and the multi-shape feature map by weighting corresponding weights of the multispectral feature map and the multi-shape feature map. The classification module is configured to restore the resolution of the fused feature map to the original image size, classify the feature map with restored resolution, and obtain a final urban green space classification result.

[0010] The present application provides a computer-readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned urban green plant identification method.

[0011] The present application provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to implement the above-mentioned urban green plant identification method.

[0012] Compared with the prior art, the present application has the following beneficial effects: the present application provides a city green plant identification method, which deeply integrates shape perception and spectral information perception technology into a classification network, constructs a multi-shape and spectral feature fusion encoder, specifically, through convolution operation using a dynamic direction convolution kernel, deep features and strip features of the image to be measured, such as vegetation types and ground object boundaries, can be accurately extracted; by using pyramid convolution to extract multi-scale features in the image, including multi-scale shape features and multi-scale spectral features, then by fusing the multi-scale spectral features with the deep features, a multi-spectral feature map is output, which can enhance the identification ability of subtle spectral differences between different vegetation, at the same time, by fusing the multi-scale shape features and strip features, the extraction accuracy of complex shape ground objects can be effectively improved; in addition, by weighting and fusing the weights corresponding to the multi-spectral feature map and the multi-shape feature map, a fused feature map is obtained, realizing deep fusion of shape and spectral features, providing high discriminative feature input for the subsequent decoder, thereby significantly improving the accuracy of city green classification. BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.

[0014] Figure 1 A flowchart of a city green plant identification method provided by an embodiment of the present application; Figure 2 A schematic diagram of the overall structure of a multi-shape and spectral feature fusion encoder MSSA-Encoder provided by an embodiment of the present application; Figure 3 A schematic diagram of the structure of a feature fusion guide module FFG provided by an embodiment of the present application; Figure 4 A schematic diagram of classification results of different deep learning models provided by an embodiment of the present application; Figure 5 A schematic diagram of a computer device for a city green plant identification method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0015] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0016] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The embodiments of the present application will be described below with reference to the drawings.

[0017] Figure 1 is a flowchart of a city green plant identification method provided by an embodiment of the present application, as shown in the embodiment of the present application, the city green plant identification method comprises: Figure 1 S1: obtaining a remote sensing image to be identified.

[0018] Illustratively, the remote sensing image can be a Gaofen-2 and Sentinel-2 remote sensing image, and the obtained Gaofen-2 and Sentinel-2 remote sensing image is subjected to radiation correction, atmospheric correction and orthorectification. The Gaofen-2 image with a resolution of 1m and the Sentinel-2 image with a resolution of 10m are obtained.

[0019] Illustratively, the three red edge bands of the Sentinel-2 image can be extracted and resampled, and finally three red edge bands with a resolution of 1m are obtained.

[0020] After obtaining the red edge band with a resolution of 1m and the panchromatic band of the high-resolution image, the Gram-Schmidt method integrated in ENVI is selected for image fusion to obtain a fused red edge image, which is then superimposed with the corrected Gaofen-2 image of S2-1 to obtain the final experimental image as a green plant image dataset, which can include red, green, blue, infrared and three near-infrared bands, a total of seven bands. All high-resolution remote sensing images in the dataset are segmented into image blocks with a size of 128*128 before use. To improve the stability of the model, the obtained images can be expanded by using mirror image, rotation and adding salt and pepper noise.

[0021] S2: extracting deep features in the remote sensing image, obtaining multi-scale spectral features in the deep features, fusing the deep features and the multi-scale spectral features to generate a multi-spectral feature map; wherein the deep features include color features and texture features.

[0022] Optionally, the deep features, strip features, multi-spectral feature map and multi-shape feature map are obtained by a multi-shape and spectral feature fusion encoder MSSA-Encoder, and the multi-shape and spectral feature fusion encoder MSSA-Encoder specifically comprises: ​the first branch, the second branch, the third branch and the fourth branch; wherein the first branch comprises two 3x3 DOConvs connected in sequence, for extracting deep features from the remote sensing image; the second branch and the third branch are two parallel 1x3 and 3x1 DOConvs, for extracting strip features from the remote sensing image; the fourth branch is a dilated spatial convolution pyramid module comprising a 1x1 convolution, three 3x3 convolutions with different dilated rates, for extracting multi-scale spectral features and multi-scale shape features from the deep features and the remote sensing image respectively.

[0023] S3: extracting strip features and multi-scale shape features from the remote sensing image, and fusing the strip features and the multi-scale shape features to generate a multi-shape feature map.

[0024] S4: determining a fused feature map by weighted fusion of weights corresponding to the multi-spectral feature map and the multi-shape feature map.

[0025] For example, the weights corresponding to the multi-spectral feature map and the multi-shape feature map are weighted fused by a feature fusion guide module FFG, which comprises a first branch and a second branch in parallel; the first branch is used to select optimal shape features from the multi-shape and multi-spectral information, and the second branch is used to obtain fused features of shallow spatial and spectral information.

[0026] S5: restoring the resolution of the fused feature map to the original image size, and then classifying the feature map with restored resolution to obtain a final urban green space classification result.

[0027] For example, before classifying and identifying the green plant remote sensing image to be identified, model parameters need to be set, which can specifically include selecting an original high-resolution image to be classified, determining the total number S of classes for classification, determining the image segmentation size img_size and the number of training samples; then, the number of down-sampling layers and the number of up-sampling layers in the encoding and decoding parts of the deep learning model are determined; then, the network learning rate, the number of optimization iterations and the model optimizer Adam are determined; finally, the number of one-dimensional vector elements output by the feature classifier can be set according to the determined total number of classes.

[0028] For example, based on the set model parameters, a semantic segmentation network MSSFNet is set, and the training data prepared in S2, i.e. the green plant image dataset, is used to complete the training of the model, and the specific structure of the network is as shown in Figure 1 The network has the following components:

[0029] (1) an input layer for receiving a remote sensing image as input; (2) Multi-shape and spectral feature fusion encoder: used to combine multi-scale deep features of different receptive fields and different shapes, and transmit the upper layer features to the lower layer through the down-sampling layer to further mine the multi-scale information of the features.

[0030] (3) Feature fusion guidance module: the multi-shape and multi-spectral information obtained from the encoder are weighted and fused to guide the decoder to make full use of the tiny shape and spectral difference information.

[0031] (4) Decoder: used to restore the resolution of the feature map to the size of the original image, and the decoder includes four up-sampling layers in total.

[0032] (5) Classifier module: used to obtain the final soft probability map through convolution, and then obtain the remote sensing image classification map.

[0033] The specific structure of the multi-shape and spectral feature fusion encoder is shown in Figure 2 In the multi-scale feature fusion encoder, the ordinary convolution kernel is replaced by DOConvs to form the local shape information enhancement module; the DOConv integrates the adaptive channel focusing mechanism to dynamically adjust the channel weight and improve the accuracy of the model. Subsequently, the module is designed to consist of four branches to improve the model's representation ability for different size and shape features. The first branch is composed of two 3x3 DOConvs for further deep feature extraction; the second and third branches are two parallel 1x3 and 3x1 DOConvs, and the two parallel strip convolutions obtain strip feature enhanced semantic information. In a narrow urban area, the extraction of strip-shaped vegetation information is of great help. The fourth branch is an upgraded pyramid convolution module, the empty space convolution pooling pyramid module, which can include: a 1x1 convolution, three 3x3 convolutions with different hole rates, and the hole rates can be selected as 6, 12, and 18. This module can extract large and small, regular shape green spaces in the urban area and fuse their different scale features. Finally, the features obtained by the strip convolution and the pyramid convolution are added and the maximum value is taken to obtain different scale features while retaining different shape features. Subsequently, the fused multi-shape features are input into the transformer encoder to focus on their different spectral features.

[0034] The specific structure of the feature fusion guidance module is shown in Figure 3As shown, the feature fusion guidance module includes two branches. Among them, the first branch is an enhanced shape feature, which extracts the optimal shape feature from the encoder multi-scale fusion block; the second branch is a global spatial-spectral feature, which is extracted from the last layer of each layer of the encoder, representing the fusion of shallow spatial and spectral information, and after obtaining the two, it is dynamically aggregated together to form an enhanced-shape-spectral feature, which is used to guide the decoding process, specifically:

[0035] In the formula, is the total weight, and are the initial weights of the feature map, and are the dynamic weights assigned by the network to the two feature maps during training.

[0036] For example, first, the acquired remote sensing image data can be preprocessed, such as radiation correction, atmospheric correction, and orthorectification processing, to lay the foundation for subsequent analysis. Next, shape perception and spectral information perception technology are integrated into the classification network. When processing remote sensing images, the network can capture the shape features of complex and variable profiles, distribution patterns, etc. of vegetation such as trees, shrubs and grasslands through the former, and identify the subtle spectral differences between different vegetation using the latter, to obtain rich vegetation information. Then, the network transmits the extracted shape and spectral information to the feature fusion guidance decoder. This decoder will fuse and reconstruct the spectral spatial features of trees, shrubs and grasslands, and extract more representative features. Finally, based on these reconstructed features, the model can classify each pixel in the remote sensing image and divide it into the corresponding vegetation category, thereby realizing remote sensing semantic segmentation of different types of vegetation and completing the precise identification and segmentation of vegetation in complex urban environments.

[0037] In the comparative experiment, the deep learning network MSSFNet in the present application is compared with the mainstream deep learning remote sensing image classification algorithm. Table 1 shows the overall accuracy (OA), average pixel accuracy (MPA), mean intersection over union (MIOU), and the accuracy comparison results of the intersection over union (IOU) and pixel accuracy (PA) of each feature of the deep learning model UNet, Deeplabv3+, PSPNet, etc. commonly used in the present application and the present remote sensing image classification.

[0038] Table 1 Accuracy comparison results of different deep learning models Figure 4 The extraction examples of the classification method on the data set are given, and it can be seen that the deep learning network on the fused image proposed in the present application performs best.

[0039] The urban green plant identification method provided in one or more embodiments of the present specification is based on the same idea. The present specification also provides a corresponding urban green plant identification device, comprising: An acquisition module is configured to acquire a remote sensing image to be identified. A multispectral feature map acquisition module is configured to extract deep features in the remote sensing image, acquire multiscale spectral features in the deep features, and fuse the deep features and the multiscale spectral features to generate a multispectral feature map. The deep features include color features and texture features. A multi-shape feature map acquisition module is configured to extract strip features and multiscale shape features in the remote sensing image, and fuse the strip features and the multiscale shape features to generate a multi-shape feature map. A fusion module is configured to fuse the multispectral feature map and the multi-shape feature map by weighting the corresponding weights to determine a fused feature map. A classification module is configured to restore the resolution of the fused feature map to the original image size, and then classify the feature map with restored resolution to obtain a final urban green classification result.

[0040] The specific limitations of the urban green plant identification device can be referred to the limitations of the urban green plant identification method described above, which will not be repeated here. Each module in the above urban green plant identification device can be realized by software, hardware and their combination. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0041] The present application also provides a computer readable storage medium, which stores a computer program. The computer program can be used to execute the urban green plant identification method provided above.

[0042] The present application also provides Figure 5 The structure diagram of the computer device is shown in FIG. 1. Figure 5 As shown in FIG. 1, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and of course can also include other hardware required by the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to realize the urban green plant identification method provided in the above embodiments.

[0043] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments of the methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0044] The technical features of the above embodiments can be combined in any way. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

Claims

1. A method for identifying urban green plants, characterized in that, The method comprises the following steps: obtaining a remote sensing image to be identified; extracting deep features in the remote sensing image, obtaining multi-scale spectral features in the deep features, fusing the deep features and the multi-scale spectral features, and generating a multi-spectral feature map; wherein the deep features include color features and texture features; extracting strip features and multi-scale shape features in the remote sensing image, and fusing the strip features and the multi-scale shape features to generate a multi-shape feature map; determining a fused feature map by weighted fusion of corresponding weights of the multi-spectral feature map and the multi-shape feature map; restoring the resolution of the fused feature map to the original image size, and then classifying the feature map with restored resolution to obtain a final urban green space classification result.

2. The urban green plant identification method of claim 1, wherein, The deep features, the strip features, the multi-spectral feature map and the multi-shape feature map are obtained through a multi-shape and spectral feature fusion encoder MSSA-Encoder, which specifically comprises: a first branch, a second branch, a third branch and a fourth branch; wherein the first branch comprises two 3x3 DOConvs connected in sequence for extracting deep features from the remote sensing image; the second branch and the third branch are two parallel 1x3 and 3x1 DOConvs for extracting strip features from the remote sensing image; and the fourth branch is a dilated spatial convolution pyramid module comprising a 1x1 convolution, three 3x3 convolutions with different dilated rates for extracting multi-scale spectral features and multi-scale shape features from the deep features and the remote sensing image respectively.

3. The method of claim 1, wherein the step of identifying the urban greenery is performed by using a machine learning algorithm. The corresponding weights of the multi-spectral feature map and the multi-shape feature map are weighted fused through a feature fusion guide module FFG, which comprises a first branch and a second branch in parallel; the optimal shape feature is selected from the multi-shape and multi-spectral information through the first branch, and the fusion feature of the shallow space and spectral information is obtained through the second branch; The shape feature and the fusion feature of the shallow space and spectral information are dynamically aggregated, and the specific calculation formula is as follows:

4. where, is the total weight, and are the initial weights of the feature maps, respectively, and are the dynamic weights assigned by the network to the two feature maps during the training process, respectively.

5. The method of claim 1, wherein the step of identifying the urban greenery is performed by using a neural network. The remote sensing image is a preprocessed Gaofen-2 image and a Sentinel-2 image; wherein the preprocessing specifically includes radiation correction, atmospheric correction and orthorectification processing.

6. An urban green plant identification device, characterized by, The method comprises the following steps: an acquisition module for acquiring a remote sensing image to be identified; a multi-spectral feature map acquisition module for extracting deep features in the remote sensing image, obtaining multi-scale spectral features in the deep features, fusing the deep features and the multi-scale spectral features, and generating a multi-spectral feature map; wherein the deep features include color features and texture features; a multi-shape feature map acquisition module for extracting strip features and multi-scale shape features in the remote sensing image, and fusing the strip features and the multi-scale shape features to generate a multi-shape feature map; a fusion module for determining a fused feature map by weighted fusion of corresponding weights of the multi-spectral feature map and the multi-shape feature map; a classification module for restoring the resolution of the fused feature map to the original image size, and then classifying the feature map with restored resolution to obtain a final urban green space classification result.

7. A computer readable storage medium characterized in that, The storage medium stores a computer program, and the computer program is executed by a processor to implement the urban green plant identification method in any one of claims 1-4.

8. A computer device, comprising: The computer program is stored in the memory and executable on the processor, and the processor executes the program to implement the urban green plant identification method in any one of claims 1-4.

Citation Information

Patent Citations

  • Urban vegetation inversion method and system based on high-resolution remote sensing image

    CN112381013A

  • Hyperspectral image recognition method and device, electronic equipment and readable storage medium

    CN113822209A

  • Hyperspectral image classification method and device, equipment and storage medium

    CN114120033A

  • Urban green land fine classification method and system based on GF-2 and open map data

    CN115984603A

  • Urban green land segmentation method, device and system and storage medium

    CN118447247A

Cited By

  • Wetland plant identification method and system based on multi-dimensional feature similarity comparison

    CN121582794A