Multi-source high-resolution remote sensing image vegetation extraction method based on multi-feature comprehensive perception
Patent Information
- Application Number
- CN202510078099.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The prior art is prone to problems such as mis-lifting, internal fragmentation, unclear boundaries and reduced model generalization capabilities when extracting vegetation in high-resolution remote sensing images, especially in different data sources and complex scenarios.
A multi-source high-resolution remote sensing image vegetation extraction method based on multi-feature comprehensive perception is adopted. Vegetation index features are selected through a random forest model, and a multi-feature comprehensive perception convolution network (MSCIN) with multi-scale feature information enhancement, global information interaction, and feature cross-fusion is constructed to achieve high-precision extraction of vegetation.
The extraction results of this method are accurate and robust under different sensors, overcome the problem of rapid attenuation of other methods in different sensors, solve internal fragmentation, misrelaxation, and missed lifting, and realize vegetation extraction under multi-source high-resolution remote sensing data.
Smart Images

Figure CN119992329A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of photogrammetry data processing, and in particular relates to a multi-source high-resolution remote sensing image vegetation extraction method based on multi-feature comprehensive perception. Background Art
[0002] Vegetation is an important part of the earth's ecosystem and is of great significance to the protection and construction of the ecological environment. Vegetation extraction based on the spectral, texture, spatial, and temporal characteristics of vegetation on high-resolution remote sensing images complements multi-source data and extracts the spatiotemporal distribution information of vegetation, which is of great significance for resource surveys, urban planning, land surveys, and forest fire monitoring. However, due to adverse effects such as background noise during remote sensing imaging, a series of problems often occur in the recognition results, such as false positives, internal fragmentation, unclear boundaries, and reduced model generalization ability.
[0003] In recent years, remote sensing technology has become an important technical means for vegetation monitoring due to its wide range and timeliness. The main methods for studying vegetation extraction at present are: threshold segmentation method based on vegetation index. This method constructs spectral indicators to extract vegetation by the difference in reflectance between the red band and the near-infrared band. This method can extract vegetation from images, but it is difficult to identify vegetation under the shade of tall buildings. In areas with sparse vegetation, the vegetation index method may not be able to accurately estimate vegetation coverage. Many machine learning methods are also used for vegetation extraction, such as maximum likelihood classification (MLC), support vector machine (SVM) and random forest (RF). Although these machine learning methods can achieve high accuracy in images, the features and threshold conditions need to be manually designed, and it is difficult to achieve automatic extraction of wide-area and multi-source data.
[0004] With the improvement of image spatial resolution and the diversification of ground information, the semantic information of the scene has become more complex, and deep learning methods have gradually become a research hotspot in artificial intelligence. At present, it has been widely used in tasks such as image recognition, image segmentation, and target detection. In the field of vegetation remote sensing, deep learning methods are mainly divided into semantic segmentation-based methods and pixel-based methods. They independently learn relevant data features in an end-to-end manner, which can meet the needs of vegetation assessment and monitoring under the growing remote sensing data. However, although the semantic segmentation method can achieve good accuracy in vegetation extraction, its essence is to aggregate homogeneous pixels, usually ignoring some detailed features, unable to take into account the detailed features inside the vegetation features, and easily confused with other symbiotic features. In addition, the semantic segmentation method requires accurate boundary information when making labels. It is difficult to distinguish mixed pixels at the boundary of the symbiotic area between vegetation and other land types. It is difficult to make sample labels on multi-source remote sensing images, which in turn leads to the problem that the semantic segmentation model is difficult to converge during the training process.
[0005] Based on pixel classification, it is only necessary to add some labeled pixels to the sample images of a specific scene, which can quickly and easily adapt to new scenes. However, most remote sensing vegetation extraction methods only extract vegetation from a single data source based on the spatial spectral characteristics of vegetation, or combine vegetation index features to extract vegetation by splicing them into the network for training. It is difficult to achieve high-precision vegetation extraction across data sources when only spectral features are used for vegetation extraction due to large differences in spectral values due to phase differences, radiation differences, etc. The spectral feature values are very different from the index feature values, and the introduction of vegetation index features to merge them in a splicing manner often causes the index features to be ignored, making it difficult to achieve complementary correction between multiple features, which will lead to problems such as missing and erroneous extraction in vegetation extraction from different data sources. In addition, the selection of vegetation index features relies solely on manual experience to achieve feature screening, lacks quantitative screening methods, and is difficult to scientifically and reasonably verify the validity of the index.
[0006] In view of this, the inventors hope to provide a method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception. Summary of the invention
[0007] The purpose of the present invention is to overcome the above-mentioned problems existing in the traditional technology and to provide a method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception.
[0008] In order to achieve the above technical objectives and the above technical effects, the present invention is implemented through the following technical solutions:
[0009] The present invention provides a method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception, comprising the following steps:
[0010] Step 1: Sample set preparation;
[0011] Step 2: Optimization of vegetation index features;
[0012] Step 3: Network model structure design;
[0013] Step 4: Multi-scale dense connection module;
[0014] Step 5: Dual-path multi-head cross-attention feature fusion module;
[0015] Step 6: Accuracy evaluation of model extraction results.
[0016] Furthermore, the sub-steps of step one are as follows:
[0017] 1) Obtain GF-2 remote sensing image data of the target area;
[0018] 2) To better show the characteristics of vegetation, the remote sensing image was displayed in ArcGIS 10.7 software with 4-3-2 bands, the characteristics of vegetation were shown in red, and the vegetation was outlined through manual visual interpretation;
[0019] 3) The preliminary delineation results are repeatedly revised, and samples are collected inside and on the edge of the delineated plot to obtain vegetation labels and background labels, and obtain sample coordinate points;
[0020] 4) Considering the spatial correlation between adjacent pixels, a 5×5 neighborhood window is used as the training sample with the pixel of the sampling point as the center to expand outward to obtain more spatial and spectral information; combining various data sources, the vegetation contour is outlined through manual visual interpretation to obtain the ground truth label data as verification data.
[0021] Furthermore, the sub-steps of step 2 are as follows:
[0022] 1) Taking vegetation index features as independent variables and sample labels as dependent variables, the random forest regression model was used to fit the nonlinear relationship between the two and perform feature optimization. The vegetation indices were selected as follows: NDVI, GNDVI, EVI, OSAVI, RVI, DVI, TVI, GVI, GI, NDGI, MCARI, and TCARI. The number of decision trees in the model was set to 20 and the depth was set to 4.
[0023] 2) According to the fitting results of the random forest model, the values are sorted. The larger the value, the more important the vegetation index feature is to the accuracy of model prediction. The top four vegetation index features are selected.
[0024] Furthermore, in step three, the input data of the model are remote sensing image spectral data and vegetation index. The cropped 5×5 samples are input into the parallel two-branch network respectively. The two networks contain three parallel sub-networks with scales of 1×1, 3×3, and 5×5 respectively. A simplified dense connection module is introduced to learn spectral and index features of different scales, and downsampling dense connections are used to perform interactive correction between features.
[0025] Furthermore, in step three, based on the self-attention of the visual transformer to obtain the maximum receptive field, a dual-path cross-attention feature fusion module is designed between the 5×5 and 3×3 branches; the shallow initial features themselves are cross-fused with the enhanced index features to establish multi-feature comprehensive perception to achieve global information complementarity, and the hybrid channel output is used to enhance the features.
[0026] Furthermore, in step three, a parallel dual-branch multi-scale network is designed, namely, a spectral multi-scale network branch and an exponential multi-scale network branch; coordinate convolution is introduced into the 5×5 and 3×3 branch networks respectively, and the convolution layer is used to optimize the extraction of spatial feature position information. At the same time, jump connections are designed inside the branch networks for extracting index features at different scales to enhance the fusion of vegetation index features at different scales and make the vegetation boundary extraction more complete; after continuous convolution, the dual-branch parallel sub-network outputs 1×1 spectral features and 1×1 exponential features respectively;
[0027] Then, the spectral features and exponential features generated by parallel splicing of networks of different scales are merged; a deep separable convolutional layer DWCConv is designed 3×3 Finely extract features to reduce parameter calculations;
[0028] Finally, after two fully connected layers and the Softmax classifier, the vegetation classification result is output.
[0029] Furthermore, in step 4, in order to realize the feature information exchange between 5×5, 3×3 and 1×1, a multi-scale dense connection module is introduced between the spectral feature and exponential feature parallel dual-branch networks; in the multi-scale dense connection module, each sub-network can receive feature information from other parallel networks, so that the network can exchange information at different scales, thereby enhancing the feature expression ability; in the dense connection block used, the features of the previous layer are passed to all the subsequent layers in a stacked manner;
[0030] In the 5×5 and 3×3 branch structures, vertical downsampling is performed to expand the width of the network; the features of the 5×5 branch are superimposed on the 3×3 and 1×1 branch networks through dense connections, and then the features of the 3×3 branch are superimposed on the 1×1 branch network;
[0031] In order to retain feature information at different scales and achieve feature fusion, each layer of the branch network is spliced with other branch networks in the channel dimension. This method not only effectively alleviates the problem of gradient disappearance, but also reduces the number of parameters, thereby enhancing the ability of feature extraction and reuse. That is to say, in the L-layer deep network structure, the network of this embodiment has L×(L+1) / 2 connections.
[0032] Furthermore, in step 5, the input feature F(H, W, C) is flattened to F(N, C) in this module, and then linear embedding is used to generate K(key), V(value), and Q(query) for each attention head. Each head calculates the attention score G of K and V by the multispectral branch and the vegetation index branch. MS and G VI ;
[0033] Then, Softmax is used to normalize the attention score and the vegetation index feature G VI With spectral characteristics Q MS Cross-multiplication makes the vegetation index branch focus on the characteristics of the spectral branch; at the same time, the spectral feature G MS and vegetation index characteristic Q VI Cross-multiplication makes the spectral feature branch allocate attention to the features of the vegetation index branch;
[0034] In the two sequence branches, branch complementary interaction is performed to achieve feature fusion; the calculation formula is as follows:
[0035]
[0036] Furthermore, in step six, four evaluation indicators are used to evaluate the accuracy: overall accuracy OA, F1 score, intersection over union ratio, and accuracy;
[0037] The overall accuracy OA refers to the ratio of the sum of the number of pixels in the category correctly identified by the model to the pixels of all identified categories in the test image;
[0038] The F1 score is the harmonic mean of precision and recall, and is used to evaluate the model's ability to handle imbalanced datasets or sensitivity to misclassified categories;
[0039] IoU is the degree of overlap between the model prediction result and the true annotated area, which is used to evaluate the accuracy and robustness of the model;
[0040] The Kappa coefficient (Kappa) for remote sensing interpretation accuracy assessment estimates the consistency between the prediction and the ground reference;
[0041] The formula for the above evaluation index is as follows:
[0042]
[0043] Among them, TP stands for true positive, FP stands for false positive, TN stands for true negative, FN stands for false negative, and N is the total number of samples; false positives represent the number of non-field pixels that are misclassified as positive, and false negatives represent positive pixels that are misclassified as negative.
[0044] The beneficial effects of the present invention are:
[0045] 1. The method of the present invention is scientifically and reasonably designed. First, the vegetation index is optimized through the random forest model to screen out the inter-class difference index that can enhance the vegetation and other ground objects. On this basis, a multi-feature comprehensive perception convolutional network (MSCIN) with multi-scale feature information enhancement, global information interaction, and feature cross-fusion is constructed. The network constructs a dual-branch parallel network of spectral features and vegetation index features. On the basis of simplifying the dense connection module, it strengthens the multi-scale feature information extraction capability while reducing the loss of detail features. In addition, in order to promote global information interaction between the original spectral information and the vegetation index features, a dual-path multi-head cross-attention fusion module is designed to expand the difference between vegetation and other ground objects and the consistency of vegetation, improve the generalization performance of the network, and realize vegetation extraction under multi-source high-resolution remote sensing data.
[0046] 2. The extraction results of the method of the present invention are robust in accuracy under different sensors, overcoming the problem of rapid attenuation of accuracy of other methods under different sensors, and at the same time solving the problems of internal fragmentation, false extraction, and missed extraction caused by sample generalization and image diversity.
[0047] Of course, any product implementing the present invention does not necessarily need to achieve all of the above advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0049] Figure 1 It is a general model structure diagram of an embodiment of the present invention;
[0050] Figure 2 It is a structural diagram of a dual-path cross-attention feature fusion module according to an embodiment of the present invention;
[0051] Figure 3 Result diagrams of vegetation extraction of different models according to embodiments of the present invention. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] like Figure 1-Figure 3As shown, this embodiment provides a multi-source high-resolution remote sensing image vegetation extraction method based on multi-feature comprehensive perception, including the following steps:
[0054] Step 1: Sample set preparation;
[0055] Step 2: Optimization of vegetation index features;
[0056] Step 3: Network model structure design;
[0057] Step 4: Multi-scale dense connection module;
[0058] Step 5: Dual-path multi-head cross-attention feature fusion module;
[0059] Step 6: Accuracy evaluation of model extraction results.
[0060] Step 1 specifically includes the following:
[0061] 1) Obtain GF-2 remote sensing image data of the target area;
[0062] 2) To better show the characteristics of vegetation, the remote sensing image was displayed in ArcGIS 10.7 software with 4-3-2 bands, the characteristics of vegetation were shown in red, and the vegetation was outlined through manual visual interpretation;
[0063] 3) The preliminary delineation results are repeatedly revised, and samples are collected inside and on the edge of the delineated plot to obtain vegetation labels and background labels, and obtain sample coordinate points;
[0064] 4) Considering the spatial correlation between adjacent pixels, a 5×5 neighborhood window is used as the training sample with the pixel of the sampling point as the center to expand outward to obtain more spatial and spectral information; combining various data sources, the vegetation contour is outlined through manual visual interpretation to obtain the ground truth label data as verification data.
[0065] Step 2 specifically includes the following:
[0066] 1) Taking vegetation index features as independent variables and sample labels as dependent variables, the random forest regression model was used to fit the nonlinear relationship between the two and perform feature optimization. The vegetation indices were selected as follows: NDVI, GNDVI, EVI, OSAVI, RVI, DVI, TVI, GVI, GI, NDGI, MCARI, and TCARI. The number of decision trees in the model was set to 20 and the depth was set to 4.
[0067] 2) According to the fitting results of the random forest model, the values are sorted. The larger the value, the more important the vegetation index feature is to the accuracy of model prediction. The top four vegetation index features are selected.
[0068] Step 3 specifically includes the following:
[0069] The input data of the model are remote sensing image spectral data and vegetation index. The cropped 5×5 samples are input into parallel two-branch networks. The two networks contain three parallel sub-networks with scales of 1×1, 3×3, and 5×5 respectively. A simplified dense connection module is introduced to learn spectral and index features of different scales, and downsampling dense connections are used to perform interactive correction between features.
[0070] Furthermore, based on the self-attention of the visual transformer to obtain the maximum receptive field, a dual-path cross-attention feature fusion module is designed between the 5×5 and 3×3 branches; the shallow initial features themselves are cross-fused with the enhanced exponential features to establish multi-feature comprehensive perception to achieve global information complementarity, and the hybrid channel output is used to enhance the features.
[0071] Furthermore, a parallel dual-branch multi-scale network is designed, namely, a spectral multi-scale network branch and an exponential multi-scale network branch; coordinate convolution is introduced into the 5×5 and 3×3 branch networks respectively, and the convolution layer is used to optimize the extraction of spatial feature position information. At the same time, jump connections are designed inside the branch networks for extracting index features at different scales to enhance the fusion of vegetation index features at different scales and make the vegetation boundary extraction more complete; after continuous convolution, the dual-branch parallel sub-network outputs 1×1 spectral features and 1×1 exponential features respectively; then, the spectral features and exponential features generated by parallel splicing of networks of different scales are merged; a deep separable convolutional layer DWCConv is designed 3×3 The features are extracted finely to reduce the amount of parameter calculation; finally, after two fully connected layers and a Softmax classifier, the vegetation classification results are output.
[0072] Step 4 specifically includes the following:
[0073] In order to realize the feature information exchange between 5×5, 3×3 and 1×1, a multi-scale dense connection module is introduced between the spectral feature and exponential feature parallel dual-branch networks; in the multi-scale dense connection module, each sub-network can receive feature information from other parallel networks, so that the network can exchange information at different scales, thereby enhancing the feature expression ability.''
[0074] In the densely connected block used, the features of the previous layer are passed to all subsequent layers in a stacked manner; in the 5×5 and 3×3 branch structures, vertical downsampling is performed to expand the width of the network; the features of the 5×5 branch are superimposed on the 3×3 and 1×1 branch networks by dense connection, and then the features of the 3×3 branch are superimposed on the 1×1 branch network. In order to retain feature information at different scales and achieve feature fusion, the branch networks of each layer are spliced with other branch networks in the channel dimension. This method not only effectively alleviates the problem of gradient vanishing, but also reduces the amount of parameters, thereby enhancing the ability of feature extraction and reuse. That is to say, in the L-layer deep network structure, the network of this embodiment has L×(L+1) / 2 connections.
[0075] Step 5 specifically includes the following:
[0076] In this module, the input features F(H, W, C) are flattened to F(N, C), and then linear embedding is used to generate K(key), V(value), and Q(query) for each attention head. Each head calculates the attention score G of K and V by the multispectral branch and the vegetation index branch. MS and G VI ;
[0077] Then, Softmax is used to normalize the attention score and the vegetation index feature G VI With spectral characteristics Q MS Cross-multiplication makes the vegetation index branch focus on the characteristics of the spectral branch; at the same time, the spectral feature G MS and vegetation index characteristic Q VI Cross-multiplication makes the spectral feature branch allocate attention to the features of the vegetation index branch;
[0078] In the two sequence branches, branch complementary interaction is performed to achieve feature fusion; the calculation formula is as follows:
[0079]
[0080] Step 6 specifically includes the following:
[0081] In order to verify the effectiveness of this method in extracting vegetation, four evaluation indicators were used for accuracy evaluation: overall accuracy OA, F1 score, intersection over union and accuracy; overall accuracy OA refers to the ratio of the sum of the number of pixels in the category correctly recognized by the model in the test image to the pixels of all recognized categories; F1 score is the harmonic mean of precision and recall, which is used to evaluate the ability of the model in dealing with unbalanced data sets or sensitivity to misclassified categories; IoU is the degree of overlap between the model prediction results and the true annotated area, which is used to evaluate the accuracy and robustness of the model; the Kappa coefficient (Kappa) of remote sensing interpretation accuracy evaluation estimates the consistency between the prediction and the ground reference.
[0082] The formula for the above evaluation index is as follows:
[0083]
[0084]
[0085] Among them, TP stands for true positive, FP stands for false positive, TN stands for true negative, FN stands for false negative, and N is the total number of samples; false positives represent the number of non-field pixels that are misclassified as positive, and false negatives represent positive pixels that are misclassified as negative.
[0086] A specific application of this embodiment is:
[0087] Obtain two GF-2 image data of Ma'anshan City, Anhui Province in October 2022, manually outline and sample them, obtain 266,547 vegetation pixels and 435,205 background pixels respectively, and make them into 5×5 standard sample data for input into the model. This embodiment was experimented with tensflower2.4.0 framework in Python3.7 environment. All experiments were conducted on a device equipped with a 1080Ti graphics card. The Adam optimizer was used to optimize the network. During the training process, all experiments used the same hyperparameters, including the number of training rounds (40 rounds), batch size (512), and initial learning rate (0.001).
[0088] As shown in Table 1 and Figure 3 As shown in the figure, with the GF-2 remote sensing image as the data source, it can be seen from the figure that the four models have good effects on vegetation extraction, and can extract vegetation relatively completely in mountainous areas. However, in some complex urban areas A(4)-(7), the mixed situation of urban roads, buildings and vegetation is prone to misclassification. Compared with other methods, the method of this embodiment has fewer cases of misclassification and omission, and the boundary extraction is more complete, reaching 86.46%, 92.85% and 85.58% in IoU, OA and Kappa coefficient respectively.
[0089] The method of this embodiment can extract vegetation relatively completely on GF7 and Planet images, with less omission and error extraction.
[0090] In the GF7 image, the method of this embodiment achieves 84.21%, 92.03%, and 86.81% in IoU, OA, and Kappa coefficient, respectively.
[0091] On the Planet image, the method of this embodiment achieved 86.96%, 91.82%, and 83.28% in IoU, OA, and Kappa coefficient, respectively.
[0092] This shows that MSICN still maintains high internal consistency and boundary accuracy for vegetation extraction across data sources and has high generalization ability.
[0093] Table 1 Comparison of accuracy verification results of different models
[0094]
[0095] In view of the problems of missing and wrong extraction in vegetation extraction from multiple data sources, a convolutional network (MSICN) for vegetation extraction with multi-feature comprehensive perception is constructed. The method first optimizes the vegetation index through the random forest model, and screens out the inter-class difference index that can improve the vegetation and other ground objects. On this basis, a multi-feature comprehensive perception convolutional network (MSCIN) with multi-scale feature information enhancement, global information interaction, and feature cross-fusion is constructed. The network strengthens the multi-scale feature information extraction capability while reducing the loss of detail features by constructing a dual-branch parallel network of spectral features and vegetation index features on the basis of simplifying dense connection modules. In addition, in order to promote global information interaction between the original spectral information and the vegetation index features, a dual-path multi-head cross-attention fusion module is designed to expand the difference between vegetation and other ground objects and the consistency of vegetation, improve the generalization performance of the network, and realize vegetation extraction under multi-source high-resolution remote sensing data. The extraction results of the method of the present invention under different sensors have robust accuracy, overcome the problem of rapid attenuation of accuracy of other methods under different sensors, and solve the problems of internal fragmentation, wrong extraction, and missing extraction caused by sample generalization and image diversity.
[0096] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multi-source high-resolution remote sensing image vegetation extraction method based on multi-feature comprehensive perception, characterized in that: The steps include: Step 1: Sample set preparation; Step 2: Optimization of vegetation index features; Step 3: Network model structure design; Step 4: Multi-scale dense connection module; Step 5: Dual-path multi-head cross-attention feature fusion module; Step 6: Accuracy evaluation of model extraction results.
2. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 1 is characterized in that: The sub-steps of step one are as follows: 1) Obtain GF-2 remote sensing image data of the target area; 2) To better show the characteristics of vegetation, the remote sensing image was displayed in ArcGIS 10.7 software with 4-3-2 bands, the characteristics of vegetation were shown in red, and the vegetation was outlined through manual visual interpretation; 3) The preliminary delineation results are repeatedly revised, and samples are collected inside and on the edge of the delineated plot to obtain vegetation labels and background labels, and obtain sample coordinate points; 4) Considering the spatial correlation between adjacent pixels, a 5×5 neighborhood window is used to expand outward from the pixel of the sampling point as the center as the training sample to obtain more spatial and spectral information; Combining various data sources, vegetation contours were outlined through manual visual interpretation to obtain ground truth label data as verification data.
3. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 2 is characterized in that: The sub-steps of step 2 are as follows: 1) Taking vegetation index features as independent variables and sample labels as dependent variables, the random forest regression model was used to fit the nonlinear relationship between the two and perform feature optimization. The vegetation indices were selected as follows: NDVI, GNDVI, EVI, OSAVI, RVI, DVI, TVI, GVI, GI, NDGI, MCARI, and TCARI. The number of decision trees in the model was set to 20 and the depth was set to 4. 2) According to the fitting results of the random forest model, the values are sorted. The larger the value, the more important the vegetation index feature is to the accuracy of model prediction. The top four vegetation index features are selected.
4. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 3 is characterized in that: In step three, the input data of the model are remote sensing image spectral data and vegetation index. The cropped 5×5 samples are input into the parallel two-branch network respectively. The two networks contain three parallel sub-networks with scales of 1×1, 3×3, and 5×5 respectively. A simplified dense connection module is introduced to learn spectral and index features of different scales, and downsampling dense connections are used to perform interactive correction between features.
5. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 4 is characterized in that: In step 3, based on the self-attention of the visual transformer to obtain the maximum receptive field, a two-way cross-attention feature fusion module is designed between the 5×5 and 3×3 branches; The shallow initial features are cross-modally fused with the enhanced index features to establish multi-feature comprehensive perception to achieve global information complementarity, and the hybrid channel output is used to enhance the features.
6. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 5 is characterized in that: In step 3, a parallel dual-branch multi-scale network is designed, namely, a spectral multi-scale network branch and an exponential multi-scale network branch; coordinate convolution is introduced into the 5×5 and 3×3 branch networks respectively, and the convolution layer is used to optimize the extraction of spatial feature position information. At the same time, jump connections are designed inside the branch networks for extracting index features at different scales to enhance the fusion of vegetation index features at different scales and make the extraction of vegetation boundaries more complete; after continuous convolution, the dual-branch parallel sub-network outputs 1×1 spectral features and 1×1 exponential features respectively; Then, the spectral features and exponential features generated by parallel splicing of networks of different scales are merged; a deep separable convolutional layer DWCConv is designed 3×3 Finely extract features to reduce parameter calculations; Finally, after two fully connected layers and the Softmax classifier, the vegetation classification result is output.
7. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 6 is characterized in that: In step 4, in order to realize the feature information exchange between 5×5, 3×3 and 1×1, a multi-scale dense connection module is introduced between the spectral feature and exponential feature parallel dual-branch networks; in the multi-scale dense connection module, each sub-network can receive feature information from other parallel networks, so that the network can exchange information at different scales, thereby enhancing the feature expression ability; in the dense connection block used, the features of the previous layer are passed to all the subsequent layers in a stacked manner; In the 5×5 and 3×3 branch structures, vertical downsampling is performed to expand the width of the network; the features of the 5×5 branch are superimposed on the 3×3 and 1×1 branch networks through dense connections, and then the features of the 3×3 branch are superimposed on the 1×1 branch network; Each layer of the branch network is spliced with other branch networks in the channel dimension.
8. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 7 is characterized in that: In step 5, the input features F(H, W, C) are flattened to F(N, C) in this module, and then linear embedding is used to generate K, V, Q for each attention head. Each head calculates the attention score G of K and V by the multispectral branch and the vegetation index branch. MS and G VI ; Then, Softmax is used to normalize the attention score and the vegetation index feature G VI With spectral characteristics Q MS Cross-multiplication makes the vegetation index branch focus on the characteristics of the spectral branch; at the same time, the spectral feature G MS and vegetation index characteristic Q VI Cross-multiplication makes the spectral feature branch allocate attention to the features of the vegetation index branch; In the two sequence branches, branch complementary interaction is performed to achieve feature fusion; The calculation formula is as follows:
9. The method for extracting vegetation from multi-source high-resolution remote sensing images based on multi-feature comprehensive perception according to claim 8, characterized in that: In step 6, four evaluation indicators are used to evaluate the accuracy: overall accuracy OA, F1 score, intersection over union ratio and accuracy; The overall accuracy OA refers to the ratio of the sum of the number of pixels in the category correctly identified by the model to the pixels of all identified categories in the test image; The F1 score is the harmonic mean of precision and recall, and is used to evaluate the model's ability to handle imbalanced datasets or sensitivity to misclassified categories; IoU is the degree of overlap between the model prediction result and the true annotated area, which is used to evaluate the accuracy and robustness of the model; The Kappa coefficient for the remote sensing interpretation accuracy assessment estimates the consistency between the prediction and the ground reference; The formula for the above evaluation index is as follows: Among them, TP stands for true positive, FP stands for false positive, TN stands for true negative, FN stands for false negative, and N is the total number of samples; false positives represent the number of non-field pixels that are misclassified as positive, and false negatives represent positive pixels that are misclassified as negative.
Citation Information
Patent Citations
Urban vegetation unmanned aerial vehicle remote sensing classification method based on multi-scale feature sensing network
CN114943902A
Convolutional neural network and Transform fused remote sensing image building extraction method
CN116071650A
Network intrusion detection method based on feature selection and hybrid neural network
CN117478402A
High-resolution image-based ecological patch extraction method
WO2024020744A1