A wetland feature set optimization method that fuses improved filtering and packing strategies

By integrating a multi-criteria filtering algorithm and an encapsulated algorithm that combines Laplacian score, distance correlation coefficient, and JM distance, the problems of multi-source data redundancy and low classification efficiency in wetland information extraction are solved, thereby improving the speed and accuracy of feature selection algorithms and reducing redundancy.

CN116704275BActive Publication Date: 2026-03-31RICE TECH (HUBEI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-08
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies for wetland information extraction suffer from problems such as multi-source data redundancy and low classification efficiency. Filtering and encapsulation feature selection algorithms each have their own advantages and disadvantages, making it difficult to achieve a balance between speed and accuracy.

Method used

A multi-criteria fusion filtering algorithm (MCF-LDJ) is adopted, which combines Laplace score (LS), distance correlation coefficient (DC) and JM distance, and integrates filtering and encapsulation strategies. The MCF-LDJ algorithm is used to initially select the feature set, and the RFECV-RF algorithm is used to further optimize it, finally obtaining the optimal feature subset.

Benefits of technology

It improves the running speed and evaluation accuracy of the feature selection algorithm, reduces the redundancy of feature subsets, and maintains classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704275B_ABST
    Figure CN116704275B_ABST
Patent Text Reader

Abstract

The application discloses a wetland feature set optimization method fusing improved filtering and packaging strategies, and steps of the method comprise the following: according to the actual situation of a research area, sample selection and a classification system are completed; multi-source remote sensing data are selected, and corresponding pretreatment is carried out on each data; feature extraction is carried out on the pretreated data, and an original feature set is constructed; three single-criterion filtering algorithms, i.e., LS, DC and JM distance, are fused to obtain an MCF-LDJ algorithm, the original feature set is preliminarily selected based on the MCF-LDJ, and a preliminary selected feature set is obtained; the preliminary selected feature set obtained in step 4 is further optimized by using a packaging algorithm, and a final optimized feature subset is obtained; the application proposes a multi-criterion fusion filtering algorithm, so that the importance of calculated features is more reasonable. The complementarity of the filtering algorithm and the packaging algorithm is utilized to improve the operation speed and evaluation accuracy of the feature selection algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing and application, and specifically to a method for feature optimization of optical and SAR images, and more particularly to a method for wetland feature set optimization that integrates improved filtering and encapsulation strategies. Background Technology

[0002] Remote sensing imagery, with its advantages of large data acquisition range, periodic observation, and remote sensing, provides a rich data source for wetland observation. Due to the complexity of real-world land scenes, it is difficult to accurately represent the characteristics of land features using a single feature. Utilizing multi-source data can provide rich features for image interpretation, offering technical support for accurate description of wetland features. Optical imagery, with its advantages of high temporal resolution, ease of acquisition, and ease of interpretation, has become the main remote sensing data source for wetland classification. However, optical sensors cannot provide information on the physical characteristics of vegetation and are susceptible to weather conditions, making it difficult to continuously acquire effective optical imagery. Synthetic Aperture Radar (SAR) with its active imaging mode, featuring all-weather, all-time, and multi-polarization capabilities, can effectively supplement the limitations of optical imagery. Furthermore, the backscattering coefficient of SAR imagery can accurately reflect the characteristics and types of land features, making it particularly suitable for wetlands with complex land cover types. Therefore, the collaborative extraction of wetland information using optical and SAR remote sensing imagery has become one of the important research directions for wetland feature information extraction and analysis.

[0003] Collaborative classification of multi-source data can increase the number of available classification features, but problems such as data redundancy and low classification efficiency also arise. Dimensionality reduction techniques are an effective means to solve these problems. Currently, commonly used dimensionality reduction techniques include feature extraction and feature selection. In comparison, feature extraction methods often perform better in terms of prediction accuracy, but the time cost is too high. Considering both extraction efficiency and accuracy, feature selection is generally used instead of feature extraction for dimensionality reduction. Based on whether the feature selection process is combined with a classifier, it is divided into three types: filtering, encapsulation, and embedded. Filtering algorithms are independent of the learning algorithm, have high computational efficiency, and strong versatility, but the selected feature subsets contain many highly correlated features. Encapsulation algorithms evaluate the classification accuracy of different feature subset combinations in the original feature space by combining various classifiers, and select the optimal feature subset. This method performs well in classification performance, but its efficiency is low. Embedded algorithms are a compromise between filtering and encapsulation algorithms. The performance of the selected feature subset is only excellent for itself, and it is prone to overfitting. Filter-based feature selection algorithms based on a single evaluation criterion are highly efficient, but their robustness is poor because they can only evaluate the merits of features from one perspective. Encapsulated feature selection algorithms offer higher accuracy in assessing feature importance, but are less time-efficient. Summary of the Invention

[0004] To address the aforementioned technical shortcomings, the purpose of this invention is to provide a wetland feature set optimization method that integrates improved filtering and encapsulation strategies. This method considers three aspects: feature locality preservation, feature correlation, and inter-class separability. It selects three single evaluation criteria—Laplacian score (LS), distance correlation coefficient (DC), and Jeffreys-Matusita distance—for fusion, proposing a multi-criterion filter algorithm (MCF-LDJ) to make the calculated feature importance more reasonable. The complementarity between the filtering and encapsulation algorithms improves the running speed and evaluation accuracy of the feature selection algorithm.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0006] This invention provides a wetland feature set optimization method that integrates improved filtering and encapsulation strategies, comprising the following steps:

[0007] Step 1: Based on the actual conditions of the study area, complete the sample selection and establish a classification system;

[0008] Step 2: Select multi-source remote sensing data and perform corresponding preprocessing on each data source;

[0009] Step 3: Extract features from the preprocessed data and construct the original feature set;

[0010] Step 4: Perform preliminary selection on the original feature set based on MCF-LDJ to obtain the initial selected feature set;

[0011] Step 5: The initial feature set obtained in Step 4 is further optimized using an encapsulation algorithm to obtain the final optimized feature subset.

[0012] Preferably, in step 1, based on wetland classification standards such as the Ramsar Convention on Wetlands and the National Technical Regulations for Wetland Resources Survey and Monitoring, and combined with the actual conditions of the study area and previous studies on the Yellow River Delta wetlands, the wetland types in the study area are divided into eight categories: mudflats, Yellow River wetlands, seawater wetlands, Spartina alterniflora, reeds, buildings, Suaeda salsa, and ponds. This paper selects training and test samples based on the principles of randomness and uniformity.

[0013] Preferably, the preprocessing in step 2 includes: radiometric calibration, atmospheric correction, and cropping of optical images; radiometric calibration, complex data conversion, speckle filtering, geocoding, and registration with optical images of SAR images.

[0014] Preferably, step 3 extracts 83-dimensional features, including: based on GF-3 fully polarimetric SAR imagery, original statistical features such as backscattering coefficient, polarization ratio, C3 parameter, T3 parameter, and radar vegetation index are extracted, as well as polarization decomposition features and polarization texture features, totaling 57-dimensional features, as shown in Table 1. Based on Sentinel-2 multispectral imagery, band features, exponential features, and texture features, totaling 26-dimensional features are extracted, as shown in Table 2.

[0015] Table 1 Ranking of Multispectral Image Features

[0016]

[0017] Table 2 Ranking of SAR Image Features

[0018]

[0019]

[0020] Preferably, the MCF-LDJ algorithm in step 4 is a fusion of three single-criteria filtering algorithms: LS, DC, and JM distance; specifically:

[0021] (1) LS algorithm

[0022] The LS algorithm places greater emphasis on the local preservation of features. Assume the sample set has T types, M samples, and each sample has N features. ri Let t be the r-th feature of the i-th sample. i Let be the type of the i-th sample. A nearest neighbor graph G can be constructed using the given samples. x = (V,E), the M×M weight matrix S of graph G is defined as follows:

[0023]

[0024] In the formula, t is a constant.

[0025] The Laplace score of the r-th feature is:

[0026]

[0027] In the formula, D = diag(SI).

[0028] (2) DC algorithm

[0029] The DC algorithm is constructed from the difference between the product of the joint characteristic function and the product of their respective marginal characteristic functions. It can describe the linear and nonlinear relationships between two variables. The distance covariance between vectors X and Y is defined as follows:

[0030]

[0031]

[0032] In the formula, f X f Y These are the characteristic functions corresponding to X and Y, respectively; f X,Y It is the joint characteristic function of X and Y.

[0033] Similarly, the distance variance V(X) is defined as:

[0034] V 2 (X)=V 2 (X,X)=||f X,Y (t,s)-f X (t)f X (s)|| 2

[0035] Therefore, the distance correlation coefficient R(X,Y) can be defined as:

[0036]

[0037] The range of R(X,Y) is [0, 1]. When R(X,Y) = 0, X and Y are independent. As the value of R(X,Y) increases, it indicates that X and Y are more correlated. For multi-feature classification, the distance correlation coefficient between each feature and other features is calculated first, and then the coefficients are summed and averaged. This average value is used as the correlation value of each feature.

[0038] (3) JM distance

[0039] The JM distance is a feature selection criterion for evaluating inter-class separability. This type of method quantitatively analyzes the classification performance of the features used, aiming to select the feature set that minimizes the classifier's error probability. The JM distance comprehensively considers the impact of class overlap on class distinction and the probability distribution of each class; therefore, it has a more reliable ability to determine class separability.

[0040] The formula for calculating the JM distance between the two categories is as follows:

[0041]

[0042] In the formula, B i,j M is the Bach distance. i and M j V are the sample mean vectors of categories i and j, respectively. i and V j Let J and M be the covariance matrices for categories i and j, respectively. When the number of categories is 2, the JM distance ranges from [0, 2]. When JM... i,j When JM = 0, it means that the two categories cannot be distinguished at all under this feature; when JM i,jWhen the value is 2, it means that the two categories can be completely distinguished under this feature. The larger the JM distance value, the higher the separability between the categories.

[0043] In remote sensing image classification, there are usually more than two types of land features. For multi-class land cover classification, the JM distance calculation formula is to take the average of the JM distances between the two classes. The calculation formula is as follows:

[0044]

[0045] In the formula, c is the number of categories; P(ω) i ), P(ω) j ) are the probability densities of categories i and j, respectively.

[0046] (4) MCF-LDJ

[0047] To comprehensively evaluate the importance of each feature from three aspects—inter-class separability, feature correlation, and feature local preservation ability—this invention fuses the above three algorithms based on the following formula:

[0048]

[0049] In the formula, α is a weighting factor, taking values ​​of [0, 1]. When the value is 0.6, S new To achieve optimal accuracy, S new A larger value indicates that the feature is the optimal representation feature; L r The LS value is the feature. R(X,Y) is the JM distance of the feature; R(X,Y) is the DC value of the feature.

[0050] Based on the above formula, sort the weight values ​​of all features in descending order, and add each feature to the random forest in turn for classification experiments until the classification accuracy no longer improves when new features are added. These added features are the initial feature subset.

[0051] Preferably, step 5 employs a combined algorithm (RFECV-RF) that integrates the cross-validation-based recursive feature elimination algorithm (RFECV) and Random Forest (RF). RFECV is a greedy optimization algorithm that searches for the optimal feature set and is also a type of combined algorithm. This algorithm consists of two steps: Recursive Feature Elimination (RFE) and Cross Validation (CV). RFE is primarily used to evaluate the importance of features; its main idea is to repeatedly train the selected model to search for the optimal feature subset. CV, after ranking the features, selects the optimal feature set that maximizes the average cross-validation accuracy based on the cross-validation accuracy. RFECV includes the following steps:

[0052] (1) Train the model using the original feature set;

[0053] (2) Calculate and evaluate the importance of each feature variable;

[0054] (3) Remove less important feature variables based on the results of cross-validation;

[0055] (4) Repeat the above steps until the preset feature subset dimension is reached.

[0056] Beneficial effects: The method proposed in this invention can significantly reduce the dimensionality of feature subsets and decrease feature redundancy while maintaining the classification accuracy of feature subsets. This invention addresses three aspects: feature locality preservation, feature relevance, and inter-class separability. It selects three single evaluation criteria—Laplacian score (LS), distance correlation coefficient (DC), and Jeffreys-Matusita distance—for fusion, proposing a multi-criterion filter algorithm (MCF-LDJ) to make the calculated feature importance more reasonable. This invention utilizes the complementarity of the filtering algorithm and the encapsulation algorithm to improve the running speed and evaluation accuracy of the feature selection algorithm. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 A flowchart of a wetland feature set optimization method that integrates filtering and encapsulation strategies provided in an embodiment of the present invention;

[0059] Figure 2 (a) is a map showing the distribution of land cover categories provided in an embodiment of the present invention;

[0060] Figure 2 (b) is a distribution diagram of the classification samples provided in an embodiment of the present invention;

[0061] Figure 3 (a) shows the classification results of the feature subset selected by the LS algorithm provided in this embodiment of the invention;

[0062] Figure 3 (b) shows the classification results of the feature subset selected by the DC algorithm provided in this embodiment of the invention;

[0063] Figure 3 (c) shows the classification results of the selected feature subset based on the JM distance provided in this embodiment of the invention;

[0064] Figure 3 (d) represents the classification result of the feature subset selected by the MCF-LDJ algorithm provided in this embodiment of the invention;

[0065] Figure 3 (e) represents the classification result of the feature subset selected by the RFECV-RF algorithm provided in this embodiment of the invention;

[0066] Figure 3 In (f), the classification result of the feature subset selected by the hybrid feature selection algorithm provided in the embodiment of the present invention is shown.

[0067] Figure 3 In the example (g), the classification result of the original feature set provided in the embodiment of the present invention is shown. Detailed Implementation

[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0069] A method for optimizing wetland feature sets, integrating improved filtering and encapsulation strategies, is named the Hybrid Feature Selection Algorithm. It includes the following steps:

[0070] Step 1: Based on the actual conditions of the study area, complete the sample selection and establish a classification system;

[0071] Based on wetland classification standards such as the Ramsar Convention on Wetlands and the National Technical Regulations for Wetland Resources Survey and Monitoring, and combined with the actual conditions of the study area and previous studies on the Yellow River Delta wetlands, the wetland types in the study area are divided into eight categories: mudflats, Yellow River wetlands, seawater wetlands, Spartina alterniflora, reeds, buildings, Suaeda salsa, and ponds. This paper selects training and test samples based on the principles of randomness and uniformity, such as... Figure 2 As shown. Figure 2 (a) in the image is a map showing the distribution of land cover categories. Figure 2 (b) in the figure is a distribution map of the classification samples.

[0072] Step 2: Select multi-source remote sensing data and perform corresponding preprocessing on each data source;

[0073] The images used were preprocessed using software such as ENVI and Gamma, including: radiometric calibration, atmospheric correction, and cropping of optical images; radiometric calibration, complex data conversion, speckle filtering, geocoding, and registration with optical images of SAR images.

[0074] Step 3: Extract features from the preprocessed data and construct the original feature set;

[0075] Step 3 extracts a total of 83 original features, including: based on GF-3 fully polarimetric SAR imagery, backscattering coefficient, polarization ratio, C3 parameter, T3 parameter, radar vegetation index and other original statistical features, as well as polarization decomposition features and polarization texture features, totaling 57 features, as shown in Table 1; based on Sentinel-2 multispectral imagery, band features, exponential features and texture features, totaling 26 features, as shown in Table 2.

[0076] Table 1 Ranking of Multispectral Image Features

[0077]

[0078]

[0079] Table 2 Ranking of SAR Image Features

[0080]

[0081] Step 4: Perform preliminary selection on the original feature set based on MCF-LDJ to obtain the initial selected feature set;

[0082] The MCF-LDJ algorithm in step 4 is a fusion of three single-criteria filtering algorithms: LS, DC, and JM distance.

[0083] (1) LS algorithm

[0084] The LS algorithm places greater emphasis on the local preservation of features. Assume the sample set has T types, M samples, and each sample has N features. ri Let t be the r-th feature of the i-th sample. i Let be the type of the i-th sample. A nearest neighbor graph G can be constructed using the given samples. x = (V,E), the M×M weight matrix S of graph G is defined as follows:

[0085]

[0086] In the formula, t is a constant.

[0087] The Laplace score of the r-th feature is:

[0088]

[0089] In the formula, D = diag(SI).

[0090] (2) DC algorithm

[0091] The DC algorithm is constructed from the difference between the product of the joint characteristic function and the product of their respective marginal characteristic functions. It can describe the linear and nonlinear relationships between two variables. The distance covariance between vectors X and Y is defined as follows:

[0092]

[0093] In the formula, f X f Y These are the characteristic functions corresponding to X and Y, respectively; f X,Y It is the joint characteristic function of X and Y.

[0094] Similarly, the distance variance V(X) is defined as:

[0095] v 2 (X)=V 2 (X,X)=||f X,Y (t,s)-f X (t)f X (s)|| 2

[0096] Therefore, the distance correlation coefficient R(X,Y) can be defined as:

[0097]

[0098] The range of R(X,Y) is [0, 1]. When R(X,Y) = 0, X and Y are independent. As the value of R(X,Y) increases, it indicates that X and Y are more correlated. For multi-feature classification, the distance correlation coefficient between each feature and other features is calculated first, and then the coefficients are summed and averaged. This average value is used as the correlation value of each feature.

[0099] (3) JM distance

[0100] The JM distance is a feature selection criterion for evaluating inter-class separability. This type of method quantitatively analyzes the classification performance of the features used, aiming to select the feature set that minimizes the classifier's error probability. The JM distance comprehensively considers the impact of class overlap on class distinction and the probability distribution of each class; therefore, it has a more reliable ability to determine class separability.

[0101] The formula for calculating the JM distance between the two categories is as follows:

[0102]

[0103] In the formula, B i,j M is the Bach distance. i and M j V are the sample mean vectors of categories i and j, respectively. i and V j Let J and M be the covariance matrices for categories i and j, respectively. When the number of categories is 2, the JM distance ranges from [0, 2]. When JM... i,j When JM = 0, it means that the two categories cannot be distinguished at all under this feature; when JM i,j When the value is 2, it means that the two categories can be completely distinguished under this feature. The larger the JM distance value, the higher the separability between the categories.

[0104] In remote sensing image classification, there are usually more than two types of land features. For multi-class land cover classification, the JM distance calculation formula is to take the average of the JM distances between the two classes. The calculation formula is as follows:

[0105]

[0106] In the formula, c is the number of categories; P(ω) i ), P(ω) j ) are the probability densities of categories i and j, respectively.

[0107] (4) MCF-LDJ

[0108] To comprehensively evaluate the importance of each feature from three aspects—inter-class separability, feature correlation, and feature local preservation ability—this invention fuses the above three algorithms based on the following formula:

[0109]

[0110] In the formula, α is a weighting factor, taking values ​​of [0, 1]. When the value is 0.6, S new To achieve optimal accuracy, S new A larger value indicates that the feature is the optimal representation feature; L r The LS value is the feature. R(X,Y) is the JM distance of the feature; R(X,Y) is the DC value of the feature.

[0111] Based on the above formula, sort the weight values ​​of all features in descending order, and add each feature to the random forest in turn for classification experiments until the classification accuracy no longer improves when new features are added. These added features are the initial feature subset.

[0112] Step 5: The initial feature set obtained in Step 4 is further optimized using an encapsulation algorithm to obtain the final optimized feature subset.

[0113] We employ a wrapper algorithm (RFECV-RF) that combines the cross-validation-based recursive feature elimination algorithm (RFECV) with random forest (RF). RFECV is a greedy optimization algorithm that searches for the optimal feature set and is also a type of wrapper algorithm. This algorithm consists of two steps: recursive feature elimination (RFE) and cross-validation (CV).

[0114] RFE (Reference-Free Evaluation) is primarily used to evaluate the importance of features. The main idea of ​​this algorithm is to repeatedly train the selected model and search for the optimal feature subset. CV (Cross-Validation) selects the best feature set that maximizes the average cross-validation accuracy after feature ranking. RFECV includes the following steps:

[0115] (1) Train the model using the original feature set;

[0116] (2) Calculate and evaluate the importance of each feature variable;

[0117] (3) Remove less important feature variables based on the results of cross-validation;

[0118] (4) Repeat the above steps until the preset feature subset dimension is reached.

[0119] Taking the Yellow River Delta wetlands as the study area, wetland features were extracted based on GF-3 fully polarimetric SAR data from November 5, 2018, and Sentinel-2A multispectral data from November 3, 2018. The original feature set was optimized using LS, DC, JM distances, the MCF-LDJ algorithm proposed in this invention, the RFECV-RF encapsulated feature selection algorithm, and the hybrid feature selection algorithm proposed in this invention. Figure 3 In the table, (a) represents the classification result of the feature subset selected by the LS algorithm; Figure 3 In the diagram, (b) represents the classification result of the feature subset selected by the DC algorithm; Figure 3 In the diagram, (c) represents the classification result of the selected feature subset based on the JM distance; Figure 3 In the table, (d) represents the classification result of the feature subset selected by the MCF-LDJ algorithm; Figure 3 In the table, (e) represents the classification result of the feature subset selected by the RFECV-RF algorithm; Figure 3 In the equation (f), the classification result of the feature subset selected by the hybrid feature selection algorithm is shown. Figure 3 In the original feature set, (g) represents the classification result.

[0120] (2) Accuracy Evaluation

[0121] To quantitatively evaluate the effectiveness of this invention, the Kappa coefficient, overall laccuracy (OA), and subset dimension are selected as evaluation indicators.

[0122] Table 36 Selection Results of Feature Selection Algorithms

[0123]

[0124]

[0125] (3) Analysis of experimental results

[0126] Experimental results demonstrate that the optimal feature subset obtained by the MCF-LDJ algorithm outperforms the three single-criteria filtering algorithms in terms of subset dimension, OA, and Kappa coefficient. The optimal feature subset obtained by the RFECV-RF algorithm outperforms the MCF-LDJ algorithm in both OA and Kappa coefficient, while maintaining the same subset dimension. The optimal feature subset obtained by the hybrid feature selection algorithm outperforms the MCF-LDJ algorithm in subset dimension, OA, and Kappa coefficient; its OA and Kappa coefficients are slightly lower than those of the RFECV-RF algorithm, but its subset dimension is reduced by 24 dimensions compared to the RFECV-RF algorithm. Therefore, the method proposed in this invention can significantly reduce the dimensionality of feature subsets and decrease feature redundancy while maintaining the accuracy of feature subset classification.

[0127] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A wetland feature set optimization method that blends improved filtering and enclosure strategies, characterized by, Comprising the following steps: Step 1: According to the actual situation of the study area, complete sample selection and establish a classification system; Step 2: Select multi-source remote sensing data, and perform corresponding preprocessing on each data; Step 3: Feature extraction is performed on the preprocessed data, and an original feature set is constructed; Step 4: Fuse the three single-criterion filtering algorithms of LS, DC and JM distance to obtain the MCF-LDJ algorithm, and perform preliminary selection on the original feature set based on MCF-LDJ to obtain the initial feature set; Step 5: The initial feature set obtained in step 4 is further optimized using the encapsulation algorithm to obtain the final optimized feature subset; The MCF-LDJ algorithm of step 4 is to fuse the three single-criterion filtering algorithms of LS, DC and JM distance; Specifically: (1) LS algorithm The LS algorithm pays more attention to the local preservation ability of features, assuming that the sample set has T types, M samples, and each sample has N features; f ri is the rth feature of the ith sample, t i is the type to which the ith sample belongs; using the given samples, a nearest neighbor graph G x =(V,E) can be constructed, and the MxM order weight matrix S of the graph G is defined as follows: In the formula, t is a constant; The Laplace score of the rth feature quantity is: In the formula, D = diag(SI); (2) DC algorithm DC algorithm is constructed by the difference between the joint characteristic function of two variables and the product of their marginal characteristic functions, which can describe the linear and nonlinear relationship between two variables. The distance covariance between vector X and vector Y is defined as follows: where f X , f Y are the characteristic functions of X, Y respectively; f X,Y is the joint characteristic function of X, Y. Similarly, the distance variance V(X) is defined as: V 2 (X) = V 2 (X,X) = ||f X,Y (t,s)-f X (t)f X (s)| 2 Therefore, the distance correlation coefficient R(X,Y) is defined as: The value range of R(X,Y) is [0, 1], when R(X,Y) = 0, X and Y are independent, and as the value of R(X,Y) increases, the correlation between X and Y becomes stronger; For multi-feature classification, first calculate the distance correlation coefficient between each feature and other features, and then take the average of the sum as the correlation value of each feature; (3) JM distance JM distance belongs to the feature selection evaluation criterion of class separability. This kind of method quantitatively analyzes the classification performance of the used features, and the purpose is to select the feature set with the minimum error probability of the classifier. JM distance comprehensively considers the influence of class overlap on class distinction and the probability distribution of each class, so JM distance has more reliable class separability discrimination ability; The calculation formula of JM distance between two classes is as follows: where B i,j is the Bhattacharyya distance, M i and M j are the sample mean vectors of classes i and j, respectively, V i and V j are the covariance matrices of classes i and j, respectively; when the number of classes is 2, the value range of JM distance is [0, 2]; when JM i,j = 0, it means that the two classes cannot be distinguished at all under this feature; when JM i,j = 2, it means that the two classes can be completely distinguished under this feature; the greater the value of JM distance, the higher the separability between classes. In remote sensing image classification, the types of ground objects are more than two. For multi-class JM distance, the average value of JM distance between two classes is taken, and the calculation formula is as follows: where c is the number of classes; P(ω i ), P(ω j ) are the probability densities of classes i and j, respectively. (4) MCF-LDJ In order to comprehensively evaluate the importance of each feature from the aspects of class separability, feature correlation and local preservation ability of feature, the above three algorithms are fused according to the following formula: In the formula, a is a weight factor, and a value of [0, 1], when the value is 0.6, S new The best precision can be obtained, S new The greater the value represents that the feature is the optimal expression feature; L r The LS value of the feature; The JM distance of the feature; R(X, Y) is the DC value of the feature; According to the above formula, the weight values of all features are sorted in descending order, and each feature is added to the random forest for classification test in turn until the classification accuracy does not improve by adding new features. These added features are the initial feature subset.

2. A hybrid improved filtering and packaging policy wetland feature set optimization method as described in claim 1, wherein: In step 1, the wetland types in the study area are divided into 8 categories, including beach, Yellow River, seawater, Spartina alterniflora, reed, building, salt marsh, and pond. Training samples and test samples are selected according to the principles of randomness and uniformity.

3. A hybrid improved filtering and packaging policy wetland feature set optimization method as described in claim 2, wherein: The preprocessing process in step 2 includes: radiometric calibration, atmospheric correction, cropping of optical images; radiometric calibration, complex data conversion, speckle filtering, geocoding, and registration with optical images of SAR images.

4. A hybrid improved filtering and packaging policy wetland feature set optimization method as described in claim 3, wherein: The features extracted in step 3 include: backscattering coefficient, polarization ratio, C3 parameter, T3 parameter, radar vegetation index, polarization decomposition feature and polarization texture feature are extracted based on GF-3 full polarization SAR image; band feature, index feature and texture feature are extracted based on Sentinel-2 multispectral image.

5. A wetland feature set optimization method that integrates improved filtering and packaging strategies as recited in claim 1, characterized by: Step 5 adopts the encapsulated algorithm of fusing the recursive elimination algorithm based on cross-validation and the random forest, the recursive elimination algorithm is a kind of greedy algorithm for searching the optimal feature set, and is also a kind of encapsulated algorithm, the recursive elimination algorithm includes two steps: recursive feature elimination and cross-validation; wherein the recursive feature elimination is used for importance evaluation of the features, the main idea of the algorithm is to repeatedly train the selected model to search the optimal feature subset; The cross-validation is to select different numbers of features through the feature importance rating, calculate the cross-validation accuracy of each feature set, and select the optimal feature set which makes the average cross-validation accuracy reach the maximum; the recursive elimination algorithm includes the following steps: (1) the model is trained using the original feature set; (2) the importance of each feature variable is calculated and evaluated; (3) the feature variables with lower importance are removed based on the cross-validation result; (4) the above steps are repeated until the preset feature subset dimension is reached.

Citation Information

Patent Citations

  • Wetland vegetation feature optimization and fusion method based on JM Relief F

    CN112949607A

  • Feature selection method for geographic object image analysis

    CN115205528A