A bispectrum adaptive threshold feature selection method based on random forest

By adopting a bispectral adaptive threshold feature selection method based on random forest, this paper solves a number of problems in feature selection in weather forecasting, achieves efficient and stable feature selection and accurate weather forecasting, supports cross-regional adaptability and feature interaction effect identification, and improves the accuracy and efficiency of weather forecasting.

CN121009791BActive Publication Date: 2026-04-24INNER MONGOLIA AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INNER MONGOLIA AGRICULTURAL UNIVERSITY
Filing Date
2025-08-21
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing feature selection methods in the field of meteorological forecasting suffer from limitations such as the assumption of linear relationships, subjectivity in threshold determination, one-sidedness of single evaluation criteria, difficulty in balancing explained variance and model complexity, insufficient dynamic adaptability, and neglect of feature interaction effects, resulting in unstable and inefficient feature selection results.

Method used

A bispectral adaptive threshold feature selection method based on random forest is adopted. The feature importance score is obtained through the random forest model, the gradient spectrum and cumulative importance contribution spectrum are calculated, and the optimal feature is determined by combining the optimized score function, so as to realize adaptive threshold determination and feature interaction effect identification.

Benefits of technology

It achieves efficient and stable feature selection, improves the accuracy of weather forecasts and the generalization ability of models, reduces computational resource consumption, supports cross-regional adaptability and the identification of feature interaction effects, and provides clear visualization and interpretation of feature importance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009791B_ABST
    Figure CN121009791B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of double spectrum adaptive threshold feature selection methods based on random forest, comprising: obtaining feature importance score by random forest model;Based on feature importance score, the first derivative of feature importance is calculated, gradient spectrum is constructed, and threshold candidate point of gradient spectrum is obtained;The cumulative importance contribution spectrum of feature is constructed, and the threshold candidate point of cumulative importance contribution spectrum is obtained;The threshold candidate point of gradient spectrum and cumulative importance contribution spectrum is integrated, each candidate threshold point is evaluated by optimization score function, and the optimal feature is determined.The present application first applies double spectrum feature selection in the field of rainfall simulation systematically, realizes the intelligent feature screening of high-dimensional meteorological data, and provides technical support for refined weather forecast;The present application can be directly integrated into existing weather forecast business system, improve the prediction accuracy through accurate feature selection, reduce disaster loss, reduce the consumption of computing resources, save operation cost, and improve the efficiency of model deployment and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rainfall simulation and weather forecasting technology, and in particular to a bispectral adaptive threshold feature selection method based on random forest. Background Technology

[0002] In the fields of rainfall simulation and weather forecasting, feature selection is a crucial step in improving model performance. Existing feature selection methods face the following technical challenges:

[0003] Limitations of the linear relationship assumption: Traditional Pearson correlation coefficient analysis, based on the linear relationship assumption, has significant limitations in handling complex nonlinear coupling relationships between meteorological elements. Meteorological variables such as temperature, humidity, pressure, and cloud cover often exhibit highly nonlinear interactions, and simple linear correlation analysis cannot fully uncover these implicit complex relationship structures.

[0004] Subjectivity in threshold determination: Existing methods often rely on human experience or fixed proportions when determining feature selection thresholds, lacking an objective data-driven mechanism. For example, commonly used strategies such as "selecting the top k features" or "relevance greater than a certain fixed value" are difficult to adapt to the differences in characteristics of different datasets, leading to instability and subjectivity in feature selection results.

[0005] The limitations of single evaluation criteria: While mutual information methods can capture nonlinear relationships, they suffer from high computational complexity and sensitivity to discretization parameters; Lasso regression is prone to over-sparseness, potentially discarding important features incorrectly; and the SHAP value method is computationally expensive and inefficient in high-dimensional feature spaces. These methods are often based on a single evaluation criterion, making it difficult to comprehensively consider the multidimensional importance of features.

[0006] Balancing interpretability and model complexity is challenging: existing methods lack effective mechanisms to balance the interpretability of a model with its computational complexity. Too many features can lead to the curse of dimensionality and overfitting, while too few features may result in the loss of important information. Finding the optimal balance between the two has been a long-standing technical challenge.

[0007] Insufficient dynamic adaptability: Most existing feature selection methods are static and lack the ability to adapt to changes in data characteristics. When faced with rainfall data from different regions, seasons, or climatic conditions, fixed feature selection strategies often fail to maintain optimal performance.

[0008] Feature interaction effects are often overlooked: Traditional methods tend to assess the importance of each feature independently, neglecting the interaction and combination effects between features. In meteorological systems, multiple seemingly unimportant features can produce significant predictive power when combined, but existing methods struggle to capture these complex feature interaction patterns.

[0009] With the deepening application of big data and artificial intelligence technologies in the meteorological field, there is an urgent need to develop feature selection methods that can: 1. automatically capture complex nonlinear relationships between features; 2. provide data-driven adaptive threshold determination mechanisms; 3. comprehensively consider feature importance assessment frameworks with multiple evaluation criteria; 4. effectively balance model interpretability and computational complexity; 5. possess good generalization ability and dynamic adaptation; and 6. be able to identify and utilize feature interaction effects. Summary of the Invention

[0010] The purpose of this invention is to propose a bispectral adaptive threshold feature selection method based on random forest to solve the problems existing in the prior art.

[0011] To achieve the above objectives, the present invention provides the following solution:

[0012] A bispectral adaptive threshold feature selection method based on random forest includes:

[0013] Feature importance scores are obtained using a random forest model;

[0014] Based on the feature importance score, the first derivative of the feature importance is calculated, a gradient spectrum is constructed, and the threshold candidate points of the gradient spectrum are obtained.

[0015] Construct the cumulative importance contribution spectrum of the features and obtain the threshold candidate points of the cumulative importance contribution spectrum;

[0016] By combining the threshold candidate points of the gradient spectrum and the cumulative importance contribution spectrum, the optimal feature is determined by evaluating each candidate threshold point through an optimized fractional function.

[0017] Optionally, the feature importance score is:

[0018]

[0019] in, This represents the feature importance vectors arranged in descending order of importance. This represents the total number of features.

[0020] Optionally, the gradient spectrum is:

[0021]

[0022] in, Represents the gradient spectrum. This represents the importance difference between adjacent features, where i represents any feature from the 1st to the (n-1st)th feature.

[0023] Optionally, obtaining the threshold candidate points of the gradient spectrum includes:

[0024] The gradient spectrum is normalized.

[0025] The local extremum detection algorithm is used to identify preset change points in the normalized gradient spectrum; where the preset change points represent significant breakpoints in feature importance and indicate potential feature selection threshold positions.

[0026] Optionally, the cumulative importance contribution spectrum is:

[0027]

[0028] in, Indicates the preceding The cumulative contribution ratio of each feature.

[0029] Optionally, the threshold candidate points for obtaining the cumulative importance contribution spectrum include:

[0030] The elbow rule is applied to detect the inflection point of the cumulative contribution curve in the cumulative importance contribution spectrum, and the inflection point is regarded as a threshold candidate point.

[0031] Optionally, the optimized score function is:

[0032]

[0033] in, The proportion of variance explained by the representative The representative feature is the proportion of the penalty term. Represents the score of the candidate threshold point.

[0034] Optionally, features in rainfall simulation and weather forecasting include:

[0035] Near-infrared water vapor, cloud water path, cloud optical thickness, cloud effective radius, cloud phase infrared, cloud top temperature, cloud top pressure, cloud top height, surface temperature, total evaporation, bare soil evaporation, forest canopy top evaporation, vegetation transpiration, potential evaporation, downward surface solar radiation, downward surface thermal radiation, net surface solar radiation, net surface thermal radiation, surface latent heat flux, surface sensible heat flux, predicted albedo, 2-meter air temperature, 2-meter dew point temperature, surface skin temperature, 10-meter height U-component wind speed, 10-meter height V-component wind speed, high vegetation leaf area index, low vegetation leaf area index, volumetric water content.

[0036] Optionally, the optimal feature includes:

[0037] Evaporation at the top of the canopy, net solar radiation at the ground surface, net thermal radiation at the ground surface, downward solar radiation at the ground surface, surface temperature, wind speed of the U component at a height of 10 meters, evaporation of bare soil, downward thermal radiation at the ground surface, wind speed of the V component at a height of 10 meters, volumetric water content of the first soil layer, and dew point temperature at 2 meters.

[0038] The beneficial effects of this invention are as follows:

[0039] This invention proposes for the first time a theoretical framework for feature selection through bispectral fusion, innovatively combining gradient analysis and elbow detection; it designs a fully data-driven adaptive threshold determination mechanism and develops an efficient bispectral analysis algorithm; this invention is the first to systematically apply bispectral feature selection in the field of rainfall simulation, realizing intelligent feature screening of high-dimensional meteorological data and providing technical support for refined weather forecasting; this invention can be directly integrated into existing meteorological forecasting operational systems, improving forecast accuracy, reducing disaster losses, lowering computational resource consumption, saving operating costs, and improving the efficiency of model deployment and maintenance through precise feature selection. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a schematic diagram of a bispectral adaptive threshold feature selection method based on random forest according to an embodiment of the present invention;

[0042] Figure 2 This is a schematic diagram of the feature importance scores in an embodiment of the present invention;

[0043] Figure 3 This is a schematic diagram illustrating the rate of change and ranking of feature importance in an embodiment of the present invention;

[0044] Figure 4 This is a schematic diagram of the cumulative importance curve of features in an embodiment of the present invention;

[0045] Figure 5 This is a schematic diagram illustrating the feature importance ranking and threshold selection based on K-means clustering in an embodiment of the present invention;

[0046] Figure 6 This is a schematic diagram of the bispectral adaptive threshold feature selection result based on random forest in an embodiment of the present invention;

[0047] Figure 7 This is a schematic diagram illustrating the sorting results of features using five methods according to embodiments of the present invention;

[0048] Figure 8 This is a schematic diagram of an ablation experiment for a specific number of features according to an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] like Figure 1 As shown, this embodiment proposes a bispectral adaptive threshold feature selection method based on random forest, including:

[0052] Feature importance scores in rainfall simulation and weather forecasting were obtained using a random forest model.

[0053] Based on the feature importance score, the first derivative of the feature importance is calculated, a gradient spectrum is constructed, and the threshold candidate points of the gradient spectrum are obtained.

[0054] Construct the cumulative importance contribution spectrum of the features and obtain the threshold candidate points of the cumulative importance contribution spectrum;

[0055] By combining the threshold candidate points of the gradient spectrum and the cumulative importance contribution spectrum, the optimal feature is determined by evaluating each candidate threshold point through an optimized fractional function.

[0056] Specifically, in this embodiment, a bispectral adaptive threshold feature selection method (RF-DSAT) based on random forest is proposed. The core technical architecture includes: a feature importance calculation module: calculating feature importance scores based on the random forest algorithm; a bispectral construction module: constructing gradient spectrum and cumulative contribution spectrum; a threshold candidate point identification module: identifying key inflection points in the spectrum; an adaptive threshold optimization module: determining the optimal threshold feature subset by optimizing the objective function; and a feature subset determination module: outputting the optimal feature subset.

[0057] In this embodiment, the features in rainfall simulation and weather forecasting include: wvni: near-infrared water vapor, cwp: cloud water path, cot: cloud optical thickness, cer: effective cloud radius, cpi: cloud phase infrared, ctt: cloud top temperature, ctp: cloud top pressure, cth: cloud top height, st: surface temperature, eva: total evaporation, efbs: bare soil evaporation, efttoc: canopy top evaporation, efvt: vegetation transpiration, peva: potential evaporation, ssrd: downward solar radiation to the surface, strd : Downward surface thermal radiation, snsr: Net solar radiation at the surface, sntr: Net thermal radiation at the surface, slhf: Latent heat flux at the surface, sshf: Sensible heat flux at the surface, fa: Predicted albedo, t2: Air temperature at 2 meters, t2d: Dew point temperature at 2 meters, ts: Surface skin temperature, u10: U-component wind speed at 10 meters altitude, v10: V-component wind speed at 10 meters altitude, laihv: Leaf area index of high vegetation, lailv: Leaf area index of low vegetation, vswl1: Volumetric water content (first soil layer, 0-10cm).

[0058] In this embodiment, the uses of determining the optimal threshold position include:

[0059] Feature selection: By selecting features that contribute more to model training and removing features that contribute less, the system reduces redundant data, improves computational efficiency, and enhances the model's generalization ability by using the optimal threshold position.

[0060] Dimensionality Reduction: A threshold can be used as a dividing point for the number or importance of features to be retained, reducing feature dimensionality, reducing data noise, and retaining key information.

[0061] Furthermore, obtaining feature importance scores through the random forest model includes:

[0062] Random Forest Model Construction:

[0063] Number of decision trees: Set to 1000 to ensure the stability of feature importance estimation;

[0064] Maximum depth: No limit, allowing the tree to grow fully to capture feature relationships;

[0065] Minimum number of split samples: set to 2 to maintain the model's learning ability;

[0066] Random feature sampling: randomly selected at each split n features (where n is the total number of features);

[0067] Bootstrap sampling: Each tree uses sampling with replacement at a sampling ratio of 0.8.

[0068] Feature importance scores are obtained using a random forest model. , in These are feature importance vectors arranged in descending order of importance. This refers to the total number of features. Random forests calculate the importance score for each feature by averaging the reduction in impurity (Gini coefficient or variance reduction) across multiple decision trees, effectively capturing non-linear feature relationships.

[0069] Specifically, based on the feature importance score, the first derivative of the feature importance is calculated, a gradient spectrum is constructed, and the threshold candidate points of the gradient spectrum are obtained, including:

[0070] Calculate the first derivative of feature importance, construct the gradient spectrum, and use formula (1), where This represents the difference in importance between adjacent features. To ensure comparability at different scales, the gradient is normalized, as shown in Equation (2). Then, a local extremum detection algorithm is used to identify significant change points in the gradient spectrum, as shown in Equation (3). These change points represent significant breakpoints in feature importance, indicating potential feature selection threshold locations.

[0071] (1)

[0072] (2)

[0073] (3)

[0074] Specifically, the cumulative importance contribution spectrum of the features is constructed, and the candidate threshold points for the cumulative importance contribution spectrum are obtained as follows:

[0075] The cumulative importance contribution of computational features ,in Indicates the preceding The cumulative contribution ratio of each feature. The elbow rule is applied to detect the inflection point of the cumulative contribution curve. The inflection point indicates a point where the cumulative contribution growth rate declines significantly, representing the equilibrium point of diminishing returns.

[0076] Specifically, the threshold candidate points, which combine the gradient spectrum and the cumulative importance contribution spectrum, are evaluated by optimizing the score function, including:

[0077] Threshold candidate points combining gradient spectrum and cumulative contribution spectrum Each candidate threshold point is evaluated by optimizing the score function, Equation (4);

[0078] (4)

[0079] The proportion of variance explained (model performance) is represented. The feature ratio penalty term represents the model complexity. The squared penalty term strengthens the control over redundant features and tends to select a more compact subset of features.

[0080] Specifically, the optimal threshold position is determined by formula (5).

[0081] (5)

[0082] The results of the bispectral adaptive threshold feature selection method based on random forest in this embodiment are as follows: Figures 2-6 The first derivative rate of change method for feature importance identified breakpoints at the 4th, 6th, 8th, 11th, and 14th ranked features, while the elbow rule method determined the inflection point at the 9th ranked feature. The optimal number of features was systematically determined to be 11. Figure 6 As shown, this comprehensive screening strategy achieves a good balance between the feature dimensionality of the model and the explanatory power of the predictor variables. The selected 11 features can explain 78.45% of the total variance of the target variable, effectively capturing most of the information structure in the original dataset and significantly reducing the data dimensionality. Figure 3 The 11 characteristics selected are: canopy top evaporation (efttoc), net surface solar radiation (snsr), net surface thermal radiation (sntr), downward surface solar radiation (ssrd), surface temperature (st), U-component wind speed at 10 meters (u10), bare soil evaporation (efbs), downward surface longwave radiation (strd), V-component wind speed at 10 meters (v10), first soil volumetric water content, 0-10 cm (vswl1), and 2-meter dew point temperature (t2d).

[0083] This embodiment first calculates the importance scores of all features using a random forest and sorts them in descending order. Here, the "threshold position k" actually represents "selecting the k most important features." Next, bispectral analysis (gradient spectrum identifies breakpoints where importance decreases significantly, and cumulative contribution spectrum finds inflection points where diminishing returns occur) is used to find candidate threshold positions. Then, an optimized score function is used to find the optimal balance between model performance (the explanatory power Ck of the top k features) and model complexity (feature quantity penalty term (k / n)²), yielding the optimal k value. This k value directly determines the optimal feature subset—the top k features in the importance ranking. The example in the document shows that when k=11, the optimal balance is achieved, and the corresponding top 11 most important features constitute the optimal feature subset, achieving 78.45% variance explanation with a relatively small number of features.

[0084] To compare the effectiveness of the RF-DSAT method, this embodiment introduces four classic feature selection techniques: Pearson correlation coefficient, mutual information, Lasso regression, and SHAP value. Figure 7 The results of five methods for ranking features are shown, with the top 11 features selected by each method indicated by the black boxes. Figure 4 The characteristic variables are: wvni: near-infrared water vapor, cwp: cloud water path, cot: cloud optical thickness, cer: effective cloud radius, cpi: cloud phase infrared, ctt: cloud top temperature, ctp: cloud top pressure, cth: cloud top height, st: surface temperature, eva: total evaporation, efbs: bare soil evaporation, efttoc: canopy top evaporation, efvt: vegetation transpiration, peva: potential evaporation, ssrd: downward surface solar radiation, strd: downward surface thermal radiation Radiance, snsr: net solar radiation at the Earth's surface, sntr: net thermal radiation at the Earth's surface, slhf: latent heat flux at the Earth's surface, sshf: sensible heat flux at the Earth's surface, fa: predicted albedo, t2: air temperature at 2 meters, t2d: dew point temperature at 2 meters, ts: surface skin temperature, u10: U-component wind speed at 10 meters altitude, v10: V-component wind speed at 10 meters altitude, laihv: leaf area index of high vegetation, lailv: leaf area index of low vegetation, vswl1: volumetric water content (first soil layer, 0-10 cm).

[0085] These feature combinations were evaluated using the MSPD model, and ablation experiments were conducted with varying numbers of features. The results are as follows: Figure 8 As shown, the random forest selection of 11 features yielded the best performance, achieving an NSE of 0.9324, a MAE of 0.4752, and an RMSE of 1.2052, significantly outperforming other methods. Ablation experiments further demonstrated that model performance decreased when the number of features decreased from 11 to 10 or increased to 12, with the best performance achieved when the number of features was 11, indicating that this number strikes a good balance between information sufficiency and redundancy control.

[0086] In the comparison among different methods, Random Forest (NSE=0.9324) performed best, followed by SHAP (NSE=0.9218), Pearson (NSE=0.9166), and Lasso (NSE=0.9153), while Mutual Information performed the worst (NSE=0.9112). Furthermore, the model using all features (NSE=0.9239) was inferior to the model using RF to select features, highlighting the necessity and effectiveness of feature selection. It is worth noting that although the RF method has a slightly higher bias, its overall performance is the best, further validating the advantages of RF-DSAT.

[0087] Advantages of adaptability and generalization ability:

[0088] Data-driven adaptability: No need to manually set fixed thresholds; the optimal feature subset is automatically determined entirely based on data characteristics; the dual verification mechanism of gradient spectrum and cumulative contribution spectrum improves the reliability of threshold determination; it has good adaptability to different datasets and application scenarios.

[0089] Cross-regional generalization ability: In tests in multiple different climate regions, the feature combinations selected by RF-DSAT all showed stable prediction performance, with an average accuracy improvement of 8-15% compared to fixed feature selection strategies.

[0090] Feature interpretability and interpretability: Feature importance visualization: Provides clear ranking and distribution visualization of feature importance; gradient spectrum and cumulative contribution spectrum provide intuitive explanation of the feature selection process; Supports visual analysis of feature interaction effects;

[0091] Domain knowledge consistency: The selected features are highly consistent with meteorological knowledge. For example, key meteorological elements such as cloud water path (CWP), net shortwave radiation (SNTR), and cloud optical thickness (COT) are all correctly identified as important features.

[0092] Transparency of the decision-making process: Fully record every step of feature selection and the basis for the decision; provide detailed statistical analysis reports; support traceability analysis of feature selection results;

[0093] The practical application value of this embodiment:

[0094] Potential for operational applications: It can be directly integrated into existing weather forecasting systems; it supports both real-time and batch processing application modes; and it provides standardized API interfaces for easy system integration.

[0095] Economic benefits: Improved forecast accuracy and reduced disaster losses through precise feature selection; reduced computational resource consumption and reduced operating costs; improved efficiency of model deployment and maintenance;

[0096] The scientific research value of this embodiment is that it provides a new methodology for meteorological feature analysis, promotes the development of automated feature engineering technology, and provides a reference for feature selection problems in other fields.

[0097] The innovative contribution of this embodiment:

[0098] Theoretical innovations: This paper proposes a theoretical framework for feature selection through dual-spectrum fusion for the first time; it innovatively combines gradient analysis and elbow detection; and it establishes a mathematical balance model to explain variance and model complexity.

[0099] The innovative methods of this embodiment include: designing a fully data-driven adaptive threshold determination mechanism; developing an efficient bispectral analysis algorithm; and proposing a feature interaction effect identification method based on random forest.

[0100] The innovative applications of this embodiment are: the first systematic application of bispectral feature selection in the field of rainfall simulation; the realization of intelligent feature screening of high-dimensional meteorological data; and the provision of technical support for refined weather forecasting.

[0101] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A bispectral adaptive threshold feature selection method based on random forest, characterized in that, include: Feature importance scores in rainfall simulation and weather forecasting were obtained using a random forest model. Based on the feature importance score, the first derivative of the feature importance is calculated, a gradient spectrum is constructed, and the threshold candidate points of the gradient spectrum are obtained. Obtaining the threshold candidate points of the gradient spectrum includes: The gradient spectrum is normalized. The local extremum detection algorithm is used to identify preset change points in the normalized gradient spectrum; where the preset change points represent significant breakpoints in feature importance and indicate potential feature selection threshold positions. Construct the cumulative importance contribution spectrum of the features and obtain the threshold candidate points of the cumulative importance contribution spectrum; The threshold candidate points for obtaining the cumulative importance contribution spectrum include: The elbow rule is applied to detect the inflection point of the cumulative contribution curve in the cumulative importance contribution spectrum, and the inflection point is regarded as a threshold candidate point. By combining the threshold candidate points of the gradient spectrum and the cumulative importance contribution spectrum, the optimal feature is determined by evaluating each candidate threshold point through an optimized fractional function.

2. The bispectral adaptive threshold feature selection method based on random forest according to claim 1, characterized in that, The feature importance score is: in, This represents the feature importance vectors arranged in descending order of importance. This represents the total number of features.

3. The bispectral adaptive threshold feature selection method based on random forest according to claim 1, characterized in that, The gradient spectrum is: in, Represents the gradient spectrum. This represents the importance difference between adjacent features, where i represents any feature from the 1st to the (n-1st)th feature.

4. The bispectral adaptive threshold feature selection method based on random forest according to claim 1, characterized in that, The cumulative importance contribution spectrum is as follows: in, Indicates the preceding The cumulative contribution ratio of each feature.

5. The bispectral adaptive threshold feature selection method based on random forest according to claim 1, characterized in that, The optimized score function is: in, Indicates the preceding The cumulative contribution ratio of each feature The score represents the feature proportion penalty term. k ) represents the score of the candidate threshold point.

6. The bispectral adaptive threshold feature selection method based on random forest according to claim 1, characterized in that, Features in rainfall simulation and weather forecasting include: Near-infrared water vapor, cloud water path, cloud optical thickness, cloud effective radius, cloud phase infrared, cloud top temperature, cloud top pressure, cloud top height, surface temperature, total evaporation, bare soil evaporation, forest canopy top evaporation, vegetation transpiration, potential evaporation, downward surface solar radiation, downward surface thermal radiation, net surface solar radiation, net surface thermal radiation, surface latent heat flux, surface sensible heat flux, predicted albedo, 2-meter air temperature, 2-meter dew point temperature, surface skin temperature, 10-meter height U-component wind speed, 10-meter height V-component wind speed, high vegetation leaf area index, low vegetation leaf area index, volumetric water content.

7. The bispectral adaptive threshold feature selection method based on random forest according to claim 1, characterized in that, The optimal features include: Evaporation at the top of the canopy, net solar radiation at the ground surface, net thermal radiation at the ground surface, downward solar radiation at the ground surface, surface temperature, wind speed of the U component at a height of 10 meters, evaporation of bare soil, downward thermal radiation at the ground surface, wind speed of the V component at a height of 10 meters, volumetric water content of the first soil layer, and dew point temperature at 2 meters.

Citation Information

Patent Citations

  • Random forest improvement method based on feature importance

    CN120106251A

  • Feature selection for reinforcement learning models

    US20250077960A1