Pollen concentration prediction method and system based on machine learning
Through machine learning-based methods, random forest and gradient enhancement models are constructed and weighted averaged using voting regressors, the problems of insufficient accuracy and high computational complexity in the prediction of pollen concentration are solved, and the prediction effects of high accuracy and low complexity are achieved.
Patent Information
- Application Number
- CN202510023841.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional statistical methods and physical models have problems of insufficient accuracy and high computational complexity in pollen concentration prediction, which is difficult to meet the needs of real-time dynamic monitoring.
Using a machine learning-based method, a random forest and gradient enhancement model is constructed by collecting multi-dimensional features, recursive feature elimination and random forest feature importance analysis combination optimization, and a weighted average is performed with a voting regressor to generate the final pollen concentration prediction result.
It significantly improves the accuracy and stability of pollen concentration prediction, reduces the computational complexity, and meets the needs of real-time dynamic monitoring.
Smart Images

Figure CN120030490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a pollen concentration prediction method and system based on machine learning. Background Art
[0002] Pollen concentration prediction plays an extremely important role in the fields of environmental science and public health, especially in the season of high incidence of allergies. Its prediction results are crucial to personal health management and public policy making. Accurate pollen concentration prediction can not only help people with allergies to make protective preparations in advance and reduce the incidence of allergic symptoms, but also provide a scientific basis for the resource allocation of medical institutions and reduce the pressure on the medical system.
[0003] At present, pollen concentration prediction mainly relies on two types of methods: one is the traditional statistical method, which uses historical data for regression analysis to predict pollen concentration. These methods are usually based on simple linear regression and multiple linear regression, combined with environmental factors such as temperature, humidity and wind speed for analysis; the other is based on physical model simulation, using meteorological models (such as WRF) to simulate aerodynamic processes, combined with vegetation distribution and pollen release patterns, to predict pollen diffusion in specific areas.
[0004] However, when dealing with pollen concentration prediction, traditional statistical methods have difficulty capturing the complex nonlinear relationship between environmental factors and pollen concentration due to their simple model structure, resulting in a lack of accuracy in the prediction results; and physical models have extremely high computational complexity, usually requiring powerful computing resources and long-term operation, making it difficult to meet the needs of real-time dynamic monitoring. Summary of the invention
[0005] In order to solve the technical problems that traditional statistical methods are difficult to capture the complex nonlinear relationship between environmental factors and pollen concentration when processing pollen concentration prediction due to the simple model structure, resulting in a lack of accuracy in the prediction results; and the physical model has extremely high computational complexity, usually requires powerful computing resources and long-term operation, resulting in difficulty in meeting the needs of real-time dynamic monitoring, the present invention provides a pollen concentration prediction method and system based on machine learning.
[0006] The technical solution provided by the embodiment of the present invention is as follows:
[0007] First aspect:
[0008] An embodiment of the present invention provides a pollen concentration prediction method based on machine learning, comprising:
[0009] S1: Collect multi-dimensional features that affect pollen concentration;
[0010] S2: Optimizing the multi-dimensional features by combining recursive feature elimination and random forest feature importance analysis to screen out excellent features whose importance scores are higher than the preset importance scores;
[0011] S3: Construct a pollen concentration prediction model based on random forest;
[0012] S4: using the excellent features as input of a pollen concentration prediction model based on random forest, and outputting a first pollen concentration prediction result;
[0013] S5: Construct a pollen concentration prediction model based on gradient boosting;
[0014] S6: using the excellent features as input of a pollen concentration prediction model based on gradient boosting, and outputting a second pollen concentration prediction result;
[0015] S7: Using a voting regressor, perform weighted average calculation on the first pollen concentration prediction result and the second pollen concentration prediction result to determine a final pollen concentration prediction result.
[0016] Second aspect:
[0017] An embodiment of the present invention provides a pollen concentration prediction system based on machine learning, comprising:
[0018] processor;
[0019] A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the pollen concentration prediction method based on machine learning as described in the first aspect is implemented.
[0020] The third aspect:
[0021] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method for predicting pollen concentration based on machine learning as described in the first aspect is implemented.
[0022] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0023] (1) In the present invention, the nonlinear modeling ability and feature interaction processing advantages of the random forest model are used to effectively improve the ability to capture complex relationships; the prediction error is further optimized by the gradient boosting model, which significantly improves the model accuracy; the prediction results of the two models are weighted averaged by the voting regressor, which fully integrates the complementary advantages of the models, thereby significantly improving the accuracy of the prediction results.
[0024] (2) In the present invention, the prediction results are quickly generated through the efficient parallel computing capability of the random forest model; the time cost of training and prediction is significantly reduced through the optimized design of the gradient boosting algorithm; the result integration process is further simplified through the lightweight weighted average of the voting regressor, thereby greatly reducing the computational complexity and effectively meeting the needs of real-time dynamic monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1 A schematic diagram of a process flow of a pollen concentration prediction method based on machine learning provided in an embodiment of the present invention;
[0027] Figure 2 A schematic diagram of the structure of a pollen concentration prediction system based on machine learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0029] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0030] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0031] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.
[0032] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0033] Reference Manual Attached Figure 1 , showing a flow chart of a pollen concentration prediction method based on machine learning provided in an embodiment of the present invention.
[0034] The embodiment of the present invention provides a pollen concentration prediction method based on machine learning, which can be implemented by a pollen concentration prediction device based on machine learning, and the pollen concentration prediction device based on machine learning can be a terminal or a server. The processing flow of the pollen concentration prediction method based on machine learning can include the following steps:
[0035] S1: Collect multi-dimensional features that affect pollen concentration.
[0036] Among them, multidimensional features refer to the key factors affecting the target variables (such as pollen concentration) extracted from multiple aspects, aiming to comprehensively describe the complex conditions such as environment, time, space, etc. that affect the target.
[0037] In a possible implementation, the multi-dimensional feature data specifically includes:
[0038] Temperature, precipitation, relative humidity, air pressure, wind speed, wind direction, normalized difference vegetation index, enhanced vegetation index, leaf area index of vegetation, altitude information of pollen monitoring points and time characteristic variables.
[0039] Specifically, multi-dimensional characteristic data, including meteorological, vegetation, geographical and temporal characteristics, are used to comprehensively describe the main environmental factors affecting pollen concentration. Specific characteristic variables include:
[0040] (1) Meteorological variables: Meteorological data including temperature, precipitation, relative humidity, air pressure, wind speed and direction were extracted within the 1km, 5km and 10km buffer zones around the pollen monitoring sites to capture the impact of local climate conditions on pollen concentration. The extracted meteorological variables included the daily meteorological averages on the day of pollen monitoring and 1-7 days before monitoring.
[0041] (2) Vegetation variables: The normalized difference vegetation index (NDVI), enhanced vegetation index (EVI), leaf area index (LAI) of high vegetation and low vegetation, and area ratio of different land use types (arable land, forest land, shrub, grassland, water body) were extracted within the 1 km, 5 km, and 10 km buffer zone around the pollen monitoring point to reflect the contribution of plant growth status and surrounding land cover to pollen concentration.
[0042] (3) Geographic variables: including the altitude information of pollen monitoring points, which is used to consider the potential impact of geographical altitude differences on pollen dispersion and deposition.
[0043] (4) Temporal characteristics: By extracting temporal characteristic variables, including the day of the year (doy), week of the year (woy), month of the year (moy), and season, the temporal variation characteristics of pollen concentration are revealed.
[0044] In this invention, by covering four dimensions, namely, meteorology, vegetation, geography and time, the main environmental factors affecting pollen concentration are fully described, ensuring the richness and coverage of data and reducing the risk of missing key influencing factors. At the same time, by collecting and integrating multi-dimensional feature data, the key factors affecting pollen concentration can be captured more comprehensively and accurately, providing sufficient input information for the model, thereby significantly improving the accuracy and stability of pollen concentration prediction.
[0045] S2: The multi-dimensional features are optimized by combining recursive feature elimination and random forest feature importance analysis to screen out excellent features whose importance scores are higher than the preset importance scores.
[0046] Among them, recursive feature elimination is a feature selection method that is used to gradually eliminate features with low contribution to model prediction from the initial feature set, and finally screen out the most important feature subset.
[0047] The core idea of random forest feature importance analysis is to quantify the importance of the feature by observing the frequency of the feature used to split nodes in the decision tree and its impact on the prediction results.
[0048] In a possible implementation manner, the S2 specifically includes:
[0049] Collect multi-dimensional features that affect pollen concentration;
[0050] Initialize the random forest model regression model and set random_state=42;
[0051] Recursively screening the multi-dimensional features by using the initialized random forest model regression model base learner through recursive feature elimination and cross-validation;
[0052] In each recursion, the first importance score of each feature is calculated, and the feature with the smallest first importance score is eliminated;
[0053] Repeat the preset number of iterations and output the first feature subset after recursive feature elimination;
[0054] In the present invention, features that contribute the least to model prediction are gradually removed by recursive screening to ensure that the preliminary feature subset (first feature subset) is more compact and important.
[0055] Initialize the random forest model;
[0056] Optionally, set the number of trees to 500, the maximum depth to 15, the minimum number of samples for node splitting to 5, and the minimum number of samples for leaf nodes to 2;
[0057] Inputting the first feature subset into the initialized random forest model, calculating the second importance score of each first feature subset; screening out the features represented by the second importance score value of each first feature subset being greater than the second preset importance score value, to form a second feature subset;
[0058] In the present invention, random forest uses the frequency of feature splitting in the decision tree and the impact on the prediction error to quantify the importance of each feature, and can quickly evaluate the contribution of the feature.
[0059] The intersection of the first feature subset and the second feature subset is taken to form a third feature set, and the features in the third feature set are sorted from high to low according to the importance score value, and a preset proportion of features are selected as the excellent features.
[0060] In the present invention, the final third feature set is ranked by importance, and features with the optimal ratio are selected to further optimize the input feature dimension of the model and ensure the high quality and high contribution of the feature subset.
[0061] Specifically, recursive feature elimination:
[0062] First, the recursive feature elimination method is used to perform a preliminary screening of features to eliminate redundant and inefficient features. This method recursively trains the model and gradually deletes features in each iteration to identify the most important features that contribute to model prediction. In the present invention, a random forest regressor is used as a base learner. The process uses a cross-validation (5-fold cross-validation) strategy to calculate the negative mean squared error (neg_mean_squared_error) after each iteration as an evaluation criterion to measure the impact of different feature sets on model performance. The process of recursive feature elimination includes the following steps:
[0063] 1: Initialize the random forest regression model (RandomForestRegressor) and set the parameter: random_state=42 to ensure the repeatability of the experiment. 2: Use RFECV to select features and specify min_features_to_select=15 to ensure that at least 15 features are selected for subsequent model training. This threshold can be adjusted according to the specific problem. 3 In each recursive process, the negative mean square error (neg_mean_squared_error) is used as the evaluation metric to evaluate the impact of the current feature set on the model performance. The goal of minimizing the negative mean square error is to select the feature set that can best improve the generalization ability of the model. 4: In the final feature selection, retain the features that have the greatest impact on pollen concentration prediction.
[0064] (1) Random Forest Feature Importance Analysis:
[0065] Next, based on the feature importance evaluation results of the random forest regression model, the remaining features are further screened. The random forest model itself can measure the contribution of each feature to the final prediction result through the built-in feature importance score. The importance score of each feature is calculated based on the information gain of the feature in all decision trees. Specifically, the model calculates the reduction of the error in the decision tree training by each feature and generates a feature importance score based on this. 1: Train a standard random forest regression model, train with the selected features, and calculate the importance score of each feature. In the present invention, the parameters used include: n_estimators = 500 (i.e. 500 trees), max_depth = 15 (maximum depth is 15), min_samples_split = 5 (the minimum number of split samples for each node is 5), min_samples_leaf = 2 (the minimum number of samples for each leaf node is 2). 2: Obtain the importance score of each feature through the feature_importances_ attribute of the model. The score indicates the degree of contribution of the feature to reducing the model error. 3: Based on the feature importance score, select features with a feature importance score greater than 0.01 for subsequent model training. This threshold is determined based on model test results and experience, and the purpose is to eliminate features that contribute less to model prediction.
[0066] Combination method optimization:
[0067] Finally, the results of recursive feature elimination and random forest feature importance analysis are combined to retain the features selected by both methods, and further select the features with the top 30% contribution. Redundant and low-contribution features are eliminated to generate the final optimized feature set X_optimized.
[0068] In summary, recursive feature elimination can gradually filter out features with low contribution to prediction, reduce feature dimensions, reduce model complexity, and reduce computing time and resource consumption. At the same time, through the feature importance analysis of random forests, it is clear which features have the greatest impact on pollen concentration prediction, enhance the interpretability of the model, and help understand the role of different environmental factors.
[0069] S3: Construct a pollen concentration prediction model based on random forest.
[0070] Among them, Random Forest is an ensemble learning algorithm consisting of multiple decision trees, mainly used for classification and regression tasks. It generates a strong learner by voting or averaging the prediction results of multiple weak learners (decision trees), thereby improving the accuracy and robustness of the model.
[0071] In a possible implementation, the pollen concentration prediction model based on random forest is specifically:
[0072]
[0073] in, represents the first pollen concentration prediction result, T represents the total number of trees, and f t () represents the prediction result of the tth decision tree, and x represents multi-dimensional feature data.
[0074] In the present invention, decision trees can handle complex nonlinear relationships, and random forests further enhance the modeling capabilities of nonlinearity and feature interactions by constructing multiple decision trees, thereby improving the accuracy of the model. At the same time, a single decision tree is prone to being overly sensitive to training data, while random forests can improve the stability of prediction results by constructing multiple trees and taking the average value, so that even if the input data fluctuates to a certain extent, the prediction results can remain reliable.
[0075] S4: Using the multi-dimensional feature data as input to the pollen concentration prediction model based on random forest, and outputting a first pollen concentration prediction result.
[0076] In the present invention, the model can more accurately describe the complex nonlinear relationship between pollen concentration and characteristic variables, thereby improving the accuracy of prediction results.
[0077] S5: Construct a pollen concentration prediction model based on gradient boosting.
[0078] Among them, gradient boosting is an ensemble learning method mainly used for regression and classification tasks. It builds a powerful prediction model by gradually training a series of weak learners (usually decision trees) so that each learner corrects the errors of its predecessor learner as much as possible.
[0079] In a possible implementation, the pollen concentration prediction model based on gradient boosting is specifically:
[0080]
[0081] in, represents the second pollen concentration prediction result, M represents the number of iterations, η represents the learning rate, h m () represents the output of the weak learner in the mth iteration.
[0082] In the present invention, gradient boosting allows the new model to fit the residual of the previous model through each iteration. The model can gradually optimize the pollen concentration prediction, capture more subtle patterns and relationships in the data, and significantly improve the accuracy of the prediction.
[0083] S6: Using the multi-dimensional feature data as input to the pollen concentration prediction model based on gradient boosting, and outputting a second pollen concentration prediction result.
[0084] In a possible implementation manner, after S6, the method further includes:
[0085] The pollen concentration prediction model based on random forest and the pollen concentration prediction model based on gradient boosting were trained through grid search cross-validation technology.
[0086] Among them, grid search cross-validation is a technique for model hyperparameter optimization, which mainly selects the optimal hyperparameter combination by traversing the specified hyperparameter combination and combining cross-validation to evaluate the model performance of each set of parameters. Its goal is to find the hyperparameter configuration that makes the model perform best on the validation data set.
[0087] In a possible implementation, the grid search cross-validation technique is used to train the pollen concentration prediction model based on random forest and the pollen concentration prediction model based on gradient boosting, specifically including:
[0088] Initializing various parameters of the random forest pollen concentration prediction model and the gradient boosting-based pollen concentration prediction model;
[0089] Build a grid searcher;
[0090] Among them, the grid searcher is a tool for hyperparameter optimization in machine learning. It can automatically search for the optimal hyperparameter combination in a specified parameter grid to achieve the best performance of the model on the validation set.
[0091] By constructing the grid searcher, all parameter combinations are traversed, the mean square error values of each pollen concentration prediction model under each set of parameters are calculated, the mean square error values are sorted from low to high, and the parameter combinations corresponding to the mean square error values with the highest sorting are selected;
[0092] Build a Bayesian optimizer;
[0093] Among them, Bayesian optimizer is a technology for optimizing black box functions, which is often used for hyperparameter tuning of machine learning models. Unlike traditional grid and random search, Bayesian optimization predicts the performance of unobserved points by constructing a probabilistic model of the objective function (such as a Gaussian process), and selects the next set of hyperparameters for evaluation based on this prediction, thereby finding the optimal parameters with fewer trials.
[0094] Based on the parameter combination corresponding to the top-ranked mean square error value, the Bayesian optimizer is used for optimization with the minimum mean square error as the goal. When the termination condition is met, the iteration is stopped and the optimal hyperparameter combination is output.
[0095] The calculation formula of the mean square error is:
[0096]
[0097] Among them, MSE represents mean square error, n represents the total number of samples, represents the predicted value of the i-th sample, y i Represents the actual value of the i-th sample.
[0098] In a possible implementation manner, the termination condition is specifically:
[0099] When the mean square error does not exceed the preset error value during 10 consecutive iterations, the iteration is stopped.
[0100] Preliminary determination of hyperparameters and rough grid search. By roughly defining the parameter range, the computational complexity of subsequent fine search can be reduced. 1: Set the parameter range of the random forest model: number of trees (n_estimators): [100, 200, 300, 500]; maximum depth (max_depth): [10, 20, None]; minimum number of node split samples (min_samples_split): [2, 5, 10]; minimum number of leaf node samples (min_samples_leaf): [1, 2, 4]. 2: Set the parameter range of the gradient boosting model: learning rate (learning_rate): [0.1, 0.05, 0.01]; maximum depth (max_depth): [3, 5, 7]; number of trees (n_estimators): [100, 200, 300, 500]; minimum number of samples for node splitting (min_samples_split): [2, 5, 10]; minimum number of samples for leaf nodes (min_samples_leaf): [1, 2, 4]. 3: Construct a grid searcher (GridSearchCV), use a 3-fold cross-validation strategy to evaluate the performance of each parameter combination, and use the mean square error (MSE) of the validation set as the evaluation indicator.
[0101] Fine-tuning: Bayesian optimization. Based on the rough grid search results, Bayesian optimization is used to further fine-tune the hyperparameters to improve model performance. 1: Optimize the random forest model range settings: n_estimators: [250,350]; min_samples_split: [3,7]; min_samples_leaf: [1,3]. 2: Optimize the random forest model range settings: learning_rate: [0.08,0.12]; n_estimators: [250,350; max_depth: [4,6]. 3: Use Gaussian process to construct the posterior distribution of the objective function and select the hyperparameter combination by the expected improvement criterion. Repeat the search 30 times and finally select the hyperparameter combination with the smallest mean square error.
[0102] Combination of multi-stage grid search and random search: By narrowing the search range in stages, combining random search and Bayesian optimization strategies, the computational complexity is further reduced and the optimization efficiency is improved. Among them, the random forest is set to 100 random samplings; the gradient boosting is set to 100 random samplings; the mean square error is minimized as the goal, and the optimal parameters are selected according to the mean square error (MSE) for each sampling. Step 1: Coarse grid search, with a large range (the parameter values are more dispersed). Step 2: Random search, randomly explore some of the parameter combinations screened in the first stage. Step 3: Bayesian optimization, further accurately search for the optimal hyperparameters based on the random search results.
[0103] Model training and validation: 1. Cross-validation strategy, using 5-fold cross-validation. The data set is divided into a training set (80%) and a validation set (20%). In each fold, the training set is used to train the model, and the validation set is used to evaluate the model performance. 2. Early stopping mechanism, set the number of early stopping tolerance iterations to 10. If the mean square error does not decrease significantly after 10 consecutive iterations, stop training early.
[0104] In the present invention, the grid search exhaustively enumerates all possible parameter combinations and comprehensively evaluates each set of parameters in combination with cross-validation to ensure that an initial parameter combination with good performance is found, thereby providing a reliable candidate range for subsequent optimization. At the same time, based on the better parameter combinations screened out by the grid search, the probability model of the objective function (such as the Gaussian process) is constructed to efficiently predict and explore the unobserved area, and the search for the optimal hyperparameter combination is gradually refined.
[0105] Furthermore, the optimization process is controlled by the iterative convergence condition of the mean square error (such as not exceeding the preset error value for 10 consecutive times), which not only avoids the waste of computing resources caused by excessive searching, but also effectively improves the optimization efficiency.
[0106] S7: Using a voting regressor, a weighted average calculation is performed on the first pollen concentration prediction result and the second pollen concentration prediction result to determine a final pollen concentration prediction result.
[0107] Among them, the Voting Regressor is an integrated learning method that combines the prediction results of multiple regression models and outputs a more stable and accurate prediction result through weighted average or simple average.
[0108] In a possible implementation manner, S7 specifically includes:
[0109] Through the voting regressor, the first pollen concentration prediction result and the second pollen concentration prediction result are weighted averaged and calculated using the following formula to determine the final pollen concentration prediction result:
[0110]
[0111] in, represents the final pollen concentration prediction result, α represents the weight parameter, represents the pollen concentration prediction result of the random forest regression model for the i-th sample, Represents the pollen concentration prediction result of the gradient boosting regression model for the i-th sample.
[0112] In the present invention, the prediction results of the random forest and gradient boosting models are weighted averaged through a voting regressor, which can fully utilize the advantages of the two models, reduce the deviation of a single model, and improve the accuracy and stability of the prediction.
[0113] Among them, Grid Search with Cross Validation is a technique for optimizing model hyperparameters. By systematically traversing predefined parameter combinations and combining cross-validation to evaluate the performance of each set of parameters, the optimal hyperparameter combination is selected to improve model performance.
[0114] It should be noted that through recursive feature elimination (RFECV) and random forest feature importance analysis, redundant features with little impact on pollen concentration were eliminated, and finally the meteorological, vegetation, geographical and temporal features with high importance were retained.
[0115] In one possible implementation, the number of decision trees of the random forest-based pollen concentration prediction model is set to 200, 500 and 100, the maximum depth of the decision tree is set to 10, 15 and 20, the minimum number of split samples of the decision tree is set to 5, and the minimum number of leaf node samples of the decision tree is set to 10, 5 and 2.
[0116] Specifically, for the random forest model: the main hyperparameters include the number of trees, the maximum depth of the tree, the minimum number of sample splits, and the minimum number of leaf node samples. The random forest model in the present invention is trained under the configuration of 200, 500 and 1000 trees respectively to enhance the model stability and accuracy. The maximum depth of the decision tree is set to 10, 15 and 20 to balance the complexity of the model and the risk of overfitting. The minimum number of split samples and the minimum number of leaf node samples are set to 5 and 10, 2 and 5 to avoid overfitting.
[0117] In one possible implementation, the number of weak learners of the pollen concentration prediction model based on gradient boosting is set to 100, 200 and 300, the learning rate is set to 0.01, 0.1 and 0.2, and the maximum depth is set to 3, 5 and 7.
[0118] Specifically, for the gradient boosting model: the optimization parameters include the number of weak learners, the learning rate and the maximum depth. The number of weak learners of the gradient boosting model in the present invention is set to 100, 200 and 300 to ensure sufficient fitting of the residual. The learning rate is set to 0.01, 0.1 and 0.2 to regulate the model update step size and convergence speed. The maximum depth is set to 3, 5 and 7 to control the complexity of the tree.
[0119] In the present invention, the hyperparameters of the random forest and gradient boosting models are optimized by using grid search cross-validation technology, and the best balance between model performance and complexity can be found, thereby significantly improving the accuracy and stability of prediction. At the same time, for the random forest model, the number and depth of decision trees can be optimized to enhance the ability to capture complex features; for the gradient boosting model, the learning rate and the number of weak learners can be adjusted to balance the fitting speed and complexity of the model, thereby significantly improving the accuracy and robustness of pollen concentration prediction.
[0120] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0121] (1) In the present invention, the nonlinear modeling ability and feature interaction processing advantages of the random forest model are used to effectively improve the ability to capture complex relationships; the prediction error is further optimized by the gradient boosting model, which significantly improves the model accuracy; the prediction results of the two models are weighted averaged by the voting regressor, which fully integrates the complementary advantages of the models, thereby significantly improving the accuracy of the prediction results.
[0122] (2) In the present invention, the prediction results are quickly generated through the efficient parallel computing capability of the random forest model; the time cost of training and prediction is significantly reduced through the optimized design of the gradient boosting algorithm; the result integration process is further simplified through the lightweight weighted average of the voting regressor, thereby greatly reducing the computational complexity and effectively meeting the needs of real-time dynamic monitoring.
[0123] Reference Manual Attached Figure 2 , showing a structural schematic diagram of a pollen concentration prediction system based on machine learning provided by the present invention.
[0124] The present invention further provides a pollen concentration prediction system 20 based on machine learning, which is applied to the above-mentioned pollen concentration prediction method based on machine learning, comprising:
[0125] Processor 201.
[0126] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the pollen concentration prediction method based on machine learning as in the method embodiment is implemented.
[0127] The machine learning-based pollen concentration prediction system 20 provided in the present invention can execute the above-mentioned machine learning-based pollen concentration prediction method and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.
[0128] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0129] (1) In the present invention, the nonlinear modeling ability and feature interaction processing advantages of the random forest model are used to effectively improve the ability to capture complex relationships; the prediction error is further optimized by the gradient boosting model, which significantly improves the model accuracy; the prediction results of the two models are weighted averaged by the voting regressor, which fully integrates the complementary advantages of the models, thereby significantly improving the accuracy of the prediction results.
[0130] (2) In the present invention, the prediction results are quickly generated through the efficient parallel computing capability of the random forest model; the time cost of training and prediction is significantly reduced through the optimized design of the gradient boosting algorithm; the result integration process is further simplified through the lightweight weighted average of the voting regressor, thereby greatly reducing the computational complexity and effectively meeting the needs of real-time dynamic monitoring.
[0131] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0132] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0133] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0134] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0135] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0136] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0137] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0138] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0139] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0140] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0142] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0143] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for predicting pollen concentration based on machine learning as described in the method embodiment is implemented.
[0144] A computer-readable storage medium provided by the present invention can implement the steps and effects of the pollen concentration prediction method based on machine learning in the above method embodiment. To avoid repetition, the present invention will not go into details.
[0145] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0146] (1) In the present invention, the nonlinear modeling ability and feature interaction processing advantages of the random forest model are used to effectively improve the ability to capture complex relationships; the prediction error is further optimized by the gradient boosting model, which significantly improves the model accuracy; the prediction results of the two models are weighted averaged by the voting regressor, which fully integrates the complementary advantages of the models, thereby significantly improving the accuracy of the prediction results.
[0147] (2) In the present invention, the prediction results are quickly generated through the efficient parallel computing capability of the random forest model; the time cost of training and prediction is significantly reduced through the optimized design of the gradient boosting algorithm; the result integration process is further simplified through the lightweight weighted average of the voting regressor, thereby greatly reducing the computational complexity and effectively meeting the needs of real-time dynamic monitoring.
[0148] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0149] There are a few points to note:
[0150] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.
[0151] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.
[0152] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.
[0153] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A pollen concentration prediction method based on machine learning, characterized in that: include: S1: Collect multi-dimensional features that affect pollen concentration; S2: Optimizing the multi-dimensional features by combining recursive feature elimination and random forest feature importance analysis to screen out excellent features whose importance scores are higher than the preset importance scores; S3: Construct a pollen concentration prediction model based on random forest; S4: using the excellent features as input of a pollen concentration prediction model based on random forest, and outputting a first pollen concentration prediction result; S5: Construct a pollen concentration prediction model based on gradient boosting; S6: using the excellent features as input of a pollen concentration prediction model based on gradient boosting, and outputting a second pollen concentration prediction result; S7: Using a voting regressor, perform weighted average calculation on the first pollen concentration prediction result and the second pollen concentration prediction result to determine a final pollen concentration prediction result.
2. The pollen concentration prediction method based on machine learning according to claim 1, characterized in that: The multi-dimensional features specifically include: Temperature, precipitation, relative humidity, air pressure, wind speed, wind direction, normalized difference vegetation index, enhanced vegetation index, leaf area index of vegetation, altitude information of pollen monitoring points and time characteristic variables.
3. The pollen concentration prediction method based on machine learning according to claim 1, characterized in that: The S2 specifically includes: Collect multi-dimensional features that affect pollen concentration; Initialize the random forest model regression model and set random_state=42; Recursively screening the multi-dimensional features by using the initialized random forest model regression model base learner through recursive feature elimination and cross-validation; In each recursion, the first importance score of each feature is calculated, and the feature with the smallest first importance score is eliminated; Repeat the iteration for a preset number of times and output the first feature subset after recursive feature elimination; Inputting the first feature subset into the initialized random forest model, calculating the second importance score of each first feature subset; screening out the features represented by the second importance score value of each first feature subset being greater than the second preset importance score value, to form a second feature subset; The intersection of the first feature subset and the second feature subset is taken to form a third feature set, and the features in the third feature set are sorted from high to low according to the importance score value, and a preset proportion of features are selected as the excellent features.
4. The method for predicting pollen concentration based on machine learning according to claim 1, characterized in that: The specific pollen concentration prediction model based on random forest is: in, represents the first pollen concentration prediction result, T represents the total number of trees, and f t () represents the prediction result of the tth decision tree, and x represents multi-dimensional feature data.
5. The method for predicting pollen concentration based on machine learning according to claim 1, characterized in that: The specific pollen concentration prediction model based on gradient boosting is: in, represents the second pollen concentration prediction result, M represents the number of iterations, η represents the learning rate, h m () represents the output of the weak learner in the mth iteration.
6. The method for predicting pollen concentration based on machine learning according to claim 1, characterized in that: After S6, it also includes: The pollen concentration prediction model based on random forest and the pollen concentration prediction model based on gradient boosting were trained through grid search cross-validation technology.
7. The method for predicting pollen concentration based on machine learning according to claim 6, characterized in that: The grid search cross-validation technology is used to train the pollen concentration prediction model based on random forest and the pollen concentration prediction model based on gradient boosting, specifically including: Initializing various parameters of the random forest pollen concentration prediction model and the gradient boosting-based pollen concentration prediction model; Build a grid searcher; Through the grid searcher, all parameter combinations are traversed, the mean square error values of each pollen concentration prediction model under each set of parameters are calculated, the mean square error values are sorted from low to high, and the parameter combinations corresponding to the mean square error values with the highest sorting are selected; Build a Bayesian optimizer; Based on the parameter combination corresponding to the top-ranked mean square error value, the Bayesian optimizer is used for optimization with the minimum mean square error value as the goal. When the termination condition is met, the iteration is stopped and the optimal hyperparameter combination is output.
8. The method for predicting pollen concentration based on machine learning according to claim 8, characterized in that: The termination conditions are specifically: When the mean square error does not exceed the preset error value during 10 consecutive iterations, the iteration is stopped.
9. The method for predicting pollen concentration based on machine learning according to claim 1, characterized in that: The S6 is specifically: The voting regressor performs a weighted average calculation on the first pollen concentration prediction result and the second pollen concentration prediction result by the following formula to determine the final pollen concentration prediction result: in, represents the final pollen concentration prediction result, α represents the weight parameter, represents the pollen concentration prediction result of the random forest regression model for the i-th sample, Represents the pollen concentration prediction result of the gradient boosting regression model for the i-th sample.
10. A pollen concentration prediction system based on machine learning, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the pollen concentration prediction method based on machine learning as described in any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Pollen discharge and concentration simulation system and method
CN120850253A