Multi-model dynamic weighted integrated optimization-based Yellow River scenery prediction method and device

Through the method of multi-model dynamic weighted integrated optimization, the accuracy and timeliness problems of the traditional Yellow River ice flood forecasting method under complex climate change were solved, and the accuracy and stability of the Yellow River ice flood forecast were improved, which is suitable for disaster prevention and mitigation work in the Yellow River basin.

CN120688030APending Publication Date: 2025-09-23HYDROLOGICAL BUREAU OF YELLOW RIVER WATER CONSERVANCY COMMISSION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510674724.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The traditional Yellow River ice flood forecasting method is limited in accuracy and timeliness when faced with complex climate and hydrological changes, and cannot meet the needs of disaster prevention and mitigation.

Method used

A method based on multi-model dynamic weighted integrated optimization is adopted. By preprocessing the Yellow River closure and ice data, key factors are extracted, and multivariate linear regression models, support vector regression models, neural networks and gradient boosting decision trees are trained. Combined with similar year retrieval and dynamic error evaluation, weighted integrated prediction is performed.

Benefits of technology

It has significantly improved the accuracy and stability of the Yellow River ice forecast, is suitable for hydrological and meteorological scenarios with limited samples and significant interannual variations, and improves the accuracy and reliability of the forecast.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688030A_ABST
    Figure CN120688030A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a Yellow River ice condition prediction method and device based on multi-model dynamic weighting integrated optimization. A specific embodiment of the method comprises the following steps: for each to-be-predicted year sample, identifying K historical year samples, taking data from which the k year samples are removed as a training set, and taking training data corresponding to the K historical year samples as a verification data set; according to the training set, the four models are retrained, and four models after training updating are obtained; respectively detecting local prediction errors of each model in similar years by using verification data sets corresponding to samples in K historical years according to the updated prediction results of the four models; determining a weighting coefficient of each model on the current training data; and according to each weighting coefficient, carrying out weighted integrated prediction on the Yellow River ice condition prediction result group. According to the embodiment, the model performance can be adaptively adjusted along with the annual scenery features, and the precision and stability of Yellow River scenery prediction are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of fake news detection, and specifically to a method and device for predicting the ice conditions of the Yellow River based on multi-model dynamic weighted integrated optimization. Background Art

[0002] The Yellow River ice flood is a natural disaster unique to the Yellow River Basin, occurring primarily in winter and spring. Affected by factors such as sudden temperature drops and frozen river channels, it can easily lead to river blockages, triggering ice flood disasters and posing a serious threat to the safety of life and property in coastal areas. In recent years, with the intensification of global climate change and the frequent occurrence of extreme weather events, the hydrological regime in the Yellow River Basin has also shown new trends. Traditional ice flood forecasting methods rely primarily on historical experience and statistical models, but their accuracy and timeliness are limited by the increasingly complex climate and hydrological changes. Therefore, there is an urgent need to introduce advanced forecasting technologies and methods to improve the accuracy and reliability of ice flood forecasts to better serve disaster prevention and mitigation efforts and ensure the sustainable social and economic development of the basin. Summary of the Invention

[0003] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] Some embodiments of the present disclosure propose a method, device, electronic device and computer-readable medium for predicting the ice conditions of the Yellow River based on multi-model dynamic weighted integrated optimization to solve the technical problems mentioned in the above background technology section.

[0005] In the first aspect, some embodiments of the present disclosure provide a method for predicting the ice conditions of the Yellow River based on multi-model dynamic weighted integrated optimization, the method comprising: pre-processing the Yellow River ice condition data within a preset time period to obtain pre-processed Yellow River ice condition data; extracting key factors from each ice condition influencing factor included in the above-mentioned pre-processed Yellow River ice condition data to obtain a key ice condition influencing factor sequence; training an initial multiple linear regression model, an initial support vector regression model, an initial neural network, and an initial gradient boosting decision tree according to the data sets corresponding to the above-mentioned key ice condition influencing factor sequence to obtain a trained multiple linear regression model, a support vector regression model, a neural network, and a gradient boosting decision tree; inputting the training data sets corresponding to the above-mentioned key ice condition influencing factor sequence into the multiple linear regression model, the support vector regression model, the neural network, and the gradient boosting decision tree, respectively. The model, neural network and gradient boosting decision tree are used to obtain the Yellow River ice situation prediction result group for the target time period; for each sample of the year to be predicted, K historical year samples are identified, and the data after excluding the K year samples are used as the training set, and the training data corresponding to the K historical year samples are used as the verification data set for retraining; according to the training set, the four models are retrained respectively to obtain the four updated models after training; using the verification data set corresponding to the K historical year samples, according to the prediction results of the four updated models, the local prediction errors of each model in similar years are detected respectively; using each local prediction error, the weighted coefficient of each model on the current training data is determined; according to each weighted coefficient, the above-mentioned Yellow River ice situation prediction result group is weighted integrated and predicted to obtain the final weighted integrated Yellow River ice situation prediction result.

[0006] In the second aspect, some embodiments of the present disclosure provide a device for predicting the ice conditions of the Yellow River based on multi-model dynamic weighted integrated optimization, the device comprising: a preprocessing unit, configured to preprocess the Yellow River ice condition data within a preset time period to obtain preprocessed Yellow River ice condition data; a key factor extraction unit, configured to extract key factors from each ice condition influencing factor included in the above-mentioned preprocessed Yellow River ice condition data to obtain a key ice condition influencing factor sequence; a single model training unit, configured to train an initial multiple linear regression model, an initial support vector regression model, an initial neural network, and an initial gradient boosting decision tree according to the data set corresponding to the above-mentioned key ice condition influencing factor sequence, to obtain a trained multiple linear regression model, a support vector regression model, a neural network, and a gradient boosting decision tree; a single model prediction unit, configured to input the training data set corresponding to the above-mentioned key ice condition influencing factor sequence into the multiple linear regression model, the support vector regression model, the neural network, and the gradient boosting decision tree, respectively. , a gradient boosting decision tree is used to obtain a group of Yellow River ice situation prediction results for the target time period; a similar year sample identification unit is configured to identify K historical year samples for each year sample to be predicted, and use the data excluding the K year samples as a training set, and use the training data corresponding to the K historical year samples as a verification data set for retraining; a model updating unit is configured to retrain the four models according to the training set to obtain four updated models; a local error detection unit is configured to use the verification data set corresponding to the K historical year samples to detect the local prediction errors of each model in similar years according to the prediction results of the four updated models; a weight determination unit is configured to use each local prediction error to determine the weight coefficient of each model on the current training data; a weighted integration unit is configured to perform weighted integrated prediction on the above-mentioned Yellow River ice situation prediction result group according to each weight coefficient to obtain a final weighted integrated Yellow River ice situation prediction result.

[0007] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0008] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.

[0009] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the Yellow River ice forecasting method based on multi-model dynamic weighted integration optimization of some embodiments of the present disclosure, based on the weighted integration strategy of similar years and dynamic error evaluation: for the first time, the "similar year" retrieval mechanism is introduced, and the historical samples closest to the characteristics of the target year are identified through weighted Euclidean distance, and each base model is locally retrained on the basis of eliminating similar samples, and the weighting coefficient is dynamically assigned based on its verification error in similar years. This method breaks through the limitations of traditional static weighting and stacking models in time adaptability, and can realize the adaptive adjustment of model performance according to annual characteristics, significantly improving the accuracy and stability of the Yellow River ice forecast, and is particularly suitable for hydrological and meteorological forecasting scenarios with limited samples and significant interannual variability. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0011] Figure 1 is a flowchart of some embodiments of the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization according to the present disclosure;

[0012] Figure 2 This is an example result diagram of variance inflation detection in the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0013] Figure 3 This is an example result diagram of robustness detection in the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0014] Figure 4 This is a fitting effect diagram of a multiple linear regression model in the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0015] Figure 5 This is a fitting effect diagram of a support vector regression model in the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0016] Figure 6 This is a fitting effect diagram of a neural network in the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0017] Figure 7This is a fitting effect diagram of a gradient boosting decision tree in the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0018] Figure 8 This is a prediction effect diagram of each model of the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0019] Figure 9 This is a weighted integrated prediction result diagram in the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0020] Figure 10 This is a comparison chart of the weighted integrated Yellow River ice forecast results and the forecast results of each model in the Yellow River ice forecast method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure;

[0021] Figure 11 Flowcharts of some embodiments of the Yellow River ice forecasting device based on multi-model dynamic weighted integrated optimization according to the present disclosure;

[0022] Figure 12 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0024] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0025] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0026] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0027] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0028] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0029] Figure 1 This is a process 100 of some embodiments of the Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization in some embodiments of the present disclosure. The Yellow River ice forecasting method based on multi-model dynamic weighted integrated optimization includes the following steps:

[0030] Step 101 , pre-processing the Yellow River river closure and ice situation data within a preset time period to obtain pre-processed Yellow River river closure and ice situation data.

[0031] In some embodiments, the execution subject (e.g., a computing device) of the Yellow River ice situation prediction method based on multi-model dynamic weighted integrated optimization pre-processes the Yellow River river closure ice situation data within a preset time period to obtain pre-processed Yellow River river closure ice situation data. For example, the preset time period may be a preset historical time period. For example, the above-mentioned execution subject may pre-process the Yellow River river closure-related ice situation data from 1987 to 2024, and sort out the following 21 influencing factors to carry out river closure forecasts, as shown in Table 1. The data from 1987 to 2020 is divided into a training set, and the data from 2021 to 2024 is divided into a validation set. For example, the first preset time period may be 1987 to 2020. The second preset time period may be 2021 to 2024. For example, pre-processing may refer to performing missing value supplementation, outlier and data standardization, and unified formatting on the Yellow River river closure ice situation data within a preset time period, calculating relevant features such as the cumulative temperature and ice flow date during the key period of the Yellow River ice situation; dividing the data into a training set and a validation set for subsequent construction of a prediction model. The pre-processing of the Yellow River closure and ice data can include the following X1-X21, a total of 21 ice influencing factors.

[0032] Table 1 Factors affecting the Yellow River closure and ice conditions

[0033]

[0034] Step 102 : extract key factors from the various ice-influencing factors included in the pre-processed Yellow River ice-closure data to obtain a key ice-influencing factor sequence.

[0035] In some embodiments, the execution entity may extract key factors from the various ice-influencing factors included in the pre-processed Yellow River closure and ice-influencing data to obtain a sequence of key ice-influencing factors.

[0036] In practice, the execution entity may extract key factors from the various ice-influencing factors included in the pre-processed Yellow River ice-closure data through the following steps:

[0037] The first step is to extract factors influencing the ice flow from the pre-processed Yellow River ice flow data to obtain an ice flow influencing factor set. Methods for extracting factors include principal component analysis (PCA), ordinary least squares (OLS), weighted least squares (WLS), generalized least squares (GLS), maximum likelihood (ML), principal axis factor (PA), and alpha factor (AF).

[0038] The second step is to conduct a preliminary screening of the above-mentioned ice condition influencing factor groups through correlation analysis to generate candidate ice condition influencing factor groups.

[0039] In practice, the above-mentioned execution entity can conduct a preliminary screening of the correlation analysis of the various ice-influencing factors included in the above-mentioned pre-processed Yellow River ice-closure data by the following steps:

[0040] First, a Pearson correlation coefficient analysis is performed on the above-mentioned group of factors influencing the winter storm to obtain a first analyzed group of factors influencing the winter storm. The Pearson correlation coefficient analysis may refer to a Pearson correlation coefficient analysis.

[0041] Second, a quality index factor analysis is performed on the above-mentioned wintering influence factor group to obtain a second analyzed wintering influence factor group. The quality index factor analysis may refer to a quality index (MI) factor analysis.

[0042] Third, a Spearman correlation coefficient analysis is performed on the above-mentioned group of factors affecting winter sports to obtain a third analyzed group of factors affecting winter sports. The Spearman correlation coefficient analysis may refer to a Spearman correlation coefficient analysis.

[0043] Fourth, a predetermined number of ice condition influencing factors are selected from the first ice condition influencing factor group, the second ice condition influencing factor group, and the third ice condition influencing factor group as candidate ice condition influencing factor groups.

[0044] As an example, the execution entity may perform a correlation analysis on the various ice-related factors included in the pre-processed Yellow River ice-related data using the standards in Table 2 below:

[0045] Table 2 Initial screening criteria for factors affecting winter weather

[0046]

[0047] For example, the execution entity can perform a correlation analysis (Pearson / Spearman / MI) on the 21 factors affecting winter snow, eliminate irrelevant factors, and retain 15 candidate factors (a group of candidate winter snow factors). The results of the correlation analysis are shown in Table 3:

[0048] Table 3 Correlation ranking of factors affecting the Yellow River closure and ice conditions

[0049]

[0050]

[0051] Taking into account the three correlation analysis rankings, six factors, X11, X12, X17, X18, X19, and X20, were eliminated.

[0052] The third step is to conduct multiple machine learning screening analyses on the above-mentioned candidate ice condition influencing factor groups to obtain a sequence of alternative ice condition influencing factors.

[0053] In practice, the above-mentioned execution entity can perform multiple machine learning screening and analysis on the above-mentioned candidate groups of impact factors through the following steps:

[0054] First, a first screening process is performed on the candidate group of ice impact factors by lasso regression to obtain a first ice impact factor sequence. For example, lasso regression can refer to Lasso regression.

[0055] Second, the candidate ice condition influencing factor group is subjected to a second screening process through random forest to obtain a second ice condition influencing factor sequence.

[0056] Third, a third screening process is performed on the candidate group of ice impact factors by a feature recursive method to obtain a third ice impact factor sequence. The feature recursive method may refer to a recursive feature elimination (RFE) method.

[0057] Fourth, a preset number of winter storm impact factors are selected from the first, second, and third winter storm impact factor sequences as candidate winter storm impact factor sequences. The preset number may be 7. That is, the top-ranked impact factors (i.e., those with the greatest impact on the prediction results) are selected from the first, second, and third winter storm impact factor sequences as candidate winter storm impact factor sequences.

[0058] As an example, we used machine learning methods (Lasso regression, random forest feature importance, and RFE) to further select the factors with the greatest predictive power for the Yellow River closure date from the remaining 15 factors influencing ice conditions. The following are the top seven factors ranked by standardized coefficients / importance using the three methods, as shown in Table 4:

[0059] Table 4 Standardized coefficient / importance ranking Correlation ranking

[0060]

[0061] Based on the results of the three machine learning methods, a total of 9 factors, including X3, X4, X6, X7, X8, X9, X14, X15, and X16, were selected for the next step of calculation.

[0062] The fourth step is to perform performance verification screening on the above-mentioned candidate ice condition influencing factor sequences to obtain the screened key ice condition influencing factor sequences.

[0063] In practice, the above-mentioned execution entities can perform performance verification and screening on the above-mentioned candidate impact factor sequences through the following steps:

[0064] First, perform variance inflation test on the above candidate flood influencing factor sequence to identify and remove highly collinear factors, and obtain the first tested flood influencing factor sequence. For example, variance inflation test (VIF test) can be performed on the candidate flood influencing factor sequence to identify and remove highly collinear factors. The factor correlation coefficient and VIF value are shown in Figure 2 The test results showed that the three factors affecting ice conditions, X7, X8, and X9, had serious collinearity problems, so these factors were eliminated.

[0065] Second, the robustness test is performed on the first detected ice impact factor sequence to remove unqualified factors and obtain the screened key ice impact factor sequence. The robustness test includes: cross-validation, sensitivity analysis, and noise introduction.

[0066] For example, cross validation, sensitivity analysis, and noise introduction are performed on the remaining six factors X3, X4, X6, X14, X15, and X16. Figure 3 In the example shown, factor X3 failed the factor stability test and was removed. After removal, the robustness test was performed again and passed. Figure 3 It shows the test results of data perturbation test, factor stability analysis, and time extrapolation ability, as well as the conclusion "failed, it is recommended to test unstable factors or increase data."

[0067] Finally, five factors were determined as prediction inputs: river closure flow (X4), ice flow date (X6), cumulative temperature ten days before river closure (X14), cumulative temperature from stable negative temperature to river closure (X15), and the duration of stable negative temperature to 119 / 46℃ (X16).

[0068] Step 103 : Based on the data set corresponding to the above-mentioned key ice impact factor sequence, an initial multiple linear regression model, an initial support vector regression model, an initial neural network, and an initial gradient boosting decision tree are trained respectively to obtain trained multiple linear regression models, support vector regression models, neural networks, and gradient boosting decision trees.

[0069] In some embodiments, the execution entity may train an initial multiple linear regression model, an initial support vector regression model, an initial neural network, and an initial gradient boosting decision tree based on the dataset corresponding to the key ice impact factor sequence, thereby obtaining a trained multiple linear regression model, support vector regression model, neural network, and gradient boosting decision tree. The neural network may be a CNN neural network model.

[0070] As an example, we train four models: the initial multiple linear regression model, the initial support vector regression model, the initial neural network model, and the gradient boosting decision tree model. The specific training methods and optimization strategies are as follows:

[0071] The multiple linear regression model uses robust regression to suppress the influence of outliers and optimizes the objective function through iterative reweighted least squares method, with the goal of minimizing the residual sum of squares under Huber weights.

[0072] The support vector regression model uses a Gaussian kernel function and automatically searches for optimal hyperparameters through Bayesian optimization. The optimization goal is to balance model complexity and prediction error. The maximum number of evaluations is set to 35, and parallel computing accelerates the search process.

[0073] The neural network constructs a double hidden layer structure (10-5 nodes), and uses a Bayesian regularization training algorithm to dynamically balance the fitting error and network complexity. The maximum number of iterations is 1000. The output layer data is individually standardized and then inverse transformed to restore the original scale.

[0074] The gradient boosting decision tree is based on a regression tree-based learner (the minimum number of samples for leaf nodes is 5). The model capacity is improved by dynamically increasing the number of iterations (200 + 100 × number of retries) and the learning rate (0.08). The goal is to gradually fit the residual gradient direction.

[0075] The general training strategy is as follows:

[0076] Retry mechanism: When the mean absolute error of the training set exceeds a threshold of 2, the training process is automatically restarted and retried up to 35 times.

[0077] Normalization: The input data is globally normalized (z-score), and the neural network output layer is additionally normalized and inversely transformed.

[0078] Dynamic optimization: SVR and GBDT improve generalization ability through adaptive adjustment of hyperparameters, while NN suppresses overfitting through regularization.

[0079] The results of fitting the training set are shown in Figure 4 、 Figure 5 、 Figure 6 、 Figure 7 . Figure 4 A graph showing the fitting effect of the multiple linear regression model. Figure 5 A graph showing the fitting effect of the support vector regression model. Figure 6 A graph showing the fitting effect of the neural network. Figure 7 It represents the fitting effect diagram of the gradient boosting decision tree, and the prediction results are Figure 8 .

[0080] In step 104, the training data sets corresponding to the above-mentioned key ice-influencing factor sequences are input into the multiple linear regression model, the support vector regression model, the neural network, and the gradient boosting decision tree, respectively, to obtain the Yellow River ice-influencing prediction result set for the target time period.

[0081] In some embodiments, the above-mentioned execution entity can input the training data set corresponding to the above-mentioned key ice-influencing factor sequence into the multiple linear regression model, the support vector regression model, the neural network, and the gradient boosting decision tree, respectively, to obtain the Yellow River ice-condition prediction result group for the target time period. The Yellow River ice-condition prediction result group includes four Yellow River ice-condition prediction results, wherein the Yellow River ice-condition prediction results correspond to the time nodes in the target time period (river closure date) (the time nodes can be years). For example, the Yellow River ice-condition prediction results include sub-prediction results for each time node from 2021 to 2024. Figure 8 In the example shown, the blue-black lines represent the actual values; the yellow lines represent the output values ​​of the multiple linear regression model; the cyan lines represent the output values ​​of the SVR support vector regression model; and the brown lines represent the output values ​​of the gradient boosting decision tree.

[0082] Step 105: for each year sample to be predicted, identify K historical year samples, use the data excluding the K year samples as the training set, and use the training data corresponding to the K historical year samples as the verification data set for retraining.

[0083] In some embodiments, the execution entity may identify K historical year samples for each year sample to be predicted, use the data excluding these K year samples as a training set, and retrain using the training data corresponding to each of the K historical year samples as a validation data set. Specifically, the four models (multivariate linear regression model, support vector regression model, neural network, and gradient boosting decision tree) are retrained using the training set.

[0084] For example, the execution entity may dynamically identify K historical year samples similar to the sample of the year to be predicted for each sample of the year to be predicted, and use the training data corresponding to the K historical year samples as a verification data set.

[0085] For example, the calculation method of the Euclidean distance metric weighted by factor importance is as follows: Suppose there are two feature vectors X of the year i =(X i1 , X i2 ,…X ip ) and X i =(X j1 , X j2 ,…X jp ), where p represents the number of features, and the corresponding regression coefficient of the training set multivariate linear regression model is β = (β1, β2, ... β p ). Then the weighted Euclidean distance d ij Defined as:

[0086]

[0087] Among them, the weight ω k It is usually taken as the absolute value of the regression coefficient, that is, ω k =|β k The intuitive explanation of this weighting method is that the absolute value of the regression coefficient reflects the importance of the corresponding feature to the predicted target. Therefore, when calculating the distance, features with higher importance are given greater weights, making the distance metric more consistent with the actual situation.

[0088] The training set is an updated training data set obtained by removing the training data of K historical year samples from the training data set corresponding to the above-mentioned key ice impact factor sequence.

[0089] Step 106: retrain the four models according to the training set to obtain four updated models.

[0090] In some embodiments, the execution entity may retrain the four models based on the training set to obtain four updated models. The four models may be a multiple linear regression model, a support vector regression model, a neural network, and a gradient boosting decision tree. The training set may be an updated training data set.

[0091] For example, the above-mentioned execution entity can retrain the multiple linear regression model, support vector regression model, neural network, and gradient boosting decision tree respectively based on the updated training data set to obtain the trained and updated multiple linear regression model, support vector regression model, neural network, and gradient boosting decision tree as the updated multiple linear regression model, updated support vector regression model, updated neural network, and updated gradient boosting decision tree. In order to improve the adaptability of the model to the characteristic structure of the predicted year, after removing the similar year series data identified in S7, the four types of base models (linear regression, support vector regression, neural network, GBDT) are locally retrained on the remaining samples to simulate the learning ability of each model under a "non-similar background".

[0092] The optimization and update strategies for each model are as follows: the multivariate linear regression model uses robust regression to reduce outlier sensitivity by minimizing the mean absolute error (MAE) and extracts a normalized weight vector for feature importance analysis. The support vector regression model automatically tunes kernel function parameters (Gaussian kernel bandwidth, penalty coefficient, and ε-insensitive region) through a Bayesian optimization framework, and combines parallel computing to accelerate hyperparameter search, balancing model complexity and generalization ability within 30 iterations. The neural network uses Bayesian regularization training to constrain the distribution of weight parameters, combined with a [10,5] two-layer hidden layer structure to effectively suppress overfitting while retaining the ability to express nonlinear features. The gradient boosting decision tree uses MAE as the splitting criterion, sets a 200-round weak learner (decision tree) ensemble, and enhances the accuracy of gradient direction correction by shrinking the step size by 0.1 and the leaflet node size (MinLeafSize = 5).

[0093] Step 107 , using the validation data set corresponding to the K historical year samples, and based on the prediction results of the four updated models, respectively detect the local prediction errors of each model in similar years.

[0094] In some embodiments, the execution entity may utilize a validation dataset corresponding to K historical year samples to detect the local prediction errors of each model in similar years based on the prediction results of the four updated models.

[0095] For example, the aforementioned execution entity can use a validation dataset corresponding to K historical year samples to test the local prediction errors of the updated multivariate linear regression model, the updated support vector regression model, the updated neural network model, and the updated gradient boosting decision tree model. Using K similar years as the validation dataset, the local prediction error of each model is evaluated, that is, the model's performance in a scenario similar to the current year.

[0096] Step 108: Using each local prediction error, determine the weight coefficient of each model on the current training data.

[0097] In some embodiments, the execution entity may use each local prediction error to determine a weighted coefficient for each model on the current training data.

[0098] As an example, the weight ω of a single prediction model j j It can be calculated by the following formula:

[0099]

[0100] Among them: MAE j is the mean absolute error of model j in the validation set of similar years; m is the total number of base models (multiple linear regression model, support vector regression model, neural network, gradient boosting decision tree) (4).

[0101] Step 109: Perform weighted integrated prediction on the above-mentioned Yellow River ice situation prediction result group according to each weighting coefficient to obtain the final weighted integrated Yellow River ice situation prediction result.

[0102] In some embodiments, the above-mentioned execution entity can perform weighted integrated prediction on the above-mentioned Yellow River ice situation prediction result group according to various weighting coefficients to obtain the final weighted integrated Yellow River ice situation prediction result.

[0103] As an example, the final weighted integrated Yellow River ice forecast results Calculate using the following formula:

[0104]

[0105] in: is the prediction result of model j for the test sample.

[0106] like Figure 9-10 As an example, Figure 9 This is the integrated prediction result graph; Figure 10 This is a comparison chart of the weighted integrated Yellow River ice forecast results and the single model forecast results.

[0107] Further reference is made to Table 5, which shows the forecast error estimates for different models:

[0108] Table 5. Forecast error evaluation of different models

[0109]

[0110]

[0111] The MAE values ​​of the four single models on the test set were all less than 3 days, indicating that these models have qualified prediction capabilities overall. However, further analysis found that the SVR and GBDT models had large deviations in the predictions for 2021 and 2023, resulting in relatively high root mean square errors (RMSE) and low coefficient of determination (R 2 ) is slightly lower, reflecting the lack of forecast stability of these two models in specific years.

[0112] In contrast, the weighted ensemble prediction model based on similar years and dynamic error assessment outperformed the single model across all evaluation metrics, demonstrating greater accuracy and stability. This model achieved a MAE of 1.72 days, significantly lower than that of the single model, and also had the lowest RMSE, demonstrating high accuracy across all years. This demonstrates that dynamically weighted ensemble predictions from different models offer excellent accuracy and can effectively improve the performance of the model for predicting river closure dates.

[0113] Further references Figure 11 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a device for predicting the ice conditions of the Yellow River based on multi-model dynamic weighted integrated optimization. Figure 1 Corresponding to the method embodiments shown, the Yellow River ice situation prediction device based on multi-model dynamic weighted integrated optimization can be specifically applied to various electronic devices.

[0114] like Figure 11As shown, some embodiments of the Yellow River ice situation prediction device 1100 based on multi-model dynamic weighted integrated optimization include: a preprocessing unit 1101, a key factor extraction unit 1102, a single model training unit 1103, a single model prediction unit 1104, a similar year sample identification unit 1105, a model updating unit 1106, a local error detection unit 1107, a weight determination unit 1108 and a weighted integration unit 1109. Among them, the preprocessing unit 1101 is configured to preprocess the Yellow River closure and ice situation data within a preset time period to obtain preprocessed Yellow River closure and ice situation data; the key factor extraction unit 1102 is configured to extract key factors from each ice situation influencing factor included in the above-mentioned preprocessed Yellow River closure and ice situation data to obtain a key ice situation influencing factor sequence; the single model training unit 1103 is configured to train the initial multiple linear regression model, the initial support vector regression model, the initial neural network, and the initial gradient boosting decision tree according to the data sets corresponding to the above-mentioned key ice situation influencing factor sequence, to obtain the trained multiple linear regression model, support vector regression model, neural network, and gradient boosting decision tree; the single model prediction unit 1104 is configured to input the training data sets corresponding to the above-mentioned key ice situation influencing factor sequence into the multiple linear regression model, support vector regression model, neural network, and gradient boosting decision tree, respectively, to obtain the Yellow River ice situation prediction for the target time period. result group; the similar year sample identification unit 1105 is configured to identify K historical year samples for each year sample to be predicted, and use the data after excluding the k year samples as the training set, and use the various training data corresponding to the K historical year samples as the verification data set for retraining; the model updating unit 1106 is configured to retrain the four models according to the training set to obtain the four updated models; the local error detection unit 1107 is configured to use the verification data set corresponding to the K historical year samples to detect the local prediction errors of each model in similar years according to the prediction results of the four updated models; the weight determination unit 1108 is configured to use each local prediction error to determine the weight coefficient of each model on the current training data; the weighted integration unit 1109 is configured to perform weighted integrated prediction on the above-mentioned Yellow River ice situation prediction result group according to each weight coefficient to obtain the final weighted integrated Yellow River ice situation prediction result.

[0115] It is understandable that the various units recorded in the Yellow River Ice Situation Prediction Device 1100 based on multi-model dynamic weighted integrated optimization are similar to those in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the Yellow River Ice Situation Prediction Device 1100 based on multi-model dynamic weighted integrated optimization and the units contained therein, and will not be repeated here.

[0116] Reference below Figure 12 , which shows a schematic structural diagram of an electronic device (such as a computing device) suitable for implementing some embodiments of the present disclosure. Figure 12 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure. Figure 12 As shown, the computer device includes a processor, a memory and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can enable the processor to execute any one of the Yellow River ice forecasting methods based on multi-model dynamic weighted integrated optimization. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium, which, when executed by the processor, can enable the processor to execute any one of the Yellow River ice forecasting methods based on multi-model dynamic weighted integrated optimization. The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 12 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present disclosure, and does not constitute a limitation on the computer device to which the solution of the present disclosure is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0117] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0118] Wherein, in one embodiment, the above-mentioned processor is used to run a computer program stored in a memory to implement the following steps: pre-processing the Yellow River river closure and ice situation data within a preset time period to obtain pre-processed Yellow River river closure and ice situation data; extracting key factors from each ice situation influencing factor included in the above-mentioned pre-processed Yellow River river closure and ice situation data to obtain a key ice situation influencing factor sequence; training an initial multiple linear regression model, an initial support vector regression model, an initial neural network, and an initial gradient boosting decision tree according to the data sets corresponding to the above-mentioned key ice situation influencing factor sequence to obtain a trained multiple linear regression model, a support vector regression model, a neural network, and a gradient boosting decision tree; inputting the training data sets corresponding to the above-mentioned key ice situation influencing factor sequence into the multiple linear regression model, the support vector regression model, the neural network, and the gradient boosting decision tree, respectively. In the neural network and gradient boosting decision tree, a group of Yellow River ice situation prediction results for the target time period is obtained; for each sample of the year to be predicted, K historical year samples are identified, and the data after excluding the K year samples are used as the training set, and the training data corresponding to the K historical year samples are used as the verification data set for retraining; according to the training set, the four models are retrained respectively to obtain four updated models after training; using the verification data set corresponding to the K historical year samples, according to the prediction results of the four updated models, the local prediction errors of each model in similar years are detected respectively; using each local prediction error, the weighted coefficient of each model on the current training data is determined; according to each weighted coefficient, the above-mentioned Yellow River ice situation prediction result group is weighted integrated prediction to obtain the final weighted integrated Yellow River ice situation prediction result.

[0119] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions. The method implemented when the program instructions are executed can refer to the various embodiments of the Yellow River ice situation prediction method based on multi-model dynamic weighted integrated optimization disclosed in the present disclosure.

[0120] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., provided on the computer device.

[0121] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0122] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for predicting ice conditions on the Yellow River based on multi-model dynamic weighted integrated optimization, characterized in that: include: Preprocessing the Yellow River closure and ice data within a preset time period to obtain preprocessed Yellow River closure and ice data; Extract key factors from the pre-processed Yellow River closure and ice data to obtain a sequence of key ice influencing factors; According to the data set corresponding to the key ice impact factor sequence, an initial multiple linear regression model, an initial support vector regression model, an initial neural network, and an initial gradient boosting decision tree are trained respectively to obtain a trained multiple linear regression model, a support vector regression model, a neural network, and a gradient boosting decision tree; The training data sets corresponding to the key ice-influencing factor sequences are input into the multiple linear regression model, support vector regression model, neural network, and gradient boosting decision tree respectively to obtain the Yellow River ice-influencing prediction result set for the target time period; For each year sample to be predicted, K historical year samples are identified, the data after excluding the K year samples are used as the training set, and the training data corresponding to the K historical year samples are used as the validation data set for retraining; According to the training set, the four models are retrained respectively to obtain four updated models; Using the validation data set corresponding to K historical year samples, the local prediction errors of each model in similar years are tested based on the prediction results of the four updated models. Using each local prediction error, determine the weight coefficient of each model on the current training data; According to each weighting coefficient, the Yellow River ice situation prediction result group is subjected to weighted integrated prediction to obtain the final weighted integrated Yellow River ice situation prediction result.

2. The method according to claim 1, characterized in that The key factors are extracted from the pre-processed Yellow River closure and ice data to obtain a sequence of key ice influencing factors, including: Extracting ice-condition influencing factors from the pre-processed Yellow River ice-condition data to obtain an ice-condition influencing factor group; Performing a correlation analysis on the ice condition influencing factor group to generate a candidate ice condition influencing factor group; Perform multiple machine learning screening analysis on the candidate group of ice condition influencing factors to obtain a sequence of candidate ice condition influencing factors; The candidate ice condition influencing factor sequences are subjected to efficacy verification screening to obtain a screened key ice condition influencing factor sequence.

3. The method according to claim 2, characterized in that The preliminary screening of the group of factors influencing the freezing situation by correlation analysis to generate a group of candidate factors influencing the freezing situation includes: Performing Pearson correlation coefficient analysis on the ice condition influencing factor group to obtain a first analysis ice condition influencing factor group; Performing a quality index factor analysis on the ice impact factor group to obtain a second analysis ice impact factor group; Performing Spearman correlation coefficient analysis on the wintering influence factor group to obtain a third analysis wintering influence factor group; A predetermined number of ice condition influencing factors are selected from the first analysis ice condition influencing factor group, the second analysis ice condition influencing factor group, and the third analysis ice condition influencing factor group as candidate ice condition influencing factor groups.

4. The method according to claim 2, characterized in that The multiple machine learning screening analysis is performed on the candidate group of ice condition influencing factors to obtain a sequence of candidate ice condition influencing factors, including: Performing a first screening process on the candidate group of ice impact factors by lasso regression to obtain a first ice impact factor sequence; Performing a second screening process on the candidate group of ice condition influencing factors through random forest to obtain a second ice condition influencing factor sequence; Performing a third screening process on the candidate group of ice impact factors in a feature recursive manner to obtain a third ice impact factor sequence; A preset number of ice condition influencing factors are selected from the first ice condition influencing factor sequence, the second ice condition influencing factor sequence, and the third ice condition influencing factor sequence as candidate ice condition influencing factor sequences.

5. The method according to claim 3, characterized in that The method of performing performance verification screening on the candidate ice condition influencing factor sequences to obtain the screened key ice condition influencing factor sequences comprises: Performing a variance inflation test on the candidate ice impact factor sequence to identify and remove highly collinear factors, thereby obtaining a first test ice impact factor sequence; A robustness test is performed on the first detected ice condition influencing factor sequence to remove unqualified factors and obtain a screened key ice condition influencing factor sequence.

6. A device for predicting ice conditions on the Yellow River based on multi-model dynamic weighted integrated optimization, characterized in that: include: A preprocessing unit is configured to preprocess the Yellow River closure and ice data within a preset time period to obtain preprocessed Yellow River closure and ice data; a key factor extraction unit configured to extract key factors from each ice-influencing factor included in the pre-processed Yellow River ice-closure data to obtain a key ice-influencing factor sequence; A single model training unit is configured to train an initial multiple linear regression model, an initial support vector regression model, an initial neural network, and an initial gradient boosting decision tree according to a data set corresponding to the sequence of key ice impact factors, to obtain trained multiple linear regression models, support vector regression models, neural networks, and gradient boosting decision trees; A single model prediction unit is configured to input the training data set corresponding to the key ice-influencing factor sequence into a multiple linear regression model, a support vector regression model, a neural network, and a gradient boosting decision tree, respectively, to obtain a set of Yellow River ice-influencing prediction results for the target time period; The similar year sample identification unit is configured to identify K historical year samples for each year sample to be predicted, use the data after excluding the K year samples as a training set, and use the training data corresponding to the K historical year samples as a validation data set for retraining; The model updating unit is configured to retrain the four models according to the training set to obtain four updated models; The local error detection unit is configured to use the validation data set corresponding to the K historical year samples and detect the local prediction errors of each model in similar years based on the prediction results of the four updated models; a weight determination unit configured to determine a weight coefficient of each model on current training data using each local prediction error; The weighted integration unit is configured to perform weighted integrated prediction on the Yellow River ice situation prediction result group according to various weighting coefficients to obtain the final weighted integrated Yellow River ice situation prediction result.

7. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer-readable medium, characterized in that A computer program is stored thereon, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.