Control Method for Observation Quality of Precipitation Data by Integrating Multiple Machine Learning Models

By generating an adversarial network to repair precipitation data anomalies and combining dynamic switching of multiple machine learning models, the problem of equipment failure in precipitation data quality control and insufficient adaptability of a single model is solved, and high-precision observation and prediction under complex meteorological conditions are achieved.

CN119442099BActive Publication Date: 2025-07-22陕西省气象信息中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411516618.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-07-22
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

The existing meteorological observation methods have problems such as equipment failure, signal interference, insufficient acquisition frequency and insufficient adaptability of a single model in terms of precipitation data quality control, especially in complex and extreme weather conditions, which lacks real-time feedback mechanisms, resulting in insufficient accuracy of prediction results.

Method used

Generative adversarial networks are used to repair the anomalies in precipitation data, and dynamic switching is combined with multiple machine learning models (such as linear regression, random forests and long-term memory networks). Through meteorological factor complexity scoring and feedback optimization mechanisms, closed-loop adjustment is formed to improve observation quality and prediction accuracy.

Benefits of technology

It realizes high-precision precipitation data observation and real-time adjustment under different meteorological conditions, provides reliable meteorological prediction support, ensuring the continuity of observation data and the accuracy of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119442099B_ABST
    Figure CN119442099B_ABST
Patent Text Reader

Abstract

The precipitation data observation quality control method integrating multiple machine learning models provided by this application relates to the field of meteorological observation technology. By obtaining precipitation data in the pre-observation area, the abnormal part is obtained through observation quality assessment; a generative adversarial network is used to repair the abnormal part in the obtained precipitation data to obtain the repaired precipitation data; based on the repaired precipitation data, meteorological factors for complexity analysis are extracted; complexity analysis is performed based on the extracted meteorological factors, and a complexity score of the meteorological factors is generated; based on the complexity score, the discrimination path of the current precipitation data is determined. Through a series of steps such as meteorological factor extraction, generative adversarial network repair, complexity analysis, and dynamic switching of the discrimination path, this method realizes high-precision observation and real-time adjustment of precipitation data under different meteorological conditions, providing reliable data support for meteorological prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of meteorological observation, and particularly to a control method for the observation quality of precipitation data integrating multiple machine learning models. Background Art

[0002] In the field of meteorological observation, the accurate observation and prediction of precipitation data are of great significance for disaster prevention and reduction, agricultural management, and water resource allocation.

[0003] However, limited by the complexity of precipitation observation and the dynamic changes of the external environment, there are many problems in the quality control of precipitation data in existing meteorological observations.

[0004] Firstly, precipitation observation data often shows missing or abnormal due to equipment failures, signal interference, or insufficient acquisition frequency, affecting the reliability of the data.

[0005] In addition, with the complication of meteorological conditions, it is difficult to adapt to changing weather conditions by only using a single model to analyze and predict precipitation data, resulting in insufficient accuracy of prediction results.

[0006] Currently, some precipitation observation methods enhance data analysis capabilities by introducing machine learning models, but there are still deficiencies in model switching and optimization.

[0007] Especially when dealing with complex and extreme weather, there is a lack of adjustment means for real-time feedback mechanisms and it is impossible to adjust the acquisition strategy and model parameters in real time according to the error size and data fluctuations.

[0008] Therefore, there is an urgent need to propose a control method for the observation quality of precipitation data integrating multiple machine learning models. Summary of the Invention

[0009] The present invention provides a control method for the observation quality of precipitation data integrating multiple machine learning models, aiming to solve the technical problem that it is difficult to adapt to changing weather conditions when using a single model to analyze and predict precipitation data in related technologies, avoid subsequent problems with poor model switching and optimization, and solve the practical problem of lacking adjustment means for real-time feedback mechanisms when dealing with complex and extreme weather.

[0010] To achieve the above object, the embodiments of the present application are implemented as follows:

[0011] In a first aspect, the embodiments of the present application provide a control method for the observation quality of precipitation data integrating multiple machine learning models, including the following:

[0012] Obtain precipitation data of the pre-observation area, conduct an observation quality assessment, and determine the abnormal part in the precipitation data based on the quality assessment result;

[0013] Use a generative adversarial network to repair the abnormal parts in the obtained precipitation data, and obtain the repaired precipitation data;

[0014] Based on the repaired precipitation data, extract meteorological factors for complexity analysis;

[0015] Conduct complexity analysis based on the extracted meteorological factors, and generate a complexity score for the meteorological factors;

[0016] Based on the complexity score, determine the discrimination path of the current precipitation data. The discrimination path includes the first discrimination path, the second discrimination path, and the third discrimination path;

[0017] The first discrimination path: The complexity score of the meteorological factors is in the low complexity interval;

[0018] The second discrimination path: The complexity score of the meteorological factors is in the medium complexity interval;

[0019] The third discrimination path: The complexity score of the meteorological factors is in the high complexity interval;

[0020] Dynamically switch to the corresponding machine learning model based on the discrimination path, and predict the precipitation data based on the switched machine learning model to generate a prediction result;

[0021] Compare the prediction result with the repaired precipitation data to calculate the prediction error;

[0022] Generate a feedback optimization mechanism based on the obtained prediction error, and generate an optimization plan according to the feedback optimization mechanism to optimize the future precipitation observation and prediction process;

[0023] Based on the error discrimination path, perform error discrimination optimization and adjustment, apply the optimization and adjustment to the next round of observation. After the next round of observation is completed, calculate the error again to form a closed loop.

[0024] Further, by obtaining precipitation data in the pre-observation area, conducting observation quality assessment to obtain the abnormal part, and determining the abnormal part in the precipitation data based on the quality assessment result; using a generative adversarial network to repair the abnormal part in the obtained precipitation data to obtain the repaired precipitation data; extracting meteorological factors for complexity analysis based on the repaired precipitation data; conducting complexity analysis based on the extracted meteorological factors and generating a complexity score for the meteorological factors; determining the discrimination path of the current precipitation data based on the complexity score; dynamically switching to the corresponding machine learning model based on the discrimination path, and predicting the precipitation data based on the switched machine learning model to generate a prediction result; comparing the prediction result with the repaired precipitation data to calculate the prediction error; generating a feedback optimization mechanism based on the obtained prediction error, and generating an optimization plan according to the feedback optimization mechanism to optimize the future precipitation observation and prediction process; performing error discrimination optimization adjustment based on the error discrimination path, applying the optimization adjustment to the next round of observation, and calculating the error again after the next round of observation is completed to form a closed loop. Through a series of steps such as meteorological factor extraction, generative adversarial network repair, complexity analysis, and dynamic switching of the discrimination path, this method effectively improves the quality control level of precipitation data observation, can achieve high-precision observation and real-time adjustment of precipitation data under different meteorological conditions, and provides reliable data support for meteorological prediction.

[0025] In some specific embodiments, obtaining precipitation data in the pre-observation area, conducting observation quality assessment, and determining the abnormal part in the precipitation data based on the quality assessment result specifically include:

[0026] For the observation quality assessment, by calculating the deviation degree of the observed values of the meteorological factors existing in the precipitation data, the quality of the precipitation data is determined, specifically as:

[0027]

[0028] In the formula, Q i represents the quality assessment value of the i-th meteorological factor, X i represents the observed value of the i-th meteorological factor, μ i represents the historical mean of the i-th meteorological factor, σ i represents the historical standard deviation of the i-th meteorological factor, and i represents the i-th meteorological factor.

[0029] Based on the quality assessment result, to determine the abnormal part in the precipitation data, by setting an abnormal threshold τ, when Q i exceeds the set abnormal threshold, it is regarded as abnormal data, specifically as:

[0030]

[0031] In the formula, |Q i$Q_i$ represents the absolute quality assessment value of the $i$-th meteorological factor, $\tau$ represents the set anomaly threshold, an output of 1 indicates that the data is determined to be abnormal, and an output of 0 indicates that the data is normal.

[0032] Furthermore, the observation quality assessment determines the overall quality of precipitation data by calculating the deviation degree of the meteorological factor observation value from its historical mean, and then identifies the anomalies in the data. This assessment method uses standardized deviation calculation, making the current observation value of the meteorological factor more clearly show the degree of difference compared with the long-term historical characteristics. In this way, the anomaly points that may be caused by sensor errors or environmental changes in the data can be quickly identified, providing a reliable basis for subsequent data repair and analysis.

[0033] In some specific embodiments, a generative adversarial network is used to repair the abnormal part of the obtained precipitation data to obtain the repaired precipitation data. The repair formula is:

[0034]

[0035] In the formula, represents the precipitation data of the $i$-th meteorological factor after repair, represents the repair value of the $i$-th meteorological factor generated by the generator, $\lambda$ is the fusion parameter, represents the original observation data of the $i$-th meteorological factor.

[0036] Among them, $0 < \lambda < 1$, which is used to control the weight distribution between the generated data and the original data.

[0037] Furthermore, through the repaired precipitation data, this method can effectively reduce the impact of abnormal data on the precipitation observation quality. The repair process uses a generative adversarial network to generate reasonable repair values, and through the fusion parameter, the weighted integration of the generated data and the original observation data is realized, so that the repaired data not only retains the characteristics of the original observation, but also can make up for the missing part brought by the abnormal value through the generated data. This repair method can not only ensure the continuity and stability of the observation data during abnormal data fluctuations, but also significantly improve the reliability of the data, providing more accurate basic data for subsequent complexity analysis and meteorological prediction.

[0038] In some specific embodiments, based on the repaired precipitation data, meteorological factors for complexity analysis are extracted. The extraction formula is:

[0039]

[0040] In the formula, represents the correlation between the $i$-th meteorological factor $X$ i and the precipitation intensity $R$, $cov(X$ i , $R)$ represents $X$ iThe covariance between and R, and σ R respectively represent the standard deviations of X i and R. Based on the calculation results, meteorological factors for complexity analysis are extracted.

[0041] Furthermore, by extracting meteorological factors highly correlated with precipitation intensity, this method can screen out the key factors most valuable for complexity analysis from the repaired precipitation data. Using correlation analysis combining covariance and standard deviation can effectively ensure that the extracted meteorological factors are not only closely related to precipitation intensity but also can more intuitively reflect the variation law of precipitation data under different meteorological conditions. This extraction method greatly reduces the interference of redundant factors on complexity scoring and subsequent model prediction, improving the efficiency of data processing and the accuracy of analysis.

[0042] In some specific embodiments, complexity analysis is performed based on the extracted meteorological factors, and a complexity score of the meteorological factors is generated. The calculation formula is:

[0043]

[0044] In the formula, C is the complexity score of the meteorological factor, X i represents the current observed value of the i-th meteorological factor, μ i represents the historical mean of the i-th meteorological factor, σ i represents the historical standard deviation of the i-th meteorological factor, w i is the weight of each meteorological factor, i represents the i-th meteorological factor, i traverses all meteorological factors from 1 to n, and each i corresponds to a specific meteorological factor; n represents the total number of samples.

[0045] Furthermore, by performing complexity analysis on the extracted meteorological factors and generating a complexity score, this method can accurately evaluate the complexity of meteorological data under multi-factor conditions. By introducing the historical mean and standard deviation to measure the deviation degree of the current observed value, and then comprehensively calculating the complexity score in combination with the weights of each meteorological factor, the scoring result can more intuitively reflect the volatility and uncertainty of the current meteorological conditions. This complexity analysis based on deviation degree and weight effectively improves the adaptability under complex meteorological conditions and provides a scientific basis for the dynamic switching of the model.

[0046] In some specific embodiments, based on the complexity score, the discrimination path of the current precipitation data is determined, specifically including:

[0047] The first discrimination path means that when the complexity score of the meteorological factor is in the low-complexity interval, in the low-complexity interval, linear regression is selected as the machine learning model according to the first discrimination path, and a prediction result is generated based on the selected linear regression model.

[0048] Furthermore, the low-complexity interval indicates that the current meteorological conditions are stable, with little precipitation and no precipitation. Under this interval state, precipitation weather phenomena exhibit linear characteristics, and a linear regression model that can quickly process linear characteristics is used for prediction. The definition of the low-complexity interval is: the complexity score C of the obtained meteorological factors is ≤ 30.

[0049] The second discrimination path indicates that the complexity score of the current meteorological factors is in the medium-complexity interval. When in the medium-complexity interval, a random forest is selected as the machine learning model according to the second determination path, and a prediction result is generated based on the selected random forest model.

[0050] Furthermore, the medium-complexity interval indicates that there are certain changes in the current meteorological conditions, but they have not reached extreme conditions. Under this interval state, precipitation weather phenomena will exhibit non-linear characteristics, and a random forest model that can handle non-linear relationships is used for prediction processing. The definition of the medium-complexity interval is: 30 < C ≤ 70.

[0051] The third discrimination path indicates that the complexity score of the current meteorological factors is in the high-complexity interval. When in the high-complexity interval, a long short-term memory network is selected as the machine learning model according to the third discrimination path, and a prediction result is generated based on the selected long short-term memory network model.

[0052] Furthermore, the high-complexity interval indicates that the meteorological conditions are very complex, with large and rapidly changing precipitation intensity. Under this interval state, while the precipitation intensity exhibits non-linear characteristics, it also changes rapidly. A long short-term memory network model that can handle complex non-linear relationships and adapt to rapidly changing weather conditions is used for prediction processing. The high-complexity interval is defined as: C > 70.

[0053] In some specific embodiments, based on the discrimination path, it is dynamically switched to the corresponding machine learning model, and the precipitation data is predicted based on the switched machine learning model to generate a prediction result. Specifically, it includes:

[0054] A discrimination interval is set for the discrimination path, and the complexity score C is set to vary within the range of [0, 100]. Based on the score, it is switched between different models. The discrimination intervals include:

[0055] C ≤ 30 represents the discrimination interval of the first discrimination path; 30 < C ≤ 70 represents the discrimination interval of the second discrimination path; C > 70 represents the discrimination interval of the third discrimination path.

[0056] The weight W of each model is defined by the complexity score C i (C). Based on the preset meteorological complexity interval, the weights of each model are adjusted to determine the contribution of each model to the final prediction result, and the weight W of each model is defined i(C) A function that varies with the meteorological complexity score C, adjusting the contribution of the model through weighted smooth transition. Let W1(C) be the weight of the linear regression model, W2(C) be the weight of the random forest model, and W3(C) be the weight of the long short-term memory network. The switching formula is as follows:

[0057]

[0058] W1(C) represents the weight of the linear regression model, with a weight of 1 when C ≤ 30 and gradually decreasing to 0; W2(C) represents the weight of the random forest model, gradually increasing when 30 < C ≤ 70, increasing from 0 to 1, and then gradually decreasing to 0 when C > 70; W3(C) represents the weight of the long short-term memory network, gradually increasing when C > 70. C represents the complexity score of the meteorological factor.

[0059] Furthermore, by determining the path based on the complexity score and switching to different machine learning models, this method can effectively adapt to changing meteorological conditions, making the selection of the prediction model more flexible. The setting of the sub-intervals of the complexity score allows for the dynamic selection of the optimal model type under simple, general, and complex meteorological conditions, avoiding the problem that a single model cannot adapt to different meteorological complexities. By assigning different weights to different models, it is possible to give priority to using linear models at low complexity, switch to non-linear models such as random forests at medium complexity, and handle complex meteorological changes through deep models such as long short-term memory networks at high complexity, ensuring the accuracy and real-time nature of the prediction results.

[0060] In some specific embodiments, the prediction result is compared with the repaired precipitation data to calculate the prediction error. The error calculation formula is:

[0061]

[0062] In the formula, MSE represents the mean square error, n represents the total number of samples, represents the prediction error at the t-th moment, represents the precipitation prediction value at the t-th moment, represents the actual precipitation at the t-th moment.

[0063] Furthermore, by comparing the prediction result with the repaired precipitation data, this method calculates the prediction error using the mean square error to quantify the deviation between the prediction result and the actual observed data. The mean square error calculation method can not only provide an intuitive measure of the overall error but also emphasize the impact of larger deviations, thereby triggering further optimization when the prediction deviation is significant.

[0064] In some specific implementations, a feedback optimization mechanism is generated based on the obtained prediction error, and an optimization scheme is generated according to the feedback optimization mechanism to optimize the future precipitation observation and prediction process, specifically including:

[0065] pass The error E is calculated, and the error threshold ∈ is set based on the error E. The path to which the current error belongs is determined based on the error threshold ∈, specifically:

[0066]

[0067] Wherein, ∈1 represents the slight error threshold, ∈2 represents the significant error threshold, and E represents the error. Path 1: when E<∈1, the error is below the slight error threshold, and no adjustment is performed according to the path to which it belongs. Path 2: when ∈1≤E<∈2, the error is above the slight threshold but has not reached the significant error threshold, and a slight adjustment is performed. A slight adjustment is to fine-tune the parameters of the machine learning model involved in the prediction. Path 3: when E≥∈2, significant adjustment measures are triggered, and significant adjustments include increasing the data collection frequency and sensor calibration.

[0068] Furthermore, by generating a feedback optimization mechanism and performing path discrimination based on the size of the prediction error, this method can flexibly take corresponding adjustment measures for different error levels, thereby achieving dynamic optimization of future precipitation observation and prediction processes. By classifying the error paths, the prediction settings can be kept unchanged when the error is small to avoid unnecessary adjustments; when the error increases to a certain extent, slight or significant adjustments are triggered to optimize the model parameters or data collection frequency respectively. This discrimination and adjustment mechanism ensures computational efficiency while flexibly responding to the uncertainty caused by error fluctuations.

[0069] In some specific implementations, based on the error discrimination path, error discrimination optimization adjustment is performed, and the optimization adjustment is applied to the next round of observations. After the next round of observations is completed, the error is calculated again to form a closed loop, specifically:

[0070] The adjustment strategy selected according to the error path is automatically adjusted, and the adjusted acquisition frequency, model configuration and sensor calibration information are applied to the next round of observations. If the error is reduced and maintained below the slight error threshold, it proves that the current optimization is effective. Based on the effective optimization, the current acquisition frequency and model configuration continue to be used. If the error is still greater than the significant error threshold, the error source is re-evaluated until the error is adjusted below the slight error threshold.

[0071] Furthermore, through the optimization adjustment and automated application based on the error discrimination path, this method forms a feedback closed-loop mechanism in the process of precipitation observation and prediction, which can gradually improve the prediction accuracy in different observation rounds. After each round of observation, automated adjustment is performed according to the error path, and the adjusted acquisition frequency, model configuration, and sensor calibration parameters are applied to the next round of observation, thereby maintaining a dynamic balance between data quality and prediction accuracy. If the error decreases after adjustment and remains below the minor error threshold, it is determined that the current optimization is effective, and the current parameter settings are continued to ensure the stability and efficiency of the observation process.

[0072] The beneficial effects of the technical solution of the present invention are as follows: The generative adversarial network is used to repair the abnormal observed values in the precipitation data. By fusing the generated data with the original data, the reliability of the data is ensured, providing high-quality observed data for subsequent analysis; Based on the extracted meteorological factors, complexity scoring is performed to dynamically identify changes in meteorological conditions and automatically switch to a machine learning model suitable for the current complexity. This dynamic switching mechanism improves the adaptability of the model, responds quickly under low-complexity conditions, and ensures prediction accuracy under high-complexity conditions; And through the adaptive feedback optimization closed-loop, the prediction error is calculated according to the mean square error, and hierarchical adjustment is performed according to the size of the error, forming an adaptive feedback closed-loop, which can fine-tune model parameters, adjust data acquisition frequency, and calibrate sensors according to different error paths to ensure continuous optimization of observation and prediction; Through hierarchical response, in the closed-loop feedback mechanism, a hierarchical response strategy is adopted to select an appropriate optimization scheme according to the change of the error, avoiding instability caused by over-adjustment. At the same time, when the error reaches below the minor threshold, the current configuration can be maintained, ensuring the stability and resource efficiency of the observation process; And, under complex meteorological conditions, it can automatically adjust the acquisition and prediction strategies based on real-time feedback, so that the quality of the observed data and the accuracy of the prediction results are continuously optimized in a dynamic environment, providing more accurate support for meteorological prediction.

[0073] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0075] Figure 1Flowchart of the method for controlling the observation quality of precipitation data by integrating multiple machine learning models provided in the embodiments of the present application. Detailed implementation manners

[0076] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0077] Please refer to Figure 1 , Figure 1 Flowchart of the method for controlling the observation quality of precipitation data by integrating multiple machine learning models provided in the embodiments of the present application.

[0078] In this embodiment, the method for controlling the observation quality of precipitation data by integrating multiple machine learning models includes step S100, step S200, step S300, step S400, step S500, step S600, step S700, step S800, and step S900.

[0079] Step S100: Obtain the precipitation data of the pre-observation area, perform observation quality assessment, and determine the abnormal part in the precipitation data based on the quality assessment result.

[0080] Herein, it specifically includes:

[0081] Observation quality assessment: Determine the quality of the precipitation data by calculating the deviation degree of the observed values of the meteorological factors existing in the precipitation data. Specifically:

[0082]

[0083] In the formula, Q i represents the quality assessment value of the i-th meteorological factor, X i represents the observed value of the i-th meteorological factor, μ i represents the historical mean of the i-th meteorological factor, σ i represents the historical standard deviation of the i-th meteorological factor, and i represents the i-th meteorological factor.

[0084] Determine the abnormal part in the precipitation data based on the quality assessment result. By setting the abnormal threshold τ, when Q i exceeds the set abnormal threshold, it is regarded as abnormal data. Specifically:

[0085]

[0086] In the formula, |Q i | represents the absolute quality assessment value of the i-th meteorological factor, τ represents the set abnormal threshold, output 1 indicates that the data is determined to be abnormal, and output 0 indicates that the data is normal.

[0087] It should be noted that in anomaly determination, the deviation value of the observed data is discriminated according to the set anomaly threshold. When the deviation degree exceeds the threshold, it is determined as an abnormal observed value. This method ensures the effectiveness and accuracy of precipitation data, effectively reduces the interference caused by factors such as data noise and observation errors, ensures high-quality data input, and provides a more accurate data basis for prediction analysis.

[0088] Step S200: Use a generative adversarial network to repair the abnormal part in the obtained precipitation data to obtain the repaired precipitation data.

[0089] Here, its repair formula is:

[0090]

[0091] In the formula, represents the precipitation data of the i-th meteorological factor after repair, represents the repair value of the i-th meteorological factor generated by the generator, and λ is the fusion parameter. represents the original observed data of the i-th meteorological factor.

[0092] Among them, 0 < λ < 1, which is used to control the weight distribution between the generated data and the original data.

[0093] To avoid misunderstanding, a brief description of how to repair the abnormal part in the precipitation data through a generative adversarial network is as follows:

[0094] First, identify the abnormal observed values in the precipitation data through quality assessment. The abnormal values are caused by sensor failures, data noise, or acquisition errors.

[0095] The abnormal data is input into the generative adversarial network. The generator model receives the feature information of the observed data and generates a reasonable repair value similar to the characteristics of historical data. The goal of the generator is to generate a "repair value" that conforms to the normal data distribution to replace the abnormal part in the observed data.

[0096] After generating the data, the discriminator model judges the authenticity of the repair data generated by the generator to determine whether it conforms to the characteristics of real data. The discriminator compares the data generated by the generator with the normal observed data and continuously optimizes the generator until the generated data is close enough to the real data.

[0097] When the repair data generated by the generator passes the verification of the discriminator, it is weighted and fused with the original observed data. The fusion parameter is used to adjust the ratio of the generated data to the original data to ensure that the repaired data not only retains the actual characteristics of the observed data but also makes up for the abnormal part of the data.

[0098] The data after being repaired and fused by the generative adversarial network is used to obtain the repaired precipitation data, which is used for subsequent complexity analysis and precipitation prediction.

[0099] It should be noted that the setting of the fusion parameter makes the repair process have a certain degree of flexibility, dynamically adjusting the fusion ratio of the generated data and the original data based on the real-time situation of the meteorological data, so as to optimize the effect of data repair.

[0100] Step S300: Extract the meteorological factors for complexity analysis based on the repaired precipitation data.

[0101] Here, its extraction formula is:

[0102]

[0103] In the formula, represents the correlation between the i-th meteorological factor X i and the precipitation intensity R, and cov(X i , R) represents the covariance between X i and R. and σ R respectively represent the standard deviations of X i and R. Based on the calculation results, extract the meteorological factors for complexity analysis.

[0104] It should be noted that by setting the correlation threshold, the meteorological factors extracted can be dynamically adjusted for different meteorological conditions, ensuring that the complexity analysis is always based on the most representative observation factors, so as to provide accurate data support even when the meteorological conditions become complex.

[0105] Step S400: Conduct complexity analysis based on the extracted meteorological factors and generate a complexity score for the meteorological factors.

[0106] Here, its calculation formula is:

[0107]

[0108] In the formula, C is the complexity score of the meteorological factor, X i represents the current observed value of the i-th meteorological factor, μ i represents the historical mean of the i-th meteorological factor, σ i represents the historical standard deviation of the i-th meteorological factor, w i is the weight of each meteorological factor, i represents the i-th meteorological factor, i traverses all meteorological factors from 1 to n, and each i corresponds to a specific meteorological factor; n represents the total number of samples.

[0109] It should be noted that the weight parameters in the complexity score can be adjusted according to the degree of influence of meteorological factors on precipitation, making the scoring method more flexible and accurate. In different weather scenarios, the weights can be adjusted according to the actual influence of each meteorological factor, realizing the real-time and adaptability of complexity analysis, thus providing a more practical reference for the subsequent model selection and prediction strategy.

[0110] Step S500: Determine the discrimination path of the current precipitation data based on the complexity score.

[0111] Here, it specifically includes:

[0112] The first discrimination path indicates that when the complexity score of the meteorological factors is in the low-complexity interval, in the low-complexity interval, linear regression is selected as the machine learning model according to the first discrimination path, and a prediction result is generated based on the selected linear regression model.

[0113] Furthermore, the low-complexity interval indicates that the current meteorological conditions change stably, with little precipitation and no precipitation. In this interval state, the precipitation weather phenomenon shows linear characteristics, and a linear regression model that can quickly process linear characteristics is used for prediction. The definition of the low-complexity interval is: the complexity score C of the obtained meteorological factors ≤ 30.

[0114] The second discrimination path indicates that when the complexity score of the current meteorological factors is in the medium-complexity interval, in the medium-complexity interval, random forest is selected as the machine learning model according to the second determination path, and a prediction result is generated based on the selected random forest model.

[0115] Furthermore, the medium-complexity interval indicates that there are certain changes in the current meteorological conditions, but they have not reached extreme situations. In this interval state, the precipitation weather phenomenon will show non-linear characteristics, and a random forest model that can handle non-linear relationships is used for prediction processing. The definition of the medium-complexity interval is: 30 < C ≤ 70.

[0116] The third discrimination path indicates that when the complexity score of the current meteorological factors is in the high-complexity interval, in the high-complexity interval, a long short-term memory network is selected as the machine learning model according to the third discrimination path, and a prediction result is generated based on the selected long short-term memory network model.

[0117] Furthermore, the high-complexity interval indicates that the meteorological conditions are very complex, with large and rapidly changing precipitation intensity. In this interval state, while the precipitation intensity shows non-linear characteristics, it also changes rapidly. A long short-term memory network model that can handle complex non-linear relationships and adapt to rapidly changing weather conditions is used for prediction processing. The high-complexity interval is defined as: C > 70.

[0118] Step S600: Dynamically switch to the corresponding machine learning model based on the discrimination path, and predict the precipitation data based on the switched machine learning model to generate a prediction result.

[0119] Here, it specifically includes:

[0120] Set a discrimination interval for the discrimination path, set the complexity score C to vary within the range of [0, 100], and switch between different models based on the score. The discrimination intervals include:

[0121] C ≤ 30 represents the discrimination interval of the first discrimination path; 30 < C ≤ 70 represents the discrimination interval of the second discrimination path; C > 70 represents the discrimination interval of the third discrimination path.

[0122] Define the weight W of each model through the complexity score C i (C), based on the preset meteorological complexity interval, adjust the weights of each model, determine the contribution of each model to the final prediction result, and define the weight W of each model i (C) as a function that changes with the meteorological complexity score C. Adjust the contribution of the model through smooth weight transition. Let W1(C) be the weight of the linear regression model, W2(C) be the weight of the random forest model, and W3(C) be the weight of the long short-term memory network. The switching formula is:

[0123]

[0124] W1(C) represents the weight of the linear regression model, with a weight of 1 when C ≤ 30, gradually decreasing to 0; W2(C) represents the weight of the random forest model, gradually increasing when 30 < C ≤ 70, increasing from 0 to 1, and then gradually decreasing to 0 when C > 70; W3(C) represents the weight of the long short-term memory network, gradually increasing when C > 70, and C represents the complexity score of the meteorological factor.

[0125] It should be noted that the smooth change mechanism of the model weights realizes the seamless transition of each model. Within different intervals of the complexity score, the weight ratio of the model is automatically adjusted, thereby determining the contribution of each model to the final prediction result. The smooth switching design can reduce the instability of the prediction when the precipitation intensity fluctuates greatly, prevent the prediction result from mutating due to model switching, and ensure the coherence and reliability of the prediction process. In addition, the mechanism of dynamically adjusting the weight based on the complexity score of the meteorological factor significantly improves the adaptive ability, especially suitable for precipitation prediction under complex and changeable weather conditions.

[0126] Step S700: Compare the prediction result with the repaired precipitation data and calculate the prediction error.

[0127] Here, the error calculation formula is:

[0128]

[0129] Wherein, MSE represents the mean square error, n represents the total number of samples, represents the prediction error at the t-th moment, represents the precipitation prediction value at the t-th moment, represents the actual precipitation at the t-th moment.

[0130] It should be noted that the application of the mean square error judges the reliability of the current model prediction according to the error size, and flexibly adjusts the model configuration based on the feedback optimization mechanism. When the error is small, the existing configuration is maintained, and when the error exceeds the preset range, the model parameters can be dynamically adjusted or the model type can be replaced, so as to improve the self-adaptability and prediction accuracy of the method.

[0131] Step S800: Generate a feedback optimization mechanism based on the obtained prediction error, and generate an optimization plan according to the feedback optimization mechanism to optimize the future precipitation observation and prediction process.

[0132] Here, it specifically includes:

[0133] Through calculate the error E, set the error threshold ∈ based on the error E, and judge the path to which the current obtained error belongs based on the error threshold ∈. Specifically:

[0134]

[0135] Wherein, ∈1 represents the slight error threshold, ∈2 represents the significant error threshold, E represents the error, Path 1, when E < ∈1, the error is below the slight error threshold and no adjustment is made according to the belonging path; Path 2, when ∈1 ≤ E < ∈2, the error is above the slight threshold but has not reached the significant error threshold, and a slight adjustment is performed. The slight adjustment is to finely adjust the parameters of the machine learning model participating in the prediction; Path 3, when E ≥ ∈2, trigger significant adjustment measures, and the significant adjustment includes increasing the data collection frequency and sensor calibration.

[0136] It should be noted that the hierarchical setting of the error threshold ensures accurate adaptive adjustment based on the error change situation. When the error exceeds the slight threshold but has not reached the significant threshold, effective compensation can be achieved through fine adjustment of the model parameters; when the error reaches the significant threshold, adjustments such as frequency increase and sensor calibration are automatically performed to make the quality and collection frequency of the data source reach a higher standard. This optimization plan ensures stability and adaptability under complex meteorological conditions through a hierarchical response mechanism, and effectively improves the accuracy of precipitation observation and prediction.

[0137] Step S900: Based on the error discrimination path, perform error discrimination optimization and adjustment, apply the optimization and adjustment to the next round of observations. After the next round of observations is completed, calculate the error again to form a closed loop.

[0138] Specifically here:

[0139] Automatically adjust according to the adjustment strategy selected by the error path, and apply the adjusted acquisition frequency, model configuration, and sensor calibration information to the next round of observations. If the error decreases and remains below the minor error threshold, it proves that the current optimization is effective, and then continue to use the current acquisition frequency and model configuration based on the effective optimization. If the error is still greater than the significant error threshold, re-evaluate the error source until the error is adjusted below the minor error threshold.

[0140] It should be noted that through multiple rounds of feedback evaluation and automatic adjustment, this closed-loop mechanism significantly improves the adaptability and self-correction ability under complex meteorological conditions. When the error does not drop within the target range, it can automatically re-evaluate the error source and continuously make fine adjustments until it reaches below the minor error threshold, thus ensuring the long-term accuracy and reliability of precipitation observations and providing high-quality data support for weather forecasting.

[0141] When implementing this method, obtain precipitation data in the pre-observation area, conduct observation quality assessment to obtain the abnormal part, and determine the abnormal part in the precipitation data based on the quality assessment results; use a generative adversarial network to repair the abnormal part in the obtained precipitation data to obtain the repaired precipitation data; based on the repaired precipitation data, extract meteorological factors for complexity analysis; conduct complexity analysis based on the extracted meteorological factors and generate a complexity score for the meteorological factors; based on the complexity score, determine the discrimination path of the current precipitation data; dynamically switch to the corresponding machine learning model based on the discrimination path, and predict the precipitation data based on the switched machine learning model to generate a prediction result; compare the prediction result with the repaired precipitation data to calculate the prediction error; generate a feedback optimization mechanism based on the obtained prediction error, and generate an optimization plan according to the feedback optimization mechanism to optimize the future precipitation observation and prediction process; based on the error discrimination path, perform error discrimination optimization and adjustment, apply the optimization and adjustment to the next round of observations. After the next round of observations is completed, calculate the error again to form a closed loop. This method effectively improves the quality control level of precipitation data observation through a series of steps such as meteorological factor extraction, generative adversarial network repair, complexity analysis, and dynamic switching of the discrimination path, and can achieve high-precision observation and real-time adjustment of precipitation data under different meteorological conditions, providing reliable data support for weather forecasting.

[0142] For easy understanding, an implementation process of changing from a heavy rain state to a moderate rain state and then from a moderate rain state to a light rain state will be presented. Specifically:

[0143] Initial observation data collection: Real-time collection of precipitation and meteorological data in the pre-observation area. The observed data includes meteorological factors such as precipitation intensity, temperature, air pressure, humidity, and wind speed. Due to the characteristics of heavy rain, the precipitation intensity and wind speed change significantly during the initial observation stage.

[0144] Conduct quality assessment on the collected data, calculate the degree to which the observed values of each meteorological factor deviate from the historical mean. For data vulnerable to environmental interference or acquisition errors, use a generative adversarial network to repair the outliers, and obtain the repaired precipitation data, providing accurate data input for subsequent complexity analysis.

[0145] Extract key meteorological factors based on the repaired precipitation data and conduct complexity analysis. The complexity score is calculated according to the volatility and non-linearity of the current observed data. Due to the significant meteorological fluctuations during the heavy rain stage, the initial state of complexity is in the high-complexity range.

[0146] In the high-complexity range where C > 70, allocate the main weights to the long short-term memory network to capture the non-linear changes in data under heavy rain conditions. Generate precipitation prediction results based on LSTM, monitor the fluctuations of precipitation intensity in real time, and continuously update the model weights according to the actual observed data.

[0147] As the heavy rain gradually weakens to moderate rain, the precipitation intensity and meteorological factor fluctuations tend to be stable, and the complexity score drops to the medium range. Dynamically switch the model weights, reduce the weights of LSTM, and gradually increase the weights of the random forest model to adapt to the relatively stable changes in meteorological conditions.

[0148] Under moderate rain conditions, rely on the random forest model to predict future precipitation. This model can capture the non-linear characteristics during moderate rain, but has better adaptability to data fluctuations, helping to improve the prediction efficiency when the complexity decreases.

[0149] When the precipitation gradually weakens to light rain, the complexity score drops to the low-complexity range. At this time, allocate the main weights to the linear regression model to reduce the computational amount and maintain the prediction accuracy. Switch to the linear regression model, with the model weights mainly based on linear regression. Since the precipitation intensity changes little during light rain, linear regression can achieve efficient prediction under low data fluctuations.

[0150] Compare the prediction results of each stage with the repaired actual observed data, calculate the mean squared error, and determine whether the error exceeds the set error threshold.

[0151] If the error value is below the minor error threshold, maintain the current model configuration and acquisition frequency; if the error is on the significant error path, trigger significant adjustments, including increasing the acquisition frequency, recalibrating the sensors, and readjusting the model parameters again until the error drops below the minor error threshold.

[0152] Based on the optimized acquisition frequency and model parameters, enter the next round of observations. After each round of observations, repeat the above process to form a feedback closed loop, and continuously optimize the prediction accuracy and data quality control during the gradual change of precipitation conditions.

[0153] Thus, this embodiment is completed.

[0154] In summary, in the embodiment of the present invention, the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A control method for the observation quality of precipitation data that integrates multiple machine learning models, characterized in that, It includes the following: Obtain precipitation data in the pre-observation area, conduct observation quality assessment, and determine the abnormal part in the precipitation data based on the quality assessment result; Use a generative adversarial network to repair the abnormal part in the obtained precipitation data to obtain the repaired precipitation data; Based on the repaired precipitation data, extract meteorological factors for complexity analysis; Conduct complexity analysis based on the extracted meteorological factors and generate a complexity score for the meteorological factors; Based on the complexity score, determine the discrimination path of the current precipitation data, where the discrimination path includes the first discrimination path, the second discrimination path, and the third discrimination path; The first discrimination path: The complexity score of the meteorological factor is in the low-complexity interval; The second discrimination path: The complexity score of the meteorological factor is in the medium-complexity interval; The third discrimination path: The complexity score of the meteorological factor is in the high-complexity interval; Dynamically switch to the corresponding machine learning model based on the discrimination path, and predict the precipitation data based on the switched machine learning model to generate a prediction result; Compare the prediction result with the repaired precipitation data to calculate the prediction error; Generate a feedback optimization mechanism based on the obtained prediction error, and generate an optimization plan according to the feedback optimization mechanism to optimize the future precipitation observation and prediction process; Based on the error discrimination path, perform error discrimination optimization adjustment, apply the optimization adjustment to the next round of observation, and calculate the error again after the next round of observation is completed to form a closed loop.

2. The method for controlling the observation quality of precipitation data by integrating multiple machine learning models according to claim 1, characterized in that Obtain precipitation data in the pre-observation area, conduct observation quality assessment, and determine the abnormal part in the precipitation data based on the quality assessment result. Specifically, it includes: Observation quality assessment: Determine the quality of the precipitation data by calculating the deviation degree of the observed values of the meteorological factors existing in the precipitation data. Specifically: Where Q i represents the quality evaluation value of the i-th meteorological factor, X i represents the observed value of the i-th meteorological factor, μ i represents the historical mean of the i-th meteorological factor, σ i represents the historical standard deviation of the i-th meteorological factor, and i represents the i-th meteorological factor; Determine the abnormal part in precipitation data based on the quality assessment results. By setting an abnormal threshold τ, when Q i exceeds the set abnormal threshold, it is regarded as abnormal data. Specifically: where, |Q i | represents the absolute quality assessment value of the i-th meteorological factor, τ represents the set anomaly threshold, an output of 1 indicates that the data is determined to be abnormal, and an output of 0 indicates that the data is normal.

3. The control method for the observation quality of precipitation data integrating multiple machine learning models according to claim 2, characterized in that, Use a generative adversarial network to repair the abnormal part in the obtained precipitation data to obtain the repaired precipitation data. Its repair formula is: Wherein, represents the precipitation data of the i-th meteorological factor after repair, represents the repaired value of the i-th meteorological factor generated by the generator, and λ is the fusion parameter, represents the original observation data of the i-th meteorological factor.

4. The method for controlling the observation quality of precipitation data by integrating multiple machine learning models according to claim 1, characterized in that, Based on the repaired precipitation data, extract meteorological factors for complexity analysis. Its extraction formula is: In the formula, represents the correlation between the i-th meteorological factor X i and the precipitation intensity R, and cov(X i , R) represents the covariance between X i and R. and σ R represent the standard deviations of X i and R respectively. Based on the calculation results, meteorological factors for complexity analysis are extracted.

5. The control method for the observation quality of precipitation data integrating multiple machine learning models according to claim 1, wherein Conduct complexity analysis based on the extracted meteorological factors and generate a complexity score for the meteorological factors. Its calculation formula is: Where C is the complexity score of meteorological factors, and X i represents the current observed value of the i-th meteorological factor, and μ i represents the historical mean of the i-th meteorological factor, and σ i represents the historical standard deviation of the i-th meteorological factor, and w i is the weight of each meteorological factor. i represents the i-th meteorological factor, and i traverses all meteorological factors from 1 to n. Each i corresponds to a specific meteorological factor; n represents the total number of samples.

6. The control method for the observation quality of precipitation data integrating multiple machine learning models according to claim 5, characterized in that, Based on the complexity score, determine the discrimination path of the current precipitation data. Specifically, it includes: The first discrimination path indicates that when the complexity score of the meteorological factor is in the low-complexity interval, if it is in the low-complexity interval, select linear regression as the machine learning model according to the first discrimination path, and generate a prediction result based on the selected linear regression model; The second discrimination path indicates that the current complexity score of the meteorological factor is in the medium-complexity interval. If it is in the medium-complexity interval, select random forest as the machine learning model according to the second determination path, and generate a prediction result based on the selected random forest model; The third discrimination path indicates that the current complexity score of the meteorological factor is in the high-complexity interval. If it is in the high-complexity interval, select a long short-term memory network as the machine learning model according to the third discrimination path, and generate a prediction result based on the selected long short-term memory network model.

7. The control method for the observation quality of precipitation data integrating multiple machine learning models according to claim 6, characterized in that, Dynamically switch to the corresponding machine learning model based on the discrimination path, and predict the precipitation data based on the switched machine learning model to generate a prediction result. Specifically, it includes: Set the discrimination interval for the discrimination path, set the complexity score C to vary in the range of [0, 100], and switch between different models based on the score. The discrimination interval includes: C≤30 represents the discrimination interval of the first discrimination path; 30<C≤70表示第二判别路径的判别区间;C> 70 represents the discrimination interval of the third discrimination path; Define the weight W of each model through the complexity score C i (C), based on a preset meteorological complexity interval, adjust the weights of each model to determine the contribution of each model to the final prediction result, and define the weight W of each model i (C) is a function that changes with the meteorological complexity score C. Adjust the contribution of the model through smooth transition of weights. Let W1(C) be the weight of the linear regression model, W2(C) be the weight of the random forest model, and W3(C) be the weight of the long short-term memory network. The switching formula is: W1(C) represents the weight of the linear regression model. When C≤30, the weight is 1 and gradually decreases to 0. W2(C) represents the weight of the random forest model. When C≤30, the weight is 1 and gradually decreases to 0.<C≤70时逐渐增加,从0增加到1,然后在C> When C>70, it gradually decreases to 0; W3(C) represents the weight of the long short-term memory network, and gradually increases when C>70, and C represents the complexity score of the meteorological factor.

8. The control method for the observation quality of precipitation data integrating multiple machine learning models according to claim 1, characterized in that, Based on the comparison between the prediction results and the repaired precipitation data, the prediction error is calculated. The error calculation formula is: Where MSE represents the mean square error, and n represents the total number of samples. represents the prediction error at the t-th moment. represents the precipitation prediction value at the t-th moment. represents the actual precipitation at the t-th moment.

9. The control method for the observation quality of precipitation data integrating multiple machine learning models according to claim 1, characterized in that, Based on the obtained prediction error, a feedback optimization mechanism is generated, and an optimization scheme is generated according to the feedback optimization mechanism to optimize the future precipitation observation and prediction process, including: By calculate the error E, set the error threshold ∈ based on the error E, and determine the path to which the currently obtained error belongs based on the error threshold ∈, specifically: Wherein, ∈1 represents the slight error threshold, ∈2 represents the significant error threshold, and E represents the error. Path 1: when E<∈1, the error is below the slight error threshold, and no adjustment is performed according to the path to which it belongs. Path 2: when ∈1≤E<∈2, the error is above the slight threshold but has not reached the significant error threshold, and a slight adjustment is performed. A slight adjustment is to fine-tune the parameters of the machine learning model involved in the prediction. Path 3: when E≥∈2, significant adjustment measures are triggered, and significant adjustments include increasing the data collection frequency and sensor calibration.

10. The control method for the observation quality of precipitation data integrating multiple machine learning models according to claim 1, characterized in that, Based on the error discrimination path, perform error discrimination optimization and adjustment, apply the optimization and adjustment to the next round of observation, and after the next round of observation is completed, calculate the error again to form a closed loop, specifically: The adjustment strategy selected according to the error path is automatically adjusted, and the adjusted acquisition frequency, model configuration and sensor calibration information are applied to the next round of observations. If the error is reduced and maintained below the slight error threshold, it proves that the current optimization is effective. Based on the effective optimization, the current acquisition frequency and model configuration continue to be used. If the error is still greater than the significant error threshold, the error source is re-evaluated until the error is adjusted below the slight error threshold.

Citation Information

Patent Citations

  • Multi-mode rainfall estimation method integrated with machine learning

    CN113361766A

  • New energy power generation prediction method based on machine learning and meteorological similar moments

    CN116562115A