Meteorological missing data set correction method and system based on correlation analysis
By collecting and standardizing meteorological and hydrological data from distributed data sources, identifying missing data, and using variable relationship models for completion calculations, the problem of lacking multi-dimensional correlation information in existing technologies has been solved. This has enabled high-precision and physically consistent data correction, improving the integrity and reliability of meteorological and hydrological data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUANENG XINJIANG ENERGY DEVELOPMENT CO LTD SOUTHERN XINJIANG CLEAN ENERGY BRANCH
- Filing Date
- 2025-11-20
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies lack a systematic integration of temporal, spatial, and physical variables when processing meteorological and hydrological data. This results in insufficient reliability and physical consistency of data missing correction results, making it difficult to meet the needs of high-precision meteorological and hydrological simulation and forecasting.
Meteorological and hydrological data are collected from distributed data sources, cleaned and standardized, missing data are identified and their physical consistency is verified, a variable relationship model trained on historical data is used for completion calculation, the results are verified and the model is updated, and multi-dimensional correction of missing data is carried out by combining time, space and historical similarity information.
It achieves high-precision, physically consistent, and adaptively optimized correction of missing meteorological and hydrological data from multiple sources, improving the integrity and reliability of the data and enhancing the practical value of meteorological forecasting and hydrological analysis.
Smart Images

Figure CN121935490A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of meteorological and hydrological data processing technology, and in particular to a method and system for correcting missing meteorological datasets based on correlation analysis. Background Technology
[0002] Meteorological and hydrological data are crucial for weather forecasting, climate analysis, and water resource management; their completeness and accuracy directly impact the reliability of related analyses and decisions. Current technologies typically acquire this data from various heterogeneous data sources, such as widely distributed automatic weather stations, remote sensing satellites, and numerical weather prediction centers, and then assimilate and fuse these data for further applications. However, these data from different platforms often differ in temporal resolution and spatial scale. Furthermore, during actual acquisition and transmission, data loss or numerical anomalies are inevitable due to equipment failures, communication interruptions, or environmental interference. Current mainstream processing methods tend to focus on a single dimension. For example, they may rely solely on autoregressive models of time series for interpolation or use only geostatistical interpolation of spatially nearby points. While such methods may achieve results in specific scenarios, they generally fail to systematically integrate the multi-dimensional correlation information between time, space, and physical variables of the data. They also lack effective verification of whether the interpolation results conform to atmospheric, hydrological, and other physical laws. As a result, when correcting data gaps under complex underlying surfaces or extreme weather processes, the reliability and physical consistency of the results are often insufficient, making it difficult to meet the application requirements of high-precision meteorological and hydrological simulation and forecasting. Summary of the Invention
[0003] In view of this, the present invention provides a method for correcting missing meteorological datasets based on correlation analysis. One or more embodiments of this specification also relate to a system for correcting missing meteorological datasets based on correlation analysis, a computing device, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.
[0004] According to a first aspect of the present invention, a method for correcting missing meteorological datasets based on correlation analysis is provided, comprising: Numerical weather forecast data, real-time meteorological data, and watershed snow cover data are collected from distributed data sources, and the collected data are cleaned and subjected to temporal and spatial standardization to obtain standardized data. The standardized data is subjected to missing data identification and physical law conformity verification to obtain the missing dataset; The missing dataset is filled in using the trained variable relationship model to generate the filled data. The variable relationship model is built and trained based on historical data. The filled data is then validated to generate validation results, and the variable relationship model is updated based on the validation results.
[0005] In some implementations, numerical weather forecast data, real-time meteorological data, and watershed snow cover data are collected from distributed data sources, and the collected data are cleaned and subjected to temporal and spatial standardization processes, including: Forecast data is obtained from the numerical weather prediction center, real-time data is collected from automatic weather stations, and snow cover monitoring data is obtained from the remote sensing platform; The collected data is filtered to obtain valid data; Perform time alignment on valid data to give all data a uniform timestamp; Spatial resampling is performed on the time-aligned data to give all data a uniform spatial resolution, resulting in standardized data.
[0006] In some implementations, missing data is identified and physical conformity checks are performed on standardized data to obtain a missing dataset, including: By scanning the data through a sliding window, continuous missing segments and random missing points are identified to obtain the missing location information; The dynamic boundary model is used to verify the physical laws of the data values. The dynamic boundary model obtains the preset range of variable values based on historical data statistics. When a data value exceeds a preset range, it is marked as physically unreasonable data and included in the missing dataset. By combining missing location information and physically invalid data, a complete missing dataset is generated.
[0007] In some implementations, the training steps for the variable relationship model include: Collect historical meteorological and hydrological data, including numerical weather prediction data, real-time meteorological data, and watershed snow cover data; Long Short-Term Memory Network is used to extract time series features, and Transformer encoder is used to capture long-term dependencies between variables. By combining causal discovery algorithms to analyze causal relationships between variables, a causal relationship network is constructed. The trained variable relationship model is obtained by training a pre-set initial variable relationship model based on a causal relationship network.
[0008] In some implementations, the steps of performing completion calculations on the missing dataset based on a variable relationship model to generate completed data specifically include: For a randomly missing dataset, complete values for meteorological and hydrological data are generated based on the following first calculation formula: in, This indicates the completed meteorological and hydrological data, and is a non-negative value. This represents the time-related weighting coefficient, which is derived from the reliability assessment of each sub-item output by the dynamic coupling correlation analysis module and is used to adjust the contribution of the time-related item. The dynamic weight coefficient of the i-th related variable is derived from the training results of the dynamic coupling correlation analysis module and is used to quantify the degree of influence of different related variables on the target variable. This represents the dimensional transformation value of the i-th related variable, obtained through a linear transformation function. Obtain, among which These are conversion coefficients derived from least squares regression analysis of historical data, ensuring that the converted dimensions match the target variable. Consistent; The value of the i-th related variable is derived from the multi-source heterogeneous data acquisition and preprocessing module and serves as the basic input data for the completion calculation. This represents the time difference between the observation time and the completion time of the i-th related variable, derived from time series analysis of the data, and is used to introduce the time decay effect; It represents the time decay constant, which is derived from historical data statistical analysis and controls the decay rate of the time effect; This represents the first numerical stability factor, set to a very small positive real number to prevent abnormal calculations of numbers with a denominator of zero; This represents the spatial correlation weighting coefficient, derived from the accuracy evaluation results of the spatial interpolation algorithm, and is used to adjust the contribution of the spatial correlation term. The distance weight of the j-th spatial neighbor point is derived from the spatial interpolation algorithm and reflects the degree of influence of spatial location on the target point. The data quality index represents the j-th spatial neighbor point. Its value is between 0 and 1. It is calculated by the missing data intelligent completion module based on the data quality markers output by the data integrity and physical rationality verification module. This represents the observed value of the j-th spatially neighboring point, distinguished from the target variable by a wavy line symbol. It is derived from real-world data collection and provides spatial correlation information. This represents the second numerical stability factor, set to a very small positive real number to ensure that the denominator is always greater than zero; This represents the historical similarity weight coefficient, derived from the confidence assessment of the pattern matching algorithm, and is used to adjust the contribution of the historical similarity item. , and The sum of is 1; The matching weight of the kth historical similar scenario is derived from the pattern matching algorithm and quantifies the similarity between the historical scenario and the current situation. The target variable value representing the kth historically similar scenario is derived from a historical database, providing a reference for historical patterns. The third numerical stability factor is set to a very small positive real number to maintain computational stability; N represents the total number of relevant variables, M represents the total number of spatially neighboring points, and P represents the total number of historically similar scenarios, all of which are positive integers. The random missing data processing unit, when performing calculations, uses a preset quality threshold. Filter nearby points in space; when identified < It automatically excludes nearby points in the space to ensure that the denominators of all fractions in the calculation are greater than zero.
[0009] In some implementations, the dynamic weighting coefficients are determined by the following formula: in, This represents the scaling parameter, which originates from hyperparameter optimization during model training and is used to balance the magnitude differences in correlation coefficients. The target variable is... Related to the i-th variable The correlation coefficient is derived from the Pearson correlation coefficient statistical analysis of historical data. It is used to quantify the strength of the linear association between variables. By taking the absolute value, it ensures that the weight calculation only considers the correlation strength and is not affected by the correlation direction; N represents the total number of related variables and is a positive integer.
[0010] In some implementations, the steps of validating the completed data, generating a validation result, and updating the variable relationship model based on the validation result specifically include: The completed data is compared with the independent observation data, the error index is calculated, and the verification result is generated. When the error index exceeds the preset threshold, the completed data is marked as data that needs to be corrected, and the user's correction input is received through the visual interface; The validated data is added to the training set, and the parameters of the variable relationship model are updated through incremental learning. Regularly generate data quality assessment reports to summarize the completion effects and model performance.
[0011] In some implementations, the method further includes: When the amount of missing data exceeds a preset threshold, an alternative completion strategy is activated. The alternative completion strategy includes using one or more combinations of time series forecasting, spatial interpolation, or historical mean methods to complete the data and generate alternative completed data.
[0012] According to a second aspect of the present invention, a meteorological missing dataset correction system based on correlation analysis is provided, comprising: The standardization processing module is used to collect numerical weather forecast data, real-time meteorological data, and watershed snow cover data from distributed data sources, and to clean and perform temporal and spatial standardization processing on the collected data to obtain standardized data. The verification module is used to identify missing data and verify compliance with physical laws in standardized data to obtain the missing dataset. The calculation module is used to perform completion calculations on the missing dataset based on the trained variable relationship model, and generate the completed data. The variable relationship model is built and trained based on historical data. After verifying the completed data, the module generates a verification result and updates the variable relationship model based on the verification result.
[0013] In some implementations, numerical weather forecast data, real-time meteorological data, and watershed snow cover data are collected from distributed data sources, and the collected data are cleaned and subjected to temporal and spatial standardization processes, including: Forecast data is obtained from the numerical weather prediction center, real-time data is collected from automatic weather stations, and snow cover monitoring data is obtained from the remote sensing platform; The collected data is filtered to obtain valid data; Perform time alignment on valid data to give all data a uniform timestamp; Spatial resampling is performed on the time-aligned data to give all data a uniform spatial resolution, resulting in standardized data.
[0014] In some implementations, missing data is identified and physical conformity checks are performed on standardized data to obtain a missing dataset, including: By scanning the data through a sliding window, continuous missing segments and random missing points are identified to obtain the missing location information; The dynamic boundary model is used to verify the physical laws of the data values. The dynamic boundary model obtains the preset range of variable values based on historical data statistics. When a data value exceeds a preset range, it is marked as physically unreasonable data and included in the missing dataset. By combining missing location information and physically invalid data, a complete missing dataset is generated.
[0015] In some implementations, the training steps for the variable relationship model include: Collect historical meteorological and hydrological data, including numerical weather prediction data, real-time meteorological data, and watershed snow cover data; Long Short-Term Memory Network is used to extract time series features, and Transformer encoder is used to capture long-term dependencies between variables. By combining causal discovery algorithms to analyze causal relationships between variables, a causal relationship network is constructed. The trained variable relationship model is obtained by training a pre-set initial variable relationship model based on a causal relationship network.
[0016] In some implementations, the steps of performing completion calculations on the missing dataset based on a variable relationship model to generate completed data specifically include: For a randomly missing dataset, complete values for meteorological and hydrological data are generated based on the following first calculation formula: in, This indicates the completed meteorological and hydrological data, and is a non-negative value. This represents the time-related weighting coefficient, which is derived from the reliability assessment of each sub-item output by the dynamic coupling correlation analysis module and is used to adjust the contribution of the time-related item. The dynamic weight coefficient of the i-th related variable is derived from the training results of the dynamic coupling correlation analysis module and is used to quantify the degree of influence of different related variables on the target variable. This represents the dimensional transformation value of the i-th related variable, obtained through a linear transformation function. Obtain, among which These are conversion coefficients derived from least squares regression analysis of historical data, ensuring that the converted dimensions match the target variable. Consistent; The value of the i-th related variable is derived from the multi-source heterogeneous data acquisition and preprocessing module and serves as the basic input data for the completion calculation. This represents the time difference between the observation time and the completion time of the i-th related variable, derived from time series analysis of the data, and is used to introduce the time decay effect; It represents the time decay constant, which is derived from historical data statistical analysis and controls the decay rate of the time effect; This represents the first numerical stability factor, set to a very small positive real number to prevent abnormal calculations of numbers with a denominator of zero; This represents the spatial correlation weighting coefficient, derived from the accuracy evaluation results of the spatial interpolation algorithm, and is used to adjust the contribution of the spatial correlation term. The distance weight of the j-th spatial neighbor point is derived from the spatial interpolation algorithm and reflects the degree of influence of spatial location on the target point. The data quality index represents the j-th spatial neighbor point. Its value is between 0 and 1. It is calculated by the missing data intelligent completion module based on the data quality markers output by the data integrity and physical rationality verification module. This represents the observed value of the j-th spatially neighboring point, distinguished from the target variable by a wavy line symbol. It is derived from real-world data collection and provides spatial correlation information. This represents the second numerical stability factor, set to a very small positive real number to ensure that the denominator is always greater than zero; This represents the historical similarity weight coefficient, derived from the confidence assessment of the pattern matching algorithm, and is used to adjust the contribution of the historical similarity item. , and The sum of is 1; The matching weight of the kth historical similar scenario is derived from the pattern matching algorithm and quantifies the similarity between the historical scenario and the current situation. The target variable value representing the kth historically similar scenario is derived from a historical database, providing a reference for historical patterns. The third numerical stability factor is set to a very small positive real number to maintain computational stability; N represents the total number of relevant variables, M represents the total number of spatially neighboring points, and P represents the total number of historically similar scenarios, all of which are positive integers. The random missing data processing unit, when performing calculations, uses a preset quality threshold. Filter nearby points in space; when identified < It automatically excludes nearby points in the space to ensure that the denominators of all fractions in the calculation are greater than zero.
[0017] In some implementations, the dynamic weighting coefficients are determined by the following formula: in, This represents the scaling parameter, which originates from hyperparameter optimization during model training and is used to balance the magnitude differences in correlation coefficients. The target variable is... Related to the i-th variable The correlation coefficient is derived from the Pearson correlation coefficient statistical analysis of historical data. It is used to quantify the strength of the linear association between variables. By taking the absolute value, it ensures that the weight calculation only considers the correlation strength and is not affected by the correlation direction; N represents the total number of related variables and is a positive integer.
[0018] In some implementations, the steps of validating the completed data, generating a validation result, and updating the variable relationship model based on the validation result specifically include: The completed data is compared with the independent observation data, the error index is calculated, and the verification result is generated. When the error index exceeds the preset threshold, the completed data is marked as data that needs to be corrected, and the user's correction input is received through the visual interface; The validated data is added to the training set, and the parameters of the variable relationship model are updated through incremental learning. Regularly generate data quality assessment reports to summarize the completion effects and model performance.
[0019] In some implementations, the system further includes: The backup completion module is used to activate a backup completion strategy when the amount of missing data exceeds a preset threshold. The backup completion strategy includes using one or more combinations of time series forecasting, spatial interpolation, or historical mean methods to complete the data and generate alternative completion data.
[0020] According to a third aspect of the present invention, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the above-described method for correcting missing meteorological datasets based on correlation analysis.
[0021] According to a fourth aspect of the present invention, a computer-readable storage medium is provided storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described method for correcting missing meteorological datasets based on correlation analysis.
[0022] According to a fifth aspect of the present invention, a computer program is provided, wherein when the computer program is executed in a computer, the computer is instructed to perform the steps of the above-described method for correcting missing meteorological datasets based on correlation analysis.
[0023] At least one embodiment of this invention collects and standardizes multiple types of meteorological and hydrological data from a distributed data source system, constructs a missing dataset through missing data identification and physical law conformity verification, and then uses a variable relationship model trained based on historical data and incorporating spatiotemporal features and causal relationships to perform completion calculations. The completion results are then verified and the model is updated, achieving high-precision, physically consistent, and adaptively optimized correction of missing meteorological and hydrological data from multiple sources. This effectively improves the integrity and reliability of the data and its practical value in meteorological forecasting and hydrological analysis. Attached Figure Description
[0024] Figure 1 This is a flowchart of a meteorological missing dataset correction method based on correlation analysis provided by the present invention; Figure 2 This is a simplified structural diagram of a meteorological missing dataset correction system based on correlation analysis provided by the present invention; Figure 3 This is a structural block diagram of a computing device provided by the present invention. Detailed Implementation
[0025] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0026] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the one or more embodiments of this specification. The singular forms “a” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items. The modifications “a” and “a plurality” as used in this disclosure are illustrative and not restrictive, and those skilled in the art will understand that they should be understood as “one or more” unless the context clearly indicates otherwise.
[0027] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0028] See Figure 1 , Figure 1 The flowchart illustrates a method for correcting missing meteorological datasets based on correlation analysis, according to some embodiments of this specification, specifically including the following steps: Numerical weather forecast data, real-time meteorological data, and watershed snow cover data are collected from distributed data sources. The collected data are then cleaned and subjected to temporal and spatial standardization to obtain standardized data.
[0029] The missing dataset is obtained by identifying missing data and verifying its conformity to physical laws.
[0030] The missing dataset is filled in using the trained variable relationship model to generate the filled data. The variable relationship model is built and trained based on historical data. The filled data is then validated to generate validation results, and the variable relationship model is updated based on the validation results.
[0031] Distributed data sources refer to systems or platforms that store data in a physically or logically dispersed manner. Examples include data acquired from numerical weather prediction centers, automatic weather stations, and remote sensing satellite platforms via API interfaces, providing multi-dimensional, multi-scale raw meteorological and hydrological information. Numerical weather prediction data refers to future weather state predictions obtained through computer simulations based on atmospheric dynamics and physics equations. Examples include gridded data from the Global Forecast System (GFS) or the European Centre for Medium-Range Weather Forecasts (ECMWF), providing high-resolution forecast fields as a reference benchmark for data completion. Real-time meteorological data refers to current or past meteorological element values actually measured by observation equipment, such as temperature, humidity, and wind speed collected from automatic weather stations, used to verify and calibrate other data sources. Watershed snow cover data refers to parameters such as snow cover extent, snow depth, or snow water equivalent within a specific river basin, obtained, for example, through inversion from MODIS or Landsat remote sensing images, reflecting the watershed's hydrological conditions. The collected data can refer to raw meteorological and hydrological observations or simulations obtained from various data sources, such as data downloaded via RESTful APIs or FTP protocols through the data acquisition module, serving as input for subsequent processing. Cleaning and spatiotemporal standardization refer to operations performed on the raw data, including filtering outliers, standardizing formats, and aligning coordinates. For example, using a sliding window to remove outliers and bilinear interpolation for spatial resampling to eliminate data heterogeneity. Standardized data refers to a dataset with a consistent spatiotemporal reference after cleaning and formatting, such as unifying all data to the WGS84 coordinate system and UTC timestamps for easier subsequent analysis. Missing data identification and physical law compliance verification refer to the process of detecting data gaps and verifying whether they conform to natural laws. For example, using statistical boundary models to determine whether temperature values are within a preset range to identify data points requiring correction. Missing datasets refer to subsets of data marked as missing or physically unreasonable, such as records containing consecutive missing segments and outliers, serving as targets for completion calculations. Trained variable relationship models refer to machine learning models that learn dependencies between variables based on historical data, such as using LSTM and Transformer networks to capture spatiotemporal correlations for estimating missing values. Completion computation can refer to the process of filling in data gaps using a pre-defined algorithm or model. For example, it might involve applying a weighted formula to fuse temporal, spatial, and historical similarity information to generate appropriate completion values. Completed data can refer to the corrected and processed dataset, such as storing the completed variable values in a database for meteorological analysis or forecasting. Variable relationship models can refer to computational frameworks that describe the mathematical relationships between meteorological and hydrological variables. For example, it might involve constructing network structures using causal discovery algorithms to characterize causal relationships between variables. Historical data can refer to accumulated meteorological and hydrological observations or simulation records, such as collecting years' worth of temperature, precipitation, and snow cover data as the basis for model training.Validation can refer to the verification step of comparing the completed results with independent observations, such as calculating the root mean square error index to evaluate the quality of the completion. Validation results can refer to error reports or quality scores generated during the validation process, such as outputting confidence indices to guide model updates.
[0032] As a concrete example: In a meteorological data completion scenario for a watershed in the Qinghai-Tibet Plateau, numerical weather prediction data in GRIB format was obtained from the China Meteorological Administration's Numerical Weather Prediction Center via HTTP API, real-time temperature and precipitation data in CSV format were collected from the network of automatic weather stations in the Tibet Autonomous Region, and MODIS snow cover products in HDF5 format were downloaded from the NASA Earthdata platform. During the data cleaning phase, the 3σ rule was used to filter out outliers. Temporal standardization aligned all data to UTC time, and spatial standardization used nearest-neighbor interpolation to unify data from different resolutions into a 1km grid. Missing data identification used a sliding window to detect gaps exceeding 3 hours, and physical verification was based on historical statistical dynamic boundaries (e.g., summer temperature range of -10℃ to 30℃). The variable relationship model adopted a hybrid LSTM-Transformer architecture implemented in PyTorch, trained using data from the past 5 years. Inputs included temperature, humidity, wind speed, and snow depth, and the output was the completed values. During completion calculation, the time weight was set to 0.4, the spatial weight to 0.3, the historical similarity weight to 0.3, and the decay constant τ = 6 hours. During verification, the data is compared with that of an independently set base station. If the mean absolute error exceeds 2°C, manual correction is received via the web interface, and incremental learning of the model is triggered.
[0033] By collecting and standardizing various types of meteorological and hydrological data from a distributed data source system, a missing dataset is constructed through missing data identification and physical law conformity verification. Then, a variable relationship model based on historical data and incorporating spatiotemporal characteristics and causal relationships is used for completion calculation. The completion results are verified and the model is updated, achieving high-precision, physically consistent, and adaptively optimized correction of missing meteorological and hydrological data from multiple sources. This effectively improves the integrity and reliability of the data and its practical value in meteorological forecasting and hydrological analysis.
[0034] The beneficial effects of one of the embodiments in this specification include at least the following: by collecting and standardizing multiple types of meteorological and hydrological data from a distributed data source system, constructing a missing dataset through missing data identification and physical law conformity verification, and then using a variable relationship model trained based on historical data and incorporating spatiotemporal characteristics and causal relationships for completion calculation, and verifying and updating the completion results, high-precision, physically consistent and adaptively optimized correction of missing meteorological and hydrological data from multiple sources is achieved, effectively improving the integrity and reliability of the data and its practical value in meteorological forecasting and hydrological analysis.
[0035] In some implementations, numerical weather forecast data, real-time meteorological data, and watershed snow cover data are collected from distributed data sources, and the collected data are cleaned and subjected to temporal and spatial standardization processing, including: obtaining forecast data from numerical weather forecast centers, collecting real-time data from automatic weather stations, and obtaining snow cover monitoring data from remote sensing platforms; filtering out invalid data from the collected data to obtain valid data; performing time alignment processing on the valid data to give all data a unified timestamp; and performing spatial resampling processing on the time-aligned data to give all data a unified spatial resolution, thus obtaining standardized data.
[0036] Numerical weather prediction centers can refer to institutions or systems responsible for running numerical weather prediction models and publishing forecast results. For example, they might acquire gridded data in GRIB format from the Global Forecast System (GFS) via FTP protocol to provide authoritative forecast data as a baseline for data completion. Forecast data can refer to the predicted values of future meteorological elements output by numerical weather prediction models, such as grid fields containing temperature, air pressure, and wind speed, converted to station data using bilinear interpolation, and used as extrapolation references in data completion. Automatic weather stations can refer to ground-based observation equipment capable of automatically measuring and transmitting meteorological parameters. For example, they might collect data through temperature, humidity, and wind speed sensors and transmit it to a central database using a 4G network to provide real-time and high-precision observations. Real-time data can refer to measurements of current meteorological conditions acquired in real time by observation equipment, such as minute-level CSV format data collected from automatic weather stations, which, after quality control marking, can be used to verify the accuracy of other data sources. Remote sensing platforms can refer to systems that remotely acquire information about the Earth's surface through satellite or airborne sensors. For example, they might use MODIS satellite imagery and the NDSI index to invert snow cover parameters to provide information on snow distribution over a wide area. Snow cover monitoring data can refer to remote sensing or ground observation data reflecting snow cover extent, depth, or snow water equivalent. For example, snowline locations can be extracted from Landsat imagery and stored in GeoTIFF format to characterize watershed hydrological conditions. Invalid data filtering refers to the process of identifying and removing outliers or erroneous records from the data. For example, applying the 3σ rule to delete temperature values outside the historical statistical range improves the reliability of the dataset. Valid data refers to a subset of data deemed reliable and usable after quality control. For example, observations within the preset interval of a dynamic boundary model are marked as clean records and can serve as the basis for subsequent processing. Time alignment refers to the operation of adjusting different time series data to a unified time reference. For example, using linear interpolation to resample minute data to hourly UTC timestamps ensures consistency in the time dimension. A unified timestamp refers to the same time reference format shared by all data points, such as using Coordinated Universal Time (UTC) and ISO 8601 standards to facilitate time series analysis and model input. Time-aligned data refers to datasets that have undergone time synchronization processing, such as aligning all variables to the hour and filling gaps through interpolation. This data can be used for spatial resampling and multivariate correlation. Spatial resampling refers to the process of transforming data from one spatial resolution or grid to another. For example, using the nearest neighbor method to resample remote sensing data from a 500-meter grid to a 1-kilometer grid can eliminate spatial scale differences. Uniform spatial resolution means that all data layers have the same pixel size or grid spacing, such as setting a 1-kilometer × 1-kilometer grid and using WGS84 projection to ensure that the data is spatially integrable.
[0037] As a concrete example: In the meteorological data standardization scenario of the upper Yangtze River basin, ECMWF GRIB format forecast data, including temperature, humidity, and wind speed grid fields, was obtained from the China Meteorological Administration Numerical Prediction Center via FTP protocol; minute-level real-time data was collected from the Sichuan Province automatic weather station network and stored in CSV format; and snow cover monitoring data was obtained from Landsat 8 imagery downloaded from the USGS remote sensing platform. Invalid data filtering applied the 3σ rule to remove temperature outliers (such as those exceeding the range of -20℃ to 40℃), obtaining valid data. Time alignment processing used linear interpolation to resample the data to hourly UTC timestamps. Spatial resampling processing used bilinear interpolation to unify data of different resolutions to a 1-kilometer grid, using the WGS84 coordinate system. The final output standardized data was used for subsequent missing data identification and completion calculations. By collecting multiple types of meteorological and hydrological data from distributed data sources and performing cleaning and spatiotemporal standardization, including invalid data filtering, time alignment, and spatial resampling, data consistency and comparability were achieved, providing high-quality input for subsequent missing data correction and improving the accuracy and reliability of meteorological and hydrological analysis.
[0038] In some implementations, the standardized data is subjected to missing data identification and physical conformity verification to obtain a missing dataset, including: scanning the data through a sliding window to identify continuous missing segments and random missing points to obtain missing location information; performing physical conformity verification on the data values based on a dynamic boundary model, where the dynamic boundary model obtains a preset range of variable values based on historical data statistics; when a data value exceeds the preset range, it is marked as physically unreasonable data and included in the missing dataset; and combining the missing location information and physically unreasonable data, a complete missing dataset is generated.
[0039] Continuous missing segments refer to intervals in a time series where data is missing at multiple consecutive time points. For example, scanning might reveal that a station has no temperature data for six consecutive hours, marking this segment as a priority for completion. Random missing points refer to isolated data gaps that occur discontinuously in a time series, such as the loss of a single moment's air pressure value due to a transmission failure. These gaps are identified and recorded through window scanning. Missing location information refers to structured information recording the specific location of data gaps, such as generating a table containing station number, variable type, and missing timestamp, to guide subsequent completion operations. Dynamic boundary models refer to mathematical models that dynamically generate preset intervals for variables based on historical statistical patterns. For example, using the mean and three standard deviations of the past 30 days' data to calculate upper and lower limits for temperature can help identify physically unreasonable values. Physical law compliance checks refer to the process of verifying whether data values conform to natural laws. For example, checking whether relative humidity is between 0 and 100% can effectively identify abnormal data caused by instrument malfunctions. Preset intervals refer to allowable intervals for variable values determined according to physical laws and historical statistics. For example, setting an upper limit for precipitation intensity based on historical percentiles can serve as a benchmark for judging the reasonableness of data. Physically irrational data can refer to observations that violate natural or statistical laws, such as a temperature of -50 degrees Celsius recorded in the equatorial region, which would be marked as invalid data after verification. A complete missing dataset can refer to a collection containing all identified missing and outlier data, such as a database integrating location information and physically irrational markers, serving as a unified input for a data completion system.
[0040] As a concrete example: In the quality control of meteorological data in the Yellow River Basin, a 24-hour sliding window is used to scan precipitation data sequences in 1-hour increments to identify continuous missing segments (e.g., no data for 8 consecutive hours at a certain station) and random missing points (e.g., missing evaporation at a single moment). A dynamic boundary model sets preset intervals for each variable based on statistical data from the past 90 days (e.g., daily temperature difference not exceeding historical maximum values). When a station's instantaneous temperature change exceeds 15 degrees Celsius, it is marked as physically unreasonable data. Finally, the missing location information and physically unreasonable data are integrated to generate a complete missing dataset containing station number, variable type, timestamp, and anomaly type, which is then submitted to the completion module for processing.
[0041] By identifying data gaps through sliding window scanning and combining them with a dynamic boundary model to verify physical conformity, a complete missing dataset containing location information and anomaly markers is systematically constructed. This provides a clear target range for subsequent accurate data completion and effectively improves the pertinence and physical rationality of data correction.
[0042] In some implementations, the training steps of the variable relationship model include: collecting historical meteorological and hydrological data, including numerical weather prediction data, real-time meteorological data, and watershed snow cover data; using a long short-term memory network to extract time series features and using a Transformer encoder to capture long-term dependencies between variables; combining causal discovery algorithms to analyze causal relationships between variables and constructing a causal relationship network; and training a pre-set initial variable relationship model based on the causal relationship network to obtain a trained variable relationship model.
[0043] Long Short-Term Memory (LSTM) networks can refer to a variant of recurrent neural networks capable of learning long-term dependencies. For example, using the hidden states and gating mechanisms of LSTM networks to process time-series data, with historical meteorological variable sequences as input and feature vectors as output, can effectively extract temporal patterns and short-term dependencies from time series data. Time-series features can refer to representative patterns or statistics extracted from time-series data, such as calculating the mean, variance, or autocorrelation coefficient through a sliding window, or automatically learned through deep learning models, used to characterize the regularity and trend of data changes over time. A Transformer encoder can refer to a neural network architecture component based on a self-attention mechanism. For example, using multi-head self-attention layers and feedforward networks to process variable sequences and calculate global dependency weights between variables can capture long-term and complex dependencies between variables. Long-term dependencies between variables can refer to the mutual influence or correlation between meteorological and hydrological variables over a long time span. For example, by analyzing the time-lag correlation between variables in historical data or using attention mechanisms for modeling, it can improve the model's predictive ability for long-term meteorological processes. Causal discovery algorithms refer to computational methods used to infer causal relationships between variables from data. For example, applying the Peter-Clark algorithm (PC) or LiNGAM (Linear Non-Gaussian Acyclic Model) to analyze conditional independence and construct causal graph structures can identify causal directions between variables, enhancing model interpretability. Causal relationships between variables can refer to the direct cause-effect relationship between one variable and another. For instance, causal discovery algorithms output directed edges representing causal directions, guiding the structural design of variable relationship models and avoiding spurious correlations. Causal relationship networks refer to network models that represent causal relationships between variables using a graph structure. For example, variables are treated as nodes, directed edges represent causal influences, and edge weights represent causal strength, visualizing causal paths between variables for model interpretation. Pre-defined initial variable relationship models refer to the initial architecture of the variable relationship model set before training begins. This includes defining the number of neural network layers, activation functions, and loss functions, using randomly initialized parameters, and serving as the starting point for the training process, which is then gradually improved through optimization.
[0044] As a concrete example: In a meteorological data modeling scenario for a river basin in South China, historical meteorological and hydrological data from the past 10 years were collected, including ECMWF numerical weather prediction data, automatic weather station data, and MODIS basin snow cover data. A Long Short-Term Memory (LSTM) network was used to extract time-series features of temperature, precipitation, and humidity, with a hidden layer dimension of 64 and an input sequence length of 30 days. A Transformer encoder with 4 heads and 256-dimensional embedding was used to capture long-term dependencies between variables. A causal discovery algorithm was applied using the PC algorithm to analyze causal relationships between variables, constructing a causal relationship network where nodes represent meteorological variables and directed edges represent causal influences. Based on this network, a pre-set initial variable relationship model was trained using the Adam optimizer and mean squared error loss function, with a learning rate of 0.001. After 100 training rounds, a trained variable relationship model was obtained, which was used for subsequent missing data completion.
[0045] By collecting historical meteorological and hydrological data and extracting features using a long short-term memory network and a Transformer encoder, a causal relationship network was constructed using a causal discovery algorithm. Based on this, an initial variable relationship model was trained, which enabled accurate modeling of the relationships between variables and capture of long-term dependencies. This improved the reliability and physical consistency of missing data completion, providing stronger data support for meteorological and hydrological analysis.
[0046] In some implementations, the step of performing imputation calculations on the missing dataset based on a variable relationship model to generate the imputed data specifically includes: for a random missing dataset, generating imputed values for meteorological and hydrological data based on the following first calculation formula: in, This indicates the completed meteorological and hydrological data, and is a non-negative value. This represents the time-related weighting coefficient, which is derived from the reliability assessment of each sub-item output by the dynamic coupling correlation analysis module and is used to adjust the contribution of the time-related item. The dynamic weight coefficient of the i-th related variable is derived from the training results of the dynamic coupling correlation analysis module and is used to quantify the degree of influence of different related variables on the target variable. This represents the dimensional transformation value of the i-th related variable, obtained through a linear transformation function. Obtain, among which These are conversion coefficients derived from least squares regression analysis of historical data, ensuring that the converted dimensions match the target variable. Consistent; The value of the i-th related variable is derived from the multi-source heterogeneous data acquisition and preprocessing module and serves as the basic input data for the completion calculation. This represents the time difference between the observation time and the completion time of the i-th related variable, derived from time series analysis of the data, and is used to introduce the time decay effect; It represents the time decay constant, which is derived from historical data statistical analysis and controls the decay rate of the time effect; This represents the first numerical stability factor, set to a very small positive real number to prevent abnormal calculations of numbers with a denominator of zero; This represents the spatial correlation weighting coefficient, derived from the accuracy evaluation results of the spatial interpolation algorithm, and is used to adjust the contribution of the spatial correlation term. The distance weight of the j-th spatial neighbor point is derived from the spatial interpolation algorithm and reflects the degree of influence of spatial location on the target point. The data quality index represents the j-th spatial neighbor point. Its value is between 0 and 1. It is calculated by the missing data intelligent completion module based on the data quality markers output by the data integrity and physical rationality verification module. This represents the observed value of the j-th spatially neighboring point, distinguished from the target variable by a wavy line symbol. It is derived from real-world data collection and provides spatial correlation information. This represents the second numerical stability factor, set to a very small positive real number to ensure that the denominator is always greater than zero; This represents the historical similarity weight coefficient, derived from the confidence assessment of the pattern matching algorithm, and is used to adjust the contribution of the historical similarity item. , and The sum of is 1; The matching weight of the kth historical similar scenario is derived from the pattern matching algorithm and quantifies the similarity between the historical scenario and the current situation. The target variable value representing the kth historically similar scenario is derived from a historical database, providing a reference for historical patterns. The third numerical stability factor is set to a very small positive real number to maintain computational stability; N represents the total number of relevant variables, M represents the total number of spatially neighboring points, and P represents the total number of historically similar scenarios, all of which are positive integers. The random missing data processing unit, when performing calculations, uses a preset quality threshold. Filter nearby points in space; when identified < It automatically excludes nearby points in the space to ensure that the denominators of all fractions in the calculation are greater than zero.
[0047] Random missing data refers to data gaps that occur randomly and discontinuously in a time series, such as the loss of precipitation records at a single moment due to a momentary transmission failure. These gaps are identified using a sliding window detection algorithm and are the primary focus of the completion calculation. Completed meteorological and hydrological variable values refer to calculated estimates used to fill data gaps. For example, temperature or humidity values output by fusing time correlation, spatial proximity, and historical similarity information using a weighted formula are used to restore the integrity of the dataset. The first calculation formula refers to the mathematical expression described in the claims used to calculate the completed value, such as a comprehensive formula including three weighted components: time correlation, spatial correlation, and historical similarity. The final completed result is generated through weighted summation. The time correlation weight coefficient refers to the weight parameter assigned to the time correlation component in the completion formula. For example, a weight value of 0.4 is obtained through grid search on the validation set and is used to adjust the contribution of time dimension information in the completion calculation. The spatial correlation weight coefficient refers to the weight parameter assigned to the spatial correlation component in the completion formula. For example, a weight value of 0.35 is determined based on cross-validation and is used to control the influence of spatially neighboring observations on the completion result.
[0048] Historical similarity weighting coefficients refer to the weighting parameters assigned to historical similarity components in the completion formula. For example, setting it to 0.25 and optimizing it through model training helps balance the reference value of historical similar scenarios in the completion calculation. Dynamic weighting coefficients for related variables refer to weighting parameters adaptively calculated based on the correlation strength between variables. For example, normalized weights are obtained by processing the absolute value of the correlation coefficient using the softmax function, reflecting the differences in the impact of different related variables on the target variable. Dimensional unification conversion values refer to values that convert variables with different dimensions to values with a unified dimension. For example, converting wind speed from meters per second to kilometers per hour using a linear function eliminates dimensional differences and ensures calculation consistency. Linear transformation functions refer to mathematical functions that achieve dimensional transformation through linear operations. For example, using the formula f(x) = ax + b to map the original value to the target dimension, where a is the transformation coefficient and b is the offset, helps maintain the linear relationship between variables. Dimensional transformation coefficients refer to multiplier parameters used in linear transformation functions to scale variables. For example, using a coefficient of 0.01 when converting pressure from hectopascals to kilopascals helps achieve numerical compatibility between different physical quantities. The time decay constant can be a parameter that controls the rate at which time correlation decays as the time difference increases. For example, it can be set to 6 hours and determined through historical data analysis to adjust the decay rate of the contribution of historical observations to the current completion.
[0049] Numerical stability factors can refer to extremely small positive numbers added to prevent the denominator from being zero, such as adding a constant of 1e-8 to the denominator to ensure the numerical stability of the calculation formula and avoid division by zero errors. Distance weights can refer to weighting coefficients calculated based on spatial distance, such as using an inverse distance weighting method to calculate the weight of spatially nearby points, with closer points having larger weights, reflecting the law of spatial correlation decay with distance. Data quality indicators can refer to quantitative indicators evaluating data reliability, such as calculating a score between 0 and 1 based on data integrity and anomaly detection results, used to select high-quality spatially nearby points in the completion calculation. Spatially nearby point observations can refer to meteorological and hydrological data collected from geographically close observation points, such as selecting temperature records from automatic weather stations within a 50-kilometer radius, used to provide reference data for spatial correlation completion. Matching degree weights can refer to weighting parameters that measure the similarity between historical and current scenarios, such as those calculated using the cosine similarity of multidimensional feature vectors, used to select the most relevant historical scenarios for completion reference. The target variable value for historically similar scenarios can refer to the observed value of the target variable in historical data when the current scenario is similar. For example, finding dates with similar meteorological conditions and obtaining their temperature records can be used for completion calculations based on historical similarity. The total number of relevant variables can refer to the number of relevant variables participating in the completion calculation. For example, selecting five relevant variables such as temperature, humidity, and air pressure can be used to determine the calculation dimension of the temporal correlation component. The total number of spatially neighboring points can refer to the number of spatially neighboring points participating in the completion calculation. For example, selecting 10 surrounding automatic weather stations as spatial reference points can be used to determine the calculation range of the spatial correlation component. The total number of historically similar scenarios can refer to the number of historically similar scenarios used for the completion calculation. For example, selecting the 20 most similar scenarios from the historical database can be used to determine the calculation sample size for the historical similarity component. The quality threshold can refer to the minimum data quality standard used to screen spatially neighboring points. For example, setting a data quality index of no less than 0.8 can be used to exclude low-quality observation points to ensure the reliability of the completion.
[0050] By integrating information from three dimensions—temporal correlation, spatial correlation, and historical similarity—a weighted calculation formula is used to complete missing data. This fully leverages the temporal and spatial correlation characteristics of meteorological data and historical similarity patterns, achieving high-precision and physically reasonable completion of missing data and significantly improving the completeness and usability of meteorological and hydrological datasets.
[0051] In some implementations, the dynamic weighting coefficients are determined by the following formula: in, This represents the scaling parameter, which originates from hyperparameter optimization during model training and is used to balance the magnitude differences in correlation coefficients. The target variable is... Related to the i-th variable The correlation coefficient is derived from the Pearson correlation coefficient statistical analysis of historical data. It is used to quantify the strength of the linear association between variables. By taking the absolute value, it ensures that the weight calculation only considers the correlation strength and is not affected by the correlation direction; N represents the total number of related variables and is a positive integer.
[0052] The second calculation formula can refer to a specific mathematical expression used to calculate dynamic weight coefficients. For example, using the softmax function to process the product of the scaling parameter and the absolute value of the correlation coefficient, outputting normalized weight values, ensures that the weight coefficients are between zero and one and sum to 1. The scaling parameter can refer to an adjustable parameter that controls the smoothness of the weight distribution. For example, the optimal value can be determined during model training through grid search or Bayesian optimization, used to adjust the strength of the correlation coefficient's influence on the weights. Hyperparameter optimization can refer to methods for adjusting hyperparameters to optimize performance during machine learning model training. For example, using cross-validation and random search strategies to find the optimal value of the scaling parameter λ, used to improve the predictive accuracy of variable relationship models. The correlation coefficient can refer to a statistical indicator that measures the strength and direction of the linear relationship between two variables. For example, calculating the Pearson correlation coefficient yields a value between negative and positive one, obtained through analysis of historical datasets, used to characterize the statistical dependency between variables. Historical data statistical analysis can refer to the process of statistical calculation and pattern recognition of historical meteorological and hydrological data. For example, calculating the correlation coefficient matrix and statistical significance between variables, implemented using Python's pandas library, to provide the statistical features required for weight calculation. The absolute value of the correlation coefficient can refer to the magnitude of the correlation coefficient without considering its sign. For example, in the second calculation formula, the absolute value of the correlation coefficient is used for exponential operations to ensure that the weights depend only on the strength of the correlation and ignore the influence of direction. Weight calculation can refer to the mathematical operation process of determining the weight coefficients of each variable. For example, the second calculation formula can be used in conjunction with the scaling parameter and the absolute value of the correlation coefficient to calculate dynamic weights. Matrix operations can be implemented through programming to allocate the contribution of variables in the completion calculation. Correlation strength can refer to the closeness of the correlation between variables. For example, it can be measured by the magnitude of the absolute value of the correlation coefficient, where zero indicates no correlation and one indicates perfect correlation. This is used to prioritize strongly correlated variables in weight calculation. Correlation direction can refer to the positive or negative characteristic of the correlation between variables. For example, positive correlation indicates that variables change in the same direction, and negative correlation indicates that they change in opposite directions. In weight calculation, the direction is ignored by taking the absolute value to avoid weight calculation bias caused by negative correlation. Physically reasonable dynamic weight coefficients can refer to weight allocation results that conform to meteorological and hydrological physical laws. For example, the weight distribution is obtained by emphasizing strongly correlated variables and ignoring weakly correlated variables to ensure the physical consistency of the completion calculation results.
[0053] By combining the scale adjustment parameter and the absolute value of the correlation coefficient with the second calculation formula to calculate the dynamic weight coefficient, and ensuring that the weight value range is reasonable and normalized, intelligent weight allocation based on the correlation strength of variables is realized, which improves the accuracy and physical consistency of missing data completion.
[0054] In some implementations, the steps of verifying the completed data, generating verification results, and updating the variable relationship model based on the verification results specifically include: comparing the completed data with the independent observation data, calculating the error index, and generating the verification results; when the error index exceeds a preset threshold, marking the completed data as data that needs correction, and receiving user correction input through a visual interface; adding the verified data to the training set, and updating the parameters of the variable relationship model through incremental learning; and periodically generating data quality assessment reports to summarize the completion effect and model performance.
[0055] Independent observational data can refer to separately collected validation data unrelated to the data source for completion, such as temperature and precipitation observations obtained from independently deployed benchmark weather stations or satellite platforms, stored in CSV format, to provide an objective benchmark for evaluating completion accuracy. Error metrics can refer to statistical measures that quantify the difference between the completed data and the true values, such as calculating the root mean square error (RMSE) or mean absolute error (MAE), with numerical results automatically generated by scripts for objectively evaluating completion quality. User-corrected input can refer to manually corrected data obtained through an interactive interface, such as users manually adjusting and submitting abnormal temperature values in a visual interface, stored in JSON format, to provide expert knowledge to improve the completion results. Validated data can refer to a dataset that has been validated and potentially corrected, such as integrating independent observation comparison results and user-input correction values, updating the main database to enhance dataset reliability and expand training samples. Incremental learning can refer to methods that gradually update model parameters without retraining the entire model, such as using online learning algorithms to dynamically adjust neural network weights with a learning rate of 0.001 to adapt the model to new data distributions. A data quality assessment report can refer to a document summarizing the effectiveness and quality of data completion. For example, it might be a PDF report containing error statistics and performance trend charts, generated periodically by automated scripts to monitor the long-term performance of the system. Completion effectiveness refers to the quality evaluation of the results after data completion, measured by metrics such as accuracy and consistency, combined with feedback from domain experts, to assess the practicality and reliability of the completion method. Model performance refers to the performance metrics of the variable relationship model in prediction or completion tasks, such as using F1 scores or explained variance, calculated through cross-validation, to guide model iteration and optimization.
[0056] As a concrete example: In a meteorological data verification scenario in the Yangtze River Delta region, the temperature data completed using the LSTM-Transformer model is compared with observation data from 10 independently set benchmark meteorological stations. The root mean square error (RMSE) is calculated as the error metric, with a preset threshold of 1.5 degrees Celsius. When the RMSE exceeds the threshold, data requiring correction is marked on the web visualization interface, and users can submit correction inputs by dragging and dropping chart points. The verified data is incorporated into the training set, and the variable relationship model parameters are updated using an incremental learning algorithm with a learning rate set to 0.0005. A data quality assessment report is automatically generated weekly, summarizing the completion effects, such as accuracy improvements, and model performance, such as changes in the F1 score, supporting meteorological forecasting business decisions.
[0057] By comparing the completed data with independent observation data to calculate error indicators, combining preset thresholds to trigger manual correction and incremental model updates, and regularly generating evaluation reports, continuous quality monitoring of the completed data and adaptive optimization of the model are achieved, thereby improving the reliability and practical value of meteorological data products.
[0058] In some implementations, the method further includes: when the amount of missing data exceeds a preset threshold, activating an alternative completion strategy, which includes using one or more combinations of time series prediction, spatial interpolation, or historical mean methods to complete the data and generate alternative completion data.
[0059] Missing data volume refers to the number or proportion of missing or invalid records in a dataset. For example, it can be calculated as a percentage of missing records relative to the total number of records, used to assess data integrity and determine the completion strategy. Alternative completion strategies refer to alternative data processing schemes used when the primary completion method is unavailable. For example, when the variable relationship model fails, a combination of time series forecasting and spatial interpolation can be used, implemented through a strategy selector, to ensure the system's continued operation under abnormal conditions. The activation process may include automatically calling the backup algorithm module when the detected missing data volume exceeds a threshold, implemented through conditional statements, to ensure the robustness of the completion system. Time series forecasting refers to statistical methods that infer future values based on historical data. For example, using the Autoregressive Integral Moving Average (ARIMA) model to predict temperature change trends, with historical data as input and predicted values as output, provides time-dimensional completion data. Spatial interpolation refers to geostatistical methods that extrapolate values at unknown locations based on spatially adjacent points. For example, using the Kriging algorithm to generate grid data based on observations from surrounding stations provides spatial-dimensional completion results. Historical mean methods refer to simple methods that use the statistical average of historical data from the same period to fill in the gaps. For example, calculating the average temperature of the same date over the past ten years as the fill value, implemented by querying a historical database using SQL, is used to provide a stable benchmark reference. Alternative fill data refers to alternative fill results generated through backup strategies. For example, when the main model fill fails, ARIMA predictions are used as temporary fills and stored as backup fields to ensure the timely output of data products.
[0060] As a concrete example: In meteorological data processing in the Yellow River Basin, when the amount of missing snow cover data for a given day exceeds a preset threshold of 40%, the system automatically activates a backup completion strategy. This strategy combines time series forecasting (using an ARIMA model to predict snow cover changes over the next three days), spatial interpolation (using kriging to generate a spatial distribution based on data from twenty surrounding stations), and historical mean methods (calculating the average snow cover rate for the same period over the past fifteen years). The completion results from these three methods are fused using a quality-weighted algorithm to generate alternative completed data, which is then stored in the database, ensuring that hydrological forecasting operations are not affected by missing data.
[0061] By enabling a backup completion strategy that includes time series prediction, spatial interpolation, and historical mean methods when the amount of missing data exceeds a preset threshold, and supporting the combined use of multiple methods, the meteorological and hydrological data completion system ensures continuous operation when the main model fails, thereby improving the robustness of the data processing workflow and the reliability of the results.
[0062] Corresponding to the above method embodiments, this specification also provides an embodiment of a meteorological missing dataset correction system based on correlation analysis. Figure 2This specification illustrates a schematic diagram of the structure of a meteorological missing dataset correction system based on correlation analysis, provided in some embodiments of this specification. For example... Figure 2 As shown, the system includes: In some implementations, numerical weather forecast data, real-time meteorological data, and watershed snow cover data are collected from distributed data sources, and the collected data are cleaned and subjected to temporal and spatial standardization processes, including: Forecast data is obtained from the numerical weather prediction center, real-time data is collected from automatic weather stations, and snow cover monitoring data is obtained from the remote sensing platform; The collected data is filtered to obtain valid data; Perform time alignment on valid data to give all data a uniform timestamp; Spatial resampling is performed on the time-aligned data to give all data a uniform spatial resolution, resulting in standardized data.
[0063] In some implementations, missing data is identified and physical conformity checks are performed on standardized data to obtain a missing dataset, including: By scanning the data through a sliding window, continuous missing segments and random missing points are identified to obtain the missing location information; The dynamic boundary model is used to verify the physical laws of the data values. The dynamic boundary model obtains the preset range of variable values based on historical data statistics. When a data value exceeds a preset range, it is marked as physically unreasonable data and included in the missing dataset. By combining missing location information and physically invalid data, a complete missing dataset is generated.
[0064] In some implementations, the training steps for the variable relationship model include: Collect historical meteorological and hydrological data, including numerical weather prediction data, real-time meteorological data, and watershed snow cover data; Long Short-Term Memory Network is used to extract time series features, and Transformer encoder is used to capture long-term dependencies between variables. By combining causal discovery algorithms to analyze causal relationships between variables, a causal relationship network is constructed. The trained variable relationship model is obtained by training a pre-set initial variable relationship model based on a causal relationship network.
[0065] In some implementations, the steps of performing completion calculations on the missing dataset based on a variable relationship model to generate completed data specifically include: For a randomly missing dataset, complete values for meteorological and hydrological data are generated based on the following first calculation formula: in, This indicates the completed meteorological and hydrological data, and is a non-negative value. This represents the time-related weighting coefficient, which is derived from the reliability assessment of each sub-item output by the dynamic coupling correlation analysis module and is used to adjust the contribution of the time-related item. The dynamic weight coefficient of the i-th related variable is derived from the training results of the dynamic coupling correlation analysis module and is used to quantify the degree of influence of different related variables on the target variable. This represents the dimensional transformation value of the i-th related variable, obtained through a linear transformation function. Obtain, among which These are conversion coefficients derived from least squares regression analysis of historical data, ensuring that the converted dimensions match the target variable. Consistent; The value of the i-th related variable is derived from the multi-source heterogeneous data acquisition and preprocessing module and serves as the basic input data for the completion calculation. This represents the time difference between the observation time and the completion time of the i-th related variable, derived from time series analysis of the data, and is used to introduce the time decay effect; It represents the time decay constant, which is derived from historical data statistical analysis and controls the decay rate of the time effect; This represents the first numerical stability factor, set to a very small positive real number to prevent abnormal calculations of numbers with a denominator of zero; This represents the spatial correlation weighting coefficient, derived from the accuracy evaluation results of the spatial interpolation algorithm, and is used to adjust the contribution of the spatial correlation term. The distance weight of the j-th spatial neighbor point is derived from the spatial interpolation algorithm and reflects the degree of influence of spatial location on the target point. The data quality index represents the j-th spatial neighbor point. Its value is between 0 and 1. It is calculated by the missing data intelligent completion module based on the data quality markers output by the data integrity and physical rationality verification module. This represents the observed value of the j-th spatially neighboring point, distinguished from the target variable by a wavy line symbol. It is derived from real-world data collection and provides spatial correlation information. This represents the second numerical stability factor, set to a very small positive real number to ensure that the denominator is always greater than zero; This represents the historical similarity weight coefficient, derived from the confidence assessment of the pattern matching algorithm, and is used to adjust the contribution of the historical similarity item. , and The sum of is 1; The matching weight of the kth historical similar scenario is derived from the pattern matching algorithm and quantifies the similarity between the historical scenario and the current situation. The target variable value representing the kth historically similar scenario is derived from a historical database, providing a reference for historical patterns. The third numerical stability factor is set to a very small positive real number to maintain computational stability; N represents the total number of relevant variables, M represents the total number of spatially neighboring points, and P represents the total number of historically similar scenarios, all of which are positive integers. The random missing data processing unit, when performing calculations, uses a preset quality threshold. Filter nearby points in space; when identified < It automatically excludes nearby points in the space to ensure that the denominators of all fractions in the calculation are greater than zero.
[0066] In some implementations, the dynamic weighting coefficients are determined by the following formula: in, This represents the scaling parameter, which originates from hyperparameter optimization during model training and is used to balance the magnitude differences in correlation coefficients. The target variable is... Related to the i-th variable The correlation coefficient is derived from the Pearson correlation coefficient statistical analysis of historical data. It is used to quantify the strength of the linear association between variables. By taking the absolute value, it ensures that the weight calculation only considers the correlation strength and is not affected by the correlation direction; N represents the total number of related variables and is a positive integer.
[0067] In some implementations, the steps of validating the completed data, generating a validation result, and updating the variable relationship model based on the validation result specifically include: The completed data is compared with the independent observation data, the error index is calculated, and the verification result is generated. When the error index exceeds the preset threshold, the completed data is marked as data that needs to be corrected, and the user's correction input is received through the visual interface; The validated data is added to the training set, and the parameters of the variable relationship model are updated through incremental learning. Regularly generate data quality assessment reports to summarize the completion effects and model performance.
[0068] In some implementations, the method further includes: When the amount of missing data exceeds a preset threshold, an alternative completion strategy is activated. The alternative completion strategy includes using one or more combinations of time series forecasting, spatial interpolation, or historical mean methods to complete the data and generate alternative completed data.
[0069] The above is an illustrative scheme of a meteorological missing dataset correction system based on correlation analysis according to this embodiment. It should be noted that the technical solution of this meteorological missing dataset correction system based on correlation analysis belongs to the same concept as the technical solution of the meteorological missing dataset correction method based on correlation analysis described above. Details not described in detail in the technical solution of the meteorological missing dataset correction system based on correlation analysis can be found in the description of the technical solution of the meteorological missing dataset correction method based on correlation analysis described above.
[0070] Figure 3 A structural block diagram of a computing device 300 according to some embodiments of this specification is shown. The components of the computing device 300 include, but are not limited to, a memory 301 and a processor 302. The processor 302 is connected to the memory 301 via a bus 303, and a database 305 is used to store data.
[0071] The computing device 300 also includes an access device 304 that enables the computing device 300 to communicate via one or more networks 306. Examples of such networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 304 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0072] In one embodiment of this specification, the aforementioned components of the computing device 300 and Figure 3 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 3The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0073] The computing device 300 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 300 can also be a mobile or stationary server.
[0074] The processor 302 executes the following computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned method for correcting missing meteorological datasets based on correlation analysis. The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the aforementioned method for correcting missing meteorological datasets based on correlation analysis belong to the same concept. Details not described in detail in the technical solution of the computing device can be found in the description of the technical solution of the aforementioned method for correcting missing meteorological datasets based on correlation analysis.
[0075] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described method for correcting missing meteorological datasets based on correlation analysis.
[0076] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the above-described method for correcting missing meteorological datasets based on correlation analysis. Details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the above-described method for correcting missing meteorological datasets based on correlation analysis.
[0077] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described method for correcting missing meteorological datasets based on correlation analysis.
[0078] The above is an illustrative example of a computer program in this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solution of the aforementioned method for correcting missing meteorological datasets based on correlation analysis. Details not described in detail in the computer program's technical solution can be found in the description of the aforementioned method for correcting missing meteorological datasets based on correlation analysis.
[0079] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0080] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0081] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0082] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0083] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this invention. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for correcting missing meteorological datasets based on correlation analysis, characterized in that, include: Numerical weather forecast data, real-time meteorological data, and watershed snow cover data are collected from distributed data sources, and the collected data are cleaned and subjected to temporal and spatial standardization to obtain standardized data. The standardized data is subjected to missing data identification and physical law conformity verification to obtain the missing dataset; The missing dataset is filled in using the trained variable relationship model to generate the filled data. The variable relationship model is built and trained based on historical data. The filled data is then validated to generate a validation result, and the variable relationship model is updated based on the validation result.
2. The method according to claim 1, characterized in that, The process of collecting numerical weather forecast data, real-time meteorological data, and watershed snow cover data from distributed data sources, and cleaning and performing temporal and spatial standardization on the collected data, includes: Forecast data is obtained from the numerical weather prediction center, real-time data is collected from automatic weather stations, and snow cover monitoring data is obtained from the remote sensing platform; The collected data is filtered to obtain valid data; Perform time alignment on valid data to give all data a uniform timestamp; Spatial resampling is performed on the time-aligned data to give all data a uniform spatial resolution, resulting in standardized data.
3. The method according to claim 1, characterized in that, The process of identifying missing data and verifying its conformity to physical laws in standardized data yields a missing dataset, including: By scanning the data through a sliding window, continuous missing segments and random missing points are identified to obtain the missing location information; The dynamic boundary model is used to verify the physical laws of the data values. The dynamic boundary model obtains the preset range of variable values based on historical data statistics. When a data value exceeds a preset range, it is marked as physically unreasonable data and included in the missing dataset. By combining missing location information and physically invalid data, a complete missing dataset is generated.
4. The method according to claim 1, characterized in that, The training steps for the variable relationship model include: Collect historical meteorological and hydrological data, including numerical weather prediction data, real-time meteorological data, and watershed snow cover data; Long Short-Term Memory Network is used to extract time series features, and Transformer encoder is used to capture long-term dependencies between variables. By combining causal discovery algorithms to analyze causal relationships between variables, a causal relationship network is constructed. The trained variable relationship model is obtained by training a pre-set initial variable relationship model based on a causal relationship network.
5. The method according to claim 1, characterized in that, The step of performing completion calculations on the missing dataset based on the variable relationship model to generate completed data specifically includes: For a randomly missing dataset, complete values for meteorological and hydrological data are generated based on the following first calculation formula: in, This indicates the completed meteorological and hydrological data, and is a non-negative value. This represents the time-related weighting coefficient, which is derived from the reliability assessment of each sub-item output by the dynamic coupling correlation analysis module and is used to adjust the contribution of the time-related item. The dynamic weight coefficient of the i-th related variable is derived from the training results of the dynamic coupling correlation analysis module and is used to quantify the degree of influence of different related variables on the target variable. This represents the dimensional transformation value of the i-th related variable, obtained through a linear transformation function. Obtain, among which These are conversion coefficients derived from least squares regression analysis of historical data, ensuring that the converted dimensions match the target variable. Consistent; The value of the i-th related variable is derived from the multi-source heterogeneous data acquisition and preprocessing module and serves as the basic input data for the completion calculation. This represents the time difference between the observation time and the completion time of the i-th related variable, derived from time series analysis of the data, and is used to introduce the time decay effect; It represents the time decay constant, which is derived from historical data statistical analysis and controls the decay rate of the time effect; This represents the first numerical stability factor, set to a very small positive real number to prevent abnormal calculations of numbers with a denominator of zero; This represents the spatial correlation weighting coefficient, derived from the accuracy evaluation results of the spatial interpolation algorithm, and is used to adjust the contribution of the spatial correlation term. The distance weight of the j-th spatial neighbor point is derived from the spatial interpolation algorithm and reflects the degree of influence of spatial location on the target point. The data quality index represents the j-th spatial neighbor point. Its value is between 0 and 1. It is calculated by the missing data intelligent completion module based on the data quality markers output by the data integrity and physical rationality verification module. This represents the observed value of the j-th spatially neighboring point, distinguished from the target variable by a wavy line symbol. It is derived from real-world data collection and provides spatial correlation information. This represents the second numerical stability factor, set to a very small positive real number to ensure that the denominator is always greater than zero; This represents the historical similarity weight coefficient, derived from the confidence assessment of the pattern matching algorithm, and is used to adjust the contribution of the historical similarity item. , and The sum of is 1; The matching weight of the kth historical similar scenario is derived from the pattern matching algorithm and quantifies the similarity between the historical scenario and the current situation. The target variable value representing the kth historically similar scenario is derived from a historical database, providing a reference for historical patterns. The third numerical stability factor is set to a very small positive real number to maintain computational stability; N represents the total number of relevant variables, M represents the total number of spatially neighboring points, and P represents the total number of historically similar scenarios, all of which are positive integers. The random missing data processing unit, when performing calculations, uses a preset quality threshold. Filter nearby points in space; when identified < It automatically excludes nearby points in the space to ensure that the denominators of all fractions in the calculation are greater than zero.
6. The method according to claim 5, characterized in that, The dynamic weighting coefficient is determined by the following formula: in, This represents the scaling parameter, which originates from hyperparameter optimization during model training and is used to balance the magnitude differences in correlation coefficients. Denotes the target variable as Related to the i-th variable The correlation coefficient is derived from the Pearson correlation coefficient statistical analysis of historical data. It is used to quantify the strength of the linear association between variables. By taking the absolute value, it ensures that the weight calculation only considers the correlation strength and is not affected by the correlation direction; N represents the total number of related variables and is a positive integer.
7. The method according to claim 1, characterized in that, The steps of validating the completed data, generating a validation result, and updating the variable relationship model based on the validation result specifically include: The completed data is compared with the independent observation data, the error index is calculated, and the verification result is generated. When the error index exceeds the preset threshold, the completed data is marked as data that needs to be corrected, and the user's correction input is received through the visual interface; The validated data is added to the training set, and the parameters of the variable relationship model are updated through incremental learning. Regularly generate data quality assessment reports to summarize the completion effects and model performance.
8. The method according to claim 1, characterized in that, The method further includes: When the amount of missing data exceeds a preset threshold, an alternative completion strategy is activated. The alternative completion strategy includes using one or more combinations of time series forecasting, spatial interpolation, or historical mean methods to complete the data and generate alternative completed data.
9. A meteorological missing dataset correction system based on correlation analysis, characterized in that, include: The standardization processing module is used to collect numerical weather forecast data, real-time meteorological data, and watershed snow cover data from distributed data sources, and to clean and perform temporal and spatial standardization processing on the collected data to obtain standardized data. The verification module is used to identify missing data and verify the conformity of physical laws to the standardized data, thereby obtaining the missing dataset. The calculation module is used to perform completion calculations on the missing dataset based on the trained variable relationship model to generate completed data. The variable relationship model is constructed and trained based on historical data. The module also verifies the completed data to generate a verification result and updates the variable relationship model based on the verification result.
10. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the meteorological missing dataset correction method based on correlation analysis as described in any one of claims 1 to 8.