Meteorological observation data automatic quality evaluation method and system

By using an automated meteorological observation data quality assessment method, the problems of low efficiency and poor accuracy in traditional methods have been solved. This method enables environmentally adaptive initialization and intelligent data processing, generating multi-dimensional meteorological quality assessment reports and improving the efficiency and accuracy of meteorological data processing.

CN120410340AActive Publication Date: 2025-08-01辽宁省生态气象和卫星遥感中心

Patent Information

Application Number
CN202510915995.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Traditional meteorological observation data quality assessment relies on manual operation, which is time-consuming and error-prone. It cannot automatically identify missing components, configure and deploy them, lacks path verification mechanisms, leading to program interruptions, and cannot automatically generate standardized reports, thus failing to meet the needs of large-scale data processing.

Method used

An automated quality assessment method is adopted, including environmental adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis, error tolerance interval analysis, and automated report generation. The data is processed automatically using the R software package. The EMI environmental meteorological condition assessment index is introduced, multi-level grouped statistical analysis is performed, and a standardized report is generated.

Benefits of technology

It has automated and standardized meteorological data quality assessment, improved efficiency and accuracy, reduced human error, provided multi-dimensional comprehensive analysis capabilities, ensured system stability and robustness, and generated detailed quality assessment reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410340A_ABST
    Figure CN120410340A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of meteorological quality evaluation, and discloses an automatic quality evaluation method and system for meteorological observation data, and the method comprises the steps: collecting observation data, and obtaining a meteorological evaluation parameter set through initial configuration; carrying out preprocessing operation to obtain a meteorological evaluation full data set; executing global statistical analysis at the same time to obtain a global statistical analysis data set; dividing the global statistical analysis data set into meteorological quality evaluation subsets according to region and depth hierarchy; a grading statistical analysis result is obtained; the method comprises the following steps of: setting a data proportion in each error allowable interval by acquiring an absolute error and a relative error; and generating meteorological quality distribution characteristics of different regions and levels to obtain a time sequence difference, and finally outputting a final meteorological quality evaluation report. Therefore, environment adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis, error allowable interval analysis and automatic report generation can be realized, and the efficiency of meteorological data quality evaluation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of meteorological quality assessment, and more particularly, to a method and system for automatic quality assessment of meteorological observation data. Background Art

[0002] Meteorological observation data is a crucial foundation for weather forecasting, climate research, and environmental monitoring. Its quality directly impacts the accuracy and reliability of meteorological services. With the continuous expansion of meteorological observation networks and increased automation, the volume of meteorological observation data is growing exponentially. Traditional manual quality assessment methods are no longer sufficient for large-scale data processing.

[0003] Currently, traditional meteorological observation data quality assessment and comparison rely on manual operations, requiring the calculation of statistical indicators item by item, which is time-consuming and prone to errors. Existing tools require manual installation of dependency packages and configuration of the environment, and are unable to automatically identify missing components and complete configuration deployment; nor can they automatically generate standardized reports. There is a lack of a path verification mechanism when loading data, which can easily lead to program interruptions due to missing files.

[0004] Therefore, how to achieve environmental adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis, error tolerance analysis, and automatic report generation to improve the efficiency and accuracy of meteorological data quality assessment has become an urgent problem to be solved. Summary of the Invention

[0005] The present invention provides a method and system for automated quality assessment of meteorological observation data, which solves the technical problem in the prior art that the ultra-early stage risk of autoimmune diseases cannot be effectively predicted.

[0006] The present invention provides a method and system for automatically evaluating the quality of meteorological observation data, comprising:

[0007] In a first aspect, a method for automatically evaluating the quality of meteorological observation data comprises the following steps:

[0008] Collect multiple meteorological observation data and perform initial configuration to obtain a set of meteorological assessment parameters;

[0009] Perform preprocessing operations on various parameters in the meteorological assessment parameter set to obtain the full meteorological assessment data set;

[0010] Performing global statistical analysis on the entire meteorological assessment data set to obtain a global statistical analysis data set;

[0011] The global statistical analysis dataset is divided into meteorological quality assessment subsets according to regional and depth levels; and the meteorological quality assessment subsets are further analyzed to obtain hierarchical statistical analysis results;

[0012] Obtain the absolute error and relative error of the meteorological evaluation manual station and automatic station data according to the hierarchical statistical analysis results, set the proportion of data within each error tolerance interval, and obtain the meteorological quality characteristic data through iterative verification;

[0013] Generate the meteorological quality distribution characteristics of different regions and levels based on the meteorological quality characteristic data to obtain the time series differences between the meteorological evaluation manual station and automatic station data, and save the generated data and the obtained data to folders respectively. Output the final meteorological quality evaluation report by matching the output formats set by different users.

[0014] Furthermore, the initialization configuration includes:

[0015] Create and define multiple initial states, respectively mark the meteorological observation data as correct, suspicious, incorrect, and missing, to obtain the marked data of the initial state of the meteorological observation data;

[0016] Construct meteorological quality evaluation indicators based on the marked data, and the meteorological quality evaluation indicators are the actual existence rate, correct rate, error rate, and suspicious rate;

[0017] Detect whether the initialization configuration of the meteorological observation data is completed in the current meteorological observation environment according to the marked data;

[0018] Load the data allocation resource configuration for the meteorological observation data that has not completed the initialization configuration according to the meteorological quality evaluation indicators. The data allocation resource configuration dynamically adjusts the memory usage and thread allocation according to the data scale and resources to complete the initialization configuration of the meteorological quality evaluation and obtain the initialized meteorological evaluation parameter set.

[0019] Furthermore, the data preprocessing includes:

[0020] Receive the configuration path parameter of the environment adaptive initialization configuration, and the configuration path parameter is used to cooperate with the meteorological observation data for the initialization configuration operation;

[0021] Automatically verify the configuration path parameter to judge whether the configuration path parameter exists and is accessible;

[0022] When the configuration path parameter exists and is accessible, load the meteorological observation data, and read and check whether the meteorological observation data is complete;

[0023] When the meteorological observation data is incomplete, perform format conversion adjustment and missing value supplementation processing on the meteorological observation data, and at the same time perform normalization processing to obtain the processed meteorological evaluation parameter set, which is the meteorological evaluation full data set.

[0024] Furthermore, the operation environment quality evaluation configuration includes:

[0025] Create multiple quality assessment codes, where the quality assessment codes include 0 - 3 assessment thresholds as quality assessment digital codes;

[0026] The quality assessment data are used to mark the current multiple meteorological observation data respectively to obtain a marking result;

[0027] Set digital ranges for the marking result, and each digital range represents four quality assessment states: excellent, good, average, and poor respectively;

[0028] Construct a meteorological quality identification system according to the digital ranges set by the quality assessment codes and the quality assessment states;

[0029] Based on the meteorological quality identification system, identify the dynamic changes of the current multiple meteorological observation data at different times to obtain the change trend of the meteorological observation state;

[0030] According to the change trend of the meteorological observation state, perform multi - dimensional statistical analysis through the loaded R software package to obtain the quality assessment state corresponding to the quality assessment code of the meteorological observation data in the multi - dimensional analysis, which is the quality state of the observation data under multi - dimensional analysis;

[0031] Among them, as the multi - dimensional analysis progresses, the data state will be re - judged; each state (such as normal, mildly abnormal, strongly abnormal) will be mapped to a quality assessment code; the quality assessment code is not fixed once and for all, but can change as the analysis dimension deepens. For example: original observation data (temperature, humidity, wind speed, etc.) → preliminary quality assessment (based on single - point inspection) → initial assessment code (such as: good) → multi - dimensional analysis (time trend, spatial comparison, climate limit, variable relationship) → state change identification (such as identified as deviating from the normal state / inconsistent) → quality assessment code update (good → abnormal) → quality assessment state result generation (abnormal → "needs to be corrected");

[0032] And judge whether the change trend amplitude between the current meteorological observation data and the meteorological observation data after dynamic change is in a gentle or rapid state to configure a hierarchical statistical analysis strategy.

[0033] Furthermore, the hierarchical statistical analysis strategy includes:

[0034] Capture the abnormal states of the dynamic changes of the meteorological observation data at different times to obtain abnormal state data;

[0035] Interpolate the meteorological time - series characteristics and spatial correlation of the abnormal state data to obtain a sharp drop or sharp rise amplitude index, and use the quality assessment code to mark the abnormal state data according to the amplitude index to obtain a secondary marking result;

[0036] Further perform climate threshold value checks, internal consistency checks, temporal consistency checks, and spatial consistency checks on the abnormal status data in the secondary marking results, and mark the data that does not conform to the meteorological inspection rules using quality assessment codes to obtain the tertiary marking results;

[0037] Perform fitting processing on the meteorological observation data with general or poor quality assessment status in the three rounds of marking results to obtain the processing results;

[0038] Integrate the meteorological observation data after the processing results with the original meteorological observation data with excellent or good quality assessment status to generate a preliminary quality assessment result and transmit it to the global statistical analysis for collaborative analysis and processing.

[0039] Furthermore, perform global statistical analysis on the entire meteorological assessment dataset to obtain a global statistical analysis dataset, including:

[0040] Set a meteorological micro-interference filtering threshold and dynamically adjust it according to the meteorological regional characteristics in combination with the preliminary quality assessment results to optimize the observation quality of each data item in the entire meteorological assessment dataset;

[0041] Obtain the extreme values, means, standard deviations, and coefficient of variation basic statistical indicators of the meteorological elements of each data item in the entire meteorological assessment dataset, and establish a global meteorological quality assessment benchmark;

[0042] Based on the global meteorological quality assessment benchmark, perform multi-level grouped statistical analysis on the meteorological data according to the time dimension, space dimension, and element dimension to obtain a hierarchical statistical analysis strategy to form an environmental quality assessment dataset;

[0043] Combine the hierarchical statistical analysis strategy with the meteorological quality identification system and the quality assessment indicators of the actual occupancy rate, correct rate, error rate, and suspicious rate at each level, and introduce the EMI environmental meteorological condition assessment index for comprehensive quantitative analysis;

[0044] The EMI environmental meteorological condition assessment index represents a comprehensive index characterizing processes such as aerosol emission, deposition, transport, and diffusion under the influence of meteorological conditions. The larger the EMI value, the more unfavorable the meteorological conditions are for the diffusion of air pollutants;

[0045] Through iterative statistical verification and threshold determination, fuse the global statistical results and the hierarchical statistical results, and finally generate a complete global statistical analysis dataset containing multi-dimensional statistical features, quality assessment identifiers, and the EMI environmental meteorological condition assessment index.

[0046] Furthermore, divide the global statistical analysis dataset into meteorological quality assessment subsets according to regions and depth levels; and further analyze the meteorological quality assessment subsets to obtain hierarchical statistical analysis results, including:

[0047] Stratify each data item in the global statistical analysis dataset by time dimension, including grouping statistics by year, season, month, and day, and combine the quality assessment marking results to form a meteorological quality assessment subset at the time level, generating time-series basic data;

[0048] Stratify each data item in the global statistical analysis dataset by space dimension, including grouping statistics by region, geographical features, and elevation, and use the regional characteristics to adjust the results to form a meteorological quality assessment subset at the space level, generating spatial distribution data;

[0049] Stratify each data item in the global statistical analysis dataset by element dimension, calculate the quality assessment indicators of meteorological elements such as temperature, humidity, air pressure, and wind speed respectively, and obtain the hierarchical statistical analysis results at the element level, generating element distribution characteristic data;

[0050] Adopt the EMI environmental meteorological condition assessment index for comprehensive analysis to obtain the hierarchical statistical analysis results;

[0051] Carry out quantitative separation analysis by quantitatively characterizing the impact of meteorological conditions on the quality assessment of data items in each meteorological quality assessment subset to obtain the quantitative separation analysis results;

[0052] The said quantitative separation analysis includes: based on the EMI environmental meteorological condition assessment index, obtain the emission change rate RE and the meteorological condition change rate RW through RE=(R1 / R0) / (E1 / E0)-1 and RW=(E1 / E0)-1; where, R0 and R1 are the measured meteorological data parameters in the comparison period respectively, and E0 and E1 are the EMI environmental meteorological condition assessment indexes in the comparison period respectively;

[0053] Analyze the contribution rate of meteorological condition changes to data quality through the EMI environmental meteorological condition assessment index to obtain the correlation between the EMI environmental meteorological condition assessment index and the meteorological observation data quality index, and quantitatively evaluate the impact degree of meteorological conditions on the quality of observation data;

[0054] Based on the spatio-temporal distribution characteristics of EMI, analyze the changing trends of meteorological conditions in different regions and the impact degree on the quality of meteorological observation data;

[0055] Integrate the two impact degrees to obtain the quantitative separation analysis results.

[0056] Furthermore, obtain the absolute error and relative error between the meteorological assessment manual station and the automatic station data according to the hierarchical statistical analysis results, and set the proportion of data within each error tolerance interval, and obtain the meteorological quality characteristic data through iterative verification, including:

[0057] The error tolerance interval analysis is based on the stratified statistical analysis results and the EMI environmental meteorological condition evaluation index, and is implemented by constructing an error statistical function. The error statistical function receives the meteorological evaluation manual station observation data and the automatic station observation data as input parameters and performs the following operations:

[0058] Calculate the absolute error value and the relative error percentage between the two sets of meteorological observation data in combination with the stratified statistical analysis results at the element level;

[0059] And use the meteorological quality evaluation subset at the time and space levels to statistically analyze the proportion of data with a relative error within the ±5% threshold range and an absolute error value within the ±10% threshold range;

[0060] And calculate the standard deviation based on the EMI environmental meteorological condition evaluation index, and finally return it to the statistical result list including the error interval proportion and the standard deviation of the absolute error value and the relative error. And generate an error interval ratio distribution table according to the calculation results, and transfer it as a key quantitative index to the meteorological quality characteristic data.

[0061] Furthermore, based on the element distribution characteristic data and the error interval ratio distribution table, dynamically bind the data fields through the ggplot2 package to obtain the probability density distribution characteristics of meteorological elements in different regions and levels, and combine the meteorological quality identification system to label the data quality level to support the basic data to cooperate with the time series characteristic analysis;

[0062] Use the probability density distribution characteristics and the time dimension stratified data, combine the error statistical results to generate a time series difference, so as to obtain the time change trend of the meteorological evaluation manual station and the automatic station data. Distinguish the data points within different error tolerance intervals through error coding, and transfer the time series characteristic data to the spatial characteristic analysis;

[0063] Based on the data characteristics obtained from the time series characteristic analysis and the spatial dimension stratified data, combine the error interval ratio distribution to obtain the geospatial thermal effect, so as to obtain the spatial distribution characteristics of the meteorological quality evaluation index, and dynamically adjust the error mapping range according to the EMI environmental meteorological condition evaluation index, and store the spatial distribution characteristics at the same time;

[0064] According to the meteorological quality characteristic data obtained from different dimension analyses of the probability density distribution, the time series characteristic analysis and the spatial characteristic analysis, generate a comprehensive report on the meteorological quality distribution characteristics in different regions and levels, integrate the time series difference analysis results of the meteorological evaluation manual station and the automatic station data, and save the generated data and the obtained data to the corresponding folders according to the analysis dimensions respectively; that is, output the final meteorological quality evaluation report by matching the output formats set by different users.

[0065] In a second aspect, a meteorological observation data automatic quality evaluation system for implementing a meteorological observation data automatic quality evaluation method includes:

[0066] Data initialization module: It is used to collect multiple meteorological observation data and perform initialization configuration to obtain a set of meteorological evaluation parameters;

[0067] Data preprocessing module: It is used to perform preprocessing operations on each parameter in the meteorological evaluation parameter set to obtain a complete meteorological evaluation data set;

[0068] Statistical analysis module: It is used to perform global statistical analysis on the complete meteorological evaluation data set to obtain a global statistical analysis data set; divide the global statistical analysis data set into meteorological quality evaluation subsets according to regions and depth levels; and further analyze the meteorological quality evaluation subsets to obtain hierarchical statistical analysis results;

[0069] Error analysis module: It is used to obtain the absolute error and relative error between the meteorological evaluation manual station and automatic station data according to the hierarchical statistical analysis results, set the proportion of data within each error tolerance interval, and obtain meteorological quality characteristic data through iterative verification;

[0070] Report generation module: It is used to generate meteorological quality distribution characteristics in different regions and levels according to the meteorological quality characteristic data to obtain the time series difference between the meteorological evaluation manual station and automatic station data, save the generated data and the obtained data to folders respectively, and output the final meteorological quality evaluation report by matching the output formats set by different users.

[0071] The beneficial effects of the present invention are as follows: Through environmental adaptive initialization configuration and intelligent dependency package management, the present invention realizes the full-process automation from environmental configuration to data processing, greatly improving work efficiency; By establishing a meteorological quality assessment identification system for multi-round marking and multi-dimensional consistency checks (climate threshold checks, internal consistency checks, time consistency checks, spatial consistency checks), it effectively reduces the problem of error-prone manual operations and enhances the accuracy and reliability of data processing; By introducing the EMI environmental meteorological condition assessment index, it realizes multi-level grouping statistical analysis in the time dimension, space dimension, and element dimension, can quantitatively separate the impact of meteorological condition changes on data quality, and provides a more comprehensive and accurate multi-dimensional comprehensive analysis ability; Through path verification mechanisms, format conversion adjustments, missing value supplementation processing, and abnormal state capture, it solves the problem of program interruption caused by file loss in traditional methods, realizes intelligent fault tolerance and exception handling, and improves the stability and robustness of the system; Based on the ggplot2 package, it realizes data visualization and automatically generates a standardized report including probability density distribution, time series feature analysis, and spatial distribution features, supporting multiple output formats, solving the technical problem of inability to automatically generate a standardized report; Through a computing resource allocation strategy that dynamically adjusts memory usage and thread allocation, and adaptively optimizes according to data scale and system resources, it realizes resource optimization and performance improvement, and improves the efficiency of large-scale meteorological data processing; By setting an allowable interval with a ±5% relative error threshold and a ±10% absolute error threshold, combined with an iterative verification mechanism, it realizes the precise quantification of the data differences between meteorological assessment manual stations and automatic stations, and finally obtains the final meteorological quality assessment report, thereby improving the error quantification and quality assessment accuracy and providing a scientific basis for meteorological data quality assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 is a schematic flowchart of a method for automated quality assessment of meteorological observation data provided in an embodiment of the present invention;

[0073] Figure 2 is a schematic diagram of the modules of a system for automated quality assessment of meteorological observation data provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] Now, the subject matter described herein will be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0075] At least one embodiment of the present invention discloses a method and system for automatic quality assessment of meteorological observation data, including:

[0076] As Figure 1 shown, a method for automatic quality assessment of meteorological observation data includes the following steps:

[0077] Step 1: Collect multiple meteorological observation data and perform initialization configuration to obtain a set of meteorological evaluation parameters;

[0078] Step 2: Perform preprocessing operations on each parameter in the set of meteorological evaluation parameters to obtain a complete set of meteorological evaluation data;

[0079] Step 3: Perform global statistical analysis on the complete set of meteorological evaluation data to obtain a global statistical analysis data set;

[0080] Step 4: Divide the global statistical analysis data set into meteorological quality assessment subsets according to regions and depth levels; and further analyze the meteorological quality assessment subsets to obtain a hierarchical statistical analysis result;

[0081] Step 5: Obtain the absolute error and relative error between the meteorological evaluation manual station and automatic station data according to the hierarchical statistical analysis result, and set the proportion of data within each error tolerance interval. After iterative verification, obtain meteorological quality characteristic data;

[0082] Step 6: Generate meteorological quality distribution characteristics for different regions and levels according to the meteorological quality characteristic data to obtain the time series difference between the meteorological evaluation manual station and automatic station data, and save the generated data and the obtained data to folders respectively. By matching the output formats set by different users, output the final meteorological quality assessment report.

[0083] Specifically: Data initialization completes the running environment configuration by automatically detecting and installing missing R packages (such as ggplot2, readxl);

[0084] Data preprocessing: Load data in xlsx format, verify the file integrity, and perform missing value processing and format conversion;

[0085] Statistical analysis engine: Global statistics: Calculate the extreme values, mean, standard deviation, and coefficient of variation of the complete data set;

[0086] Stratified analysis: Divide the data according to regions and depth levels, and calculate the proportion of absolute error and tolerance interval (±5% / 10%);

[0087] Time series comparison: Use the monitoring date as the abscissa to generate a line chart of the difference in soil moisture between the manual / automatic stations;

[0088] Visualization: Display the data distribution density through a violin plot and present the time series difference through a line chart;

[0089] Report generation: Automatically generate text / Word format reports based on templates, including statistical tables and chart references.

[0090] The above operations execute the following algorithm:

[0091] Automated chart generation algorithm: Dynamically bind data fields based on the ggplot2 package and batch output area stratification charts. Process control: Users only need to set the input path, statistical parameters, and output format, and the main function (main.R) automatically schedules the execution of each module;

[0092] The output results include: statistical tables (output_table.xlsx), charts (plots folder), and analysis reports (.txt / .docx).

[0093] Example 1

[0094] This example provides a complete automated quality assessment method for meteorological observation data. This method is implemented in R language and includes six core steps: environment adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis, error tolerance interval analysis, automated visualization generation, and intelligent report generation.

[0095] 1.1 The automated quality assessment method for meteorological observation data in this example adopts a modular design. This architecture mainly includes 5 functional modules, corresponding to 5 core steps respectively: data initialization module, data preprocessing module, statistical analysis module, error analysis module, and report generation module. Information transfer and functional cooperation are achieved through data streams between modules, forming a complete data processing chain.

[0096] 1.2 The environment adaptive initialization step first detects whether the required software packages have been installed in the current R environment. The main software packages that the system needs to detect include: dplyr, tidyr, data.table related to data processing, readxl, openxlsx related to data reading, ggplot2, plotly related to visualization, rmarkdown, knitr related to report generation, etc.

[0097] By creating a list of required software packages and then checking each package one by one to see if it is installed, for packages that are not installed, the system will automatically perform the installation operation. At the same time, the system will create an initial state system, including four status codes: correct (0), suspicious (1), error (2), and missing measurement (3). Then, quality assessment indicators will be constructed to calculate indicators such as the actual existence rate, correct rate, suspicious rate, and error rate. Finally, a computing resource allocation strategy will be configured to dynamically adjust resource usage according to the system memory and the number of CPU cores. Through automated environment configuration, the problem of manually installing dependency packages item by item in the traditional method is solved, greatly improving work efficiency; a standardized initial state system is established, providing a unified standard for subsequent data quality assessment; the dynamic resource allocation strategy improves the efficiency of large-scale data processing.

[0098] 1.3 The intelligent data preprocessing step receives the output configuration of the environment adaptive initialization in Embodiment 1.2, performs path legality verification and file accessibility detection, reads meteorological observation data in Excel format, and performs format standardization conversion and missing value processing on the data.

[0099] First, the data path parameters set by the user are verified. Check whether the file exists, is readable, and whether the file format is supported. When the file does not exist, the system will record the error information and prompt the user to check the path settings, and at the same time try to search for files with the same name in the adjacent directory; when the file is not readable, the system will check the file permissions and try to modify the access permissions. If it is still unable to access, the user will be prompted to contact the system administrator; when the file format is not supported, the system will provide format conversion suggestions and try to use the general text reading method to parse the data.

[0100] Then, an appropriate loading method is automatically selected according to the file format, supporting multiple formats such as xlsx, xls, csv, etc. For files with unknown formats, the system will try to read them in text format and perform intelligent format recognition. Then, the loaded data is checked for format to ensure that it contains necessary fields such as station number, observation time, temperature, humidity, air pressure, wind speed, etc. When necessary fields are missing, the system will generate field mapping suggestions and allow the user to manually specify the field correspondence. Subsequently, format standardization conversion is performed, converting the observation time to the date and time type and converting meteorological elements to numerical types. For data that fails to convert, the system will record the exception information and mark it as an error status. Finally, missing value detection and processing are performed. For missing values in the observation time, the records are directly deleted. For missing values in meteorological elements, time series interpolation methods are used to fill them, and the interpolated data is marked as suspicious using the quality control code. When the interpolation fails, the system will maintain the original missing measurement state and record the processing result in the log.

[0101] Through the path verification mechanism and exception handling, the problem of program interruption caused by missing files in traditional methods is solved; intelligent format conversion and missing value handling improve the automation of data processing; four types of quality inspections (climate threshold value inspection, internal consistency inspection, time consistency inspection, and spatial consistency inspection) effectively identify abnormal data and improve data quality.

[0102] 1.4 Multi-dimensional statistical analysis steps Receive the preprocessed full meteorological assessment dataset in Embodiment 1.3 and execute a three-layer statistical analysis framework: First, perform global statistical analysis, then divide the data subset by region and depth levels, and finally perform hierarchical statistical analysis on each subset.

[0103] First, perform global statistical analysis, calculate basic statistical indicators such as extreme values, means, standard deviations, and coefficients of variation of each meteorological element, and establish a global meteorological quality assessment benchmark. Then, perform stratification in the time dimension, conduct grouped statistics by year, season, month, and day, and combine the quality assessment marking results to form a meteorological quality assessment subset at the time level. Next, perform stratification in the spatial dimension, conduct grouped statistics by station, and form a meteorological quality assessment subset at the spatial level. Subsequently, perform stratification in the element dimension, calculate the quality assessment indicators of each meteorological element respectively, including accuracy rate, error rate, suspicious rate, etc. Finally, calculate the EMI environmental meteorological condition assessment index, which represents the comprehensive index of processes such as aerosol emission, sedimentation, transport, and diffusion under the influence of meteorological conditions.

[0104] Through multi-level grouped statistical analysis, comprehensive analysis in the time dimension, spatial dimension, and element dimension is achieved; the introduction of the EMI environmental meteorological condition assessment index can quantitatively separate the influence of meteorological condition changes on data quality and provides a more comprehensive and accurate multi-dimensional comprehensive analysis ability.

[0105] 1.5 Error tolerance interval analysis steps Receive the hierarchical statistical analysis results in Embodiment 1.4, calculate the absolute error and relative error between the meteorological assessment manual station and automatic station data, and count the proportion of data within each error tolerance interval.

[0106] First, construct an error statistical function that receives the meteorological assessment manual station observation data and automatic station observation data as input parameters. Then, calculate the absolute error value and relative error percentage between the two groups of meteorological observation data in combination with the hierarchical statistical analysis results at the element level. Next, use the meteorological quality assessment subsets at the time and spatial levels to count the proportion of data with a relative error within the ±5% threshold range and an absolute error value within the ±10% threshold range. Subsequently, calculate the standard deviation based on the EMI environmental meteorological condition assessment index and analyze the error distribution characteristics under different EMI conditions. Finally, through an iterative verification mechanism, verify and optimize the error statistical results to generate an error interval ratio distribution table.

[0107] By setting the tolerance intervals of ±5% relative error threshold and ±10% absolute error threshold, combined with the iterative verification mechanism, the precise quantification of the data differences between the manual and automatic meteorological assessment stations is achieved; the error analysis based on the EMI index can identify the influence degree of meteorological conditions on data quality, improving the accuracy of error quantification and quality assessment.

[0108] 1.6 Automated Visualization Generation Step The error statistical results in Example 1.5 and the statistical analysis results in Example 1.4 are received, and various types of visualization charts are automatically generated by dynamically binding data fields through the ggplot2 package.

[0109] First, based on the feature distribution characteristic data and the error interval ratio distribution table, the probability density distribution characteristic charts of meteorological elements in different regions and levels are generated by dynamically binding data fields through the ggplot2 package, and the data quality grades are marked in combination with the meteorological quality identification system. Then, using the probability density distribution characteristics and the time - dimension stratified data, combined with the error statistical results, a time - series difference chart is generated to show the time - varying trends of the data between the manual and automatic meteorological assessment stations, and the data points within different error tolerance intervals are distinguished by error coding. Next, based on the data characteristics obtained from the time - series feature analysis and the space - dimension stratified data, combined with the error interval ratio distribution, a geospatial heat effect chart is generated to show the spatial distribution characteristics of the meteorological quality assessment indicators, and the error mapping range is dynamically adjusted according to the EMI environmental meteorological condition assessment index. Finally, the generated visualization charts are saved to the corresponding folders according to the analysis dimensions respectively.

[0110] Based on the ggplot2 package, data visualization is realized, and standardized charts including probability density distribution, time - series feature analysis, and spatial distribution characteristics are automatically generated; through visual elements such as color, size, and shape, the data distribution, change trends, and spatial characteristics are intuitively displayed, solving the technical problem of inability to automatically generate standardized visualizations.

[0111] 1.7 Intelligent Report Generation Step The visualization charts in Example 1.6 and the error analysis results in Example 1.5 are integrated, and a complete meteorological quality assessment report is generated according to the output format set by the user.

[0112] First, according to the evaluation purpose and data complexity, select a suitable template from the preset report template library, which includes combinations of three purposes (briefing type, standard type, professional type) and three levels of complexity (simple, medium, complex). Then, based on the error analysis results, automatically extract key quality issues, judge the quality status of each element, and generate a report summary text, including the overall evaluation result, the evaluation results of each element, and improvement suggestions. Next, based on the error analysis results and quality evaluation indicators, generate a standardized statistical result table, including an error statistics table, an error interval distribution table, and an error statistics table under EMI conditions. Subsequently, based on the error statistics results and quality evaluation indicators, provide maintenance suggestions and optimization plans for meteorological observation stations, and judge the maintenance priorities according to the error conditions of each station. Finally, integrate the report summary, statistical tables, visualization charts, and maintenance suggestions into the selected template, and use the rmarkdown package to render and generate the final report according to the output format (Word, HTML, or PDF) set by the user.

[0113] By automatically extracting key quality issues and outliers and putting forward targeted improvement suggestions, the technical problem of being unable to automatically generate a standardized report is solved; it supports multiple output formats to meet the needs of different users; and based on the error statistics results, it provides targeted maintenance suggestions for stations, providing a scientific basis for meteorological data quality assessment.

[0114] The six steps of this embodiment form a complete data processing chain: the environment adaptation initialization in Embodiment 1.2 provides the basic environment and configuration for all subsequent steps; the intelligent data preprocessing in Embodiment 1.3 receives the environment configuration and outputs preprocessed data for statistical analysis; the multi-dimensional statistical analysis in Embodiment 1.4 receives the preprocessed data and outputs statistical results for error analysis and visualization; the error tolerance interval analysis in Embodiment 1.5 receives the statistical results and outputs error analysis results for visualization and report generation; the automated visualization generation in Embodiment 1.6 receives the statistical and error analysis results and outputs charts for report generation; the intelligent report generation in Embodiment 1.7 integrates the outputs of all previous steps and generates the final evaluation report.

[0115] The whole process from environment configuration to report generation is automated without manual intervention, greatly improving work efficiency; through the exception capture and automatic recovery mechanism, it ensures that the system can operate stably in the face of various abnormal situations; it can dynamically adjust the processing strategy according to different data characteristics and system resources to adapt to various operating environments; through hierarchical statistical analysis in three dimensions of time, space, and elements, it provides a comprehensive perspective for data quality assessment; it automatically generates various types of visualization charts to intuitively display the data quality distribution characteristics; through the computing resource allocation strategy, it optimizes memory usage and thread allocation to improve the efficiency of large-scale data processing.

[0116] Embodiment 2

[0117] This embodiment describes in detail the specific implementation of the environmental adaptation initialization step in the method for automated quality assessment of meteorological observation data. This step is the foundation of the entire method, providing the necessary software environment and configuration parameters for subsequent data processing, and forms a continuation relationship with Example 1.

[0118] 2.1 Environment Adaptive Initialization First, check whether the required software packages are installed in the current R environment. This embodiment adopts an intelligent detection mechanism, which not only checks whether the software packages are installed, but also checks whether the versions meet the requirements.

[0119] First, define a list of required packages and their minimum version requirements. This includes data processing packages such as dplyr, tidyr, and data.table; data reading packages such as readxl and openxlsx; visualization packages such as ggplot2 and plotly; and report generation packages such as rmarkdown and knitr. Then, check each package to see if it is installed. If not, automatically install it. For installed packages, check whether their version meets the minimum requirements. If not, automatically update them. Next, load all required packages and record the installation and loading results. Finally, check to see if all packages have been successfully installed. If the installation fails, give the corresponding error message and return to the previous step to repeat the automatic installation process by updating the installation path.

[0120] 2.2 Environment Adaptation Initialization also includes creating and defining quality assessment codes to mark the quality status of meteorological observation data. This embodiment defines four basic quality assessment codes and encapsulates them in a structured object.

[0121] First, the basic quality assessment codes are defined, including correct (0), suspicious (1), wrong (2), and missing (3). Then, the quality assessment status is defined, including excellent (only correct data), good (correct data and a small amount of suspicious data), general (correct, suspicious, and a small amount of wrong data), and poor (all types of data, including a large amount of missing and wrong data). Next, the quality assessment thresholds are set, including the accuracy threshold for good quality (0.9), the accuracy threshold for general quality (0.7), and the accuracy threshold for poor quality (0.5). Finally, the quality assessment code, assessment status, and control threshold are encapsulated into a complete quality assessment system object. A quality assessment code system with 0-3 assessment thresholds and four-level quality assessment status (excellent, good, general, and poor) is established, providing a unified standard for subsequent data quality assessment; the structured quality assessment system is easy to expand and maintain.

[0122] 2.3 Environment adaptive initialization also includes constructing a quality assessment index calculation function for calculating the quality assessment indexes of meteorological observation data. In this embodiment, four basic quality assessment indexes are defined: actual existence rate, correct rate, error rate, and suspicious rate.

[0123] First, create a quality assessment index calculation function that receives a data set containing quality assessment fields as input. Then, verify the input parameters to ensure that the specified quality assessment fields exist in the data. Next, calculate basic statistics, including the number of groups of data to be observed, the number of groups of actual data, the number of groups of correct data, the number of groups of suspicious data, the number of groups of error data, and the number of groups of missing data. Subsequently, calculate the quality assessment indexes, including the actual existence rate (the number of groups of actual data / the number of groups of data to be observed), the correct rate (the number of groups of correct data / the number of groups of actual data), the suspicious rate (the number of groups of suspicious data / the number of groups of actual data), the error rate (the number of groups of error data / the number of groups of actual data), and the missing rate (the number of groups of missing data / the number of groups of data to be observed). Finally, determine the quality assessment status of the data set based on the correct rate and return the complete calculation results. A standardized quality assessment index calculation method is provided to ensure the consistency and comparability of the assessment results; automated quality status determination reduces the problem of error-prone manual operations; functional design facilitates reuse in subsequent steps.

[0124] The last step of 2.4 environment adaptive initialization is to configure the computing resource allocation strategy, which dynamically adjusts the memory usage and thread allocation according to the data scale and system resources.

[0125] First, obtain the system resource information, including the available memory size and the number of CPU cores, and use corresponding methods to obtain the resource information for different operating systems. Then, dynamically adjust the resource allocation strategy according to the data scale and system resources, set the memory limit to 80% of the available memory, and set the number of threads to the number of CPU cores minus 1 (reserving one core for the system). Next, set the size of data block processing for batch processing of large-scale data. Subsequently, configure the parallel computing environment. If the system supports multi-threading, create a parallel computing cluster and export the necessary variables and functions. Finally, encapsulate all the resource configuration information into a resource policy object. By dynamically adjusting the computing resource allocation strategy of memory usage and thread allocation and performing adaptive optimization according to the data scale and system resources, resource optimization and performance improvement are achieved; the configuration of the parallel computing environment improves the efficiency of large-scale meteorological data processing.

[0126] First, perform the detection and installation of R software packages to ensure that all necessary software packages are correctly installed and loaded. Then, construct the initial state system to create a standardized quality assessment framework. Next, construct the quality assessment index calculation function to provide a unified quality assessment method. Subsequently, configure the computing resource allocation strategy to optimize the system performance. Finally, integrate all the initialization results into an environment configuration object, including the software package status, quality assessment system, quality assessment function, and resource strategy, for use in subsequent steps.

[0127] By integrating four sub-steps, a complete environment adaptive initialization process is formed, providing a unified basic environment for all subsequent steps; the modular design facilitates maintenance and expansion; the adaptive mechanism can adapt to various operating environments according to different operating systems and hardware configurations. Finally, the above four steps are integrated into a complete environment adaptive initialization function to provide a basic environment for the intelligent data preprocessing in Embodiment 3.

[0128] As the basis of the entire system, the environment configuration object output by this embodiment will be directly used by the intelligent data preprocessing in Embodiment 3, providing necessary tools and standards for data loading, format conversion, and quality inspection. At the same time, the initial state system established in this embodiment will run through the multi-dimensional statistical analysis in Embodiment 4, the error tolerance interval analysis in Embodiment 5, and even the intelligent report generation in Embodiment 7, ensuring the consistency of the entire evaluation process. Thus, it can automatically detect and install the required R software packages without manual intervention; ensure the normal operation of the system; create a complete initial state system and quality assessment index calculation function; dynamically adjust the memory usage and thread allocation according to system resources to improve data processing efficiency; can adapt to various operating environments according to different operating systems and hardware configurations; both the initial state system and the quality assessment index calculation function are designed to be extensible, and new control codes and assessment indicators can be added as needed.

[0129] Embodiment 3

[0130] This embodiment details the specific implementation of the intelligent data preprocessing step in the meteorological observation data automated quality assessment method. This step receives the output of the environment adaptive initialization configuration in Embodiment 2 and performs operations such as path verification, data loading, format conversion, and missing value processing to provide a standardized data basis for the multi-dimensional statistical analysis in Embodiment 4.

[0131] 3.1 Intelligent data preprocessing First, verify the data path parameters set by the user to ensure that the data file exists and is accessible. At the same time, adopt the try-catch mechanism for exception capture and automatic recovery to improve the robustness of the system.

[0132] First, the path is maintained in a valid string format to ensure the correctness of the input parameters. Then, it is checked whether the file at the specified path exists to avoid file non-existence errors in subsequent processing. Next, an attempt is made to open a file connection to verify that the file is readable and ensure normal file access permissions. Subsequently, the file extension is checked to verify the supported file format types (xlsx, xls, csv) to ensure that the data can be correctly loaded later.

[0133] When the file does not exist or is inaccessible, the system will automatically initiate a response strategy: First, attempt to search for a file with the same name in the current directory and common data directories. If found, the path will be automatically updated. Second, provide a user interaction interface that allows the user to reselect the correct file path. Third, check the connection status of the network path and attempt to re-establish a connection for data files stored on the network. Finally, record detailed error information and suggested solutions to help the user quickly locate and resolve path issues. Finally, encapsulate the verification result into a path information object, including the file path, format type, and verification status, for use in subsequent data loading steps.

[0134] 3.2 After the data path verification passes, the next step is to load the data and detect its format. This embodiment supports multiple data formats (xlsx, xls, csv) and can automatically select an appropriate loading method based on the file format.

[0135] First, based on the file format verified in Embodiment 3.1, automatically select the corresponding data loading method. Use the readxl package for Excel files and the read.csv function for CSV files. Then, after loading the data, verify that the data is not an empty package and contains valid records. Next, check that the data format contains the necessary columns (station number, observation time, temperature, humidity, air pressure, wind speed) to ensure that the data structure meets the requirements of subsequent processing. Subsequently, attempt to handle the case of mismatched column names by automatically identifying and converting common column name variants through an intelligent matching mechanism. If the main loading method fails, the system will attempt to reload using an alternative method to increase the success rate of data loading. Finally, return the successfully loaded original data for use in the format standardization step. The intelligent data loading function can automatically select an appropriate loading method based on the file format; automatic column name matching reduces loading failures caused by data format differences; the alternative loading mechanism increases the success rate of data loading and the fault tolerance of the system.

[0136] 3.3 After the data is successfully loaded, it is necessary to perform format standardization conversion on the data to ensure that the data types of each field are correct and provide a standardized data format for the statistical analysis in Embodiment 4.

[0137] First, convert the observation time field to the standard date and time type. The system will automatically attempt multiple common date format patterns, including formats such as year-month-day hour:minute:second and year / month / day hour:minute:second, to ensure correct recognition and conversion of various date and time formats. Then, convert the meteorological element fields (temperature, humidity, air pressure, wind speed) to numerical types to ensure the correctness of subsequent numerical calculations. Next, convert the station number to a string type to maintain the consistency of the station identifier. For exceptions that occur during the conversion process, the system will record warning messages and attempt to retain the original values to avoid data loss due to format conversion failures. Finally, return the dataset with standardized formats, where all fields have the correct data types. Automated format conversion reduces the workload of manual processing; automatic recognition of multiple date formats improves the success rate of conversion; standardized data formats provide a reliable data foundation for subsequent statistical analysis.

[0138] After the data format is standardized, it is necessary to detect and process missing values to ensure data integrity and provide high-quality data for the statistical analysis in Example 4.

[0139] First, detect the number of missing values in each field, count the number and missing rate of missing values in each field, and generate a missing value report. Then, process the missing values in the observation time. Since the observation time is a key time identifier and cannot be reasonably supplemented, directly delete the records containing missing values in the observation time. Next, process the missing values in the meteorological elements. Create corresponding quality assessment fields for each meteorological element and use the initial state system established in Example 2 for marking. Subsequently, group by station and perform time series interpolation on the meteorological element data for each station. Use the linear interpolation method to fill in the missing values and mark the interpolated data as a suspicious state. For missing values that cannot be interpolated, maintain the missing measurement state mark. Finally, return the processed dataset and the missing value report to provide a basis for subsequent quality inspections. The time series interpolation method can reasonably fill in the missing values of meteorological elements; the initial state marking ensures the traceability of the interpolated data; the missing value report provides an important reference for data quality assessment. Intelligent data preprocessing also includes four quality inspections on meteorological data: climate limit value inspection, internal consistency inspection, time consistency inspection, and spatial consistency inspection. These inspections can identify abnormal data and improve data quality.

[0140] First, perform climate threshold value checks, set the reasonable ranges for each meteorological element (temperature: -40°C to 50°C, humidity: 0% to 100%, air pressure: 850 hPa to 1050 hPa, wind speed: 0 m / s to 60 m / s), mark the data exceeding the threshold values as errors, and mark the data close to the threshold values as suspicious. Then, perform internal consistency checks to verify the physical relationships between different meteorological elements. For example, when the temperature is below 0°C, the humidity should not be 100%, and mark the data that does not conform to the physical laws as suspicious. Next, perform time consistency checks, calculate the change amplitudes of meteorological elements at adjacent times, set reasonable change thresholds (temperature: 5°C, humidity: 20%, air pressure: 10 hPa, wind speed: 10 m / s), and mark the data with abnormal changes as suspicious. Subsequently, perform spatial consistency checks, compare the meteorological elements at different stations at the same time, calculate the mean value and standard deviation, and mark the data deviating by 3 times the standard deviation as suspicious. Finally, return the dataset that has undergone the four quality checks, and all abnormal data have been appropriately marked.

[0141] Through multiple rounds of marking and multi-dimensional consistency checks, the problem of error-prone manual operations has been effectively reduced, and the accuracy and reliability of data processing have been enhanced; the four quality checks cover the main aspects of meteorological data quality assessment and provide comprehensive data quality assurance. Finally, integrate the above five steps into a complete data preprocessing function to provide high-quality preprocessed data for the multi-dimensional statistical analysis in Example 4.

[0142] First, perform data path verification and exception handling to ensure that the data files can be accessed normally. Then, perform data loading and format detection to obtain the original data and verify the data structure. Next, perform standardized conversion of data formats to ensure that all fields have the correct data types. Subsequently, perform missing value detection and handling to improve the integrity of the data. Finally, perform the four quality checks to identify and mark abnormal data. The entire process adopts an exception capture mechanism to ensure that clear error messages can be given when problems occur. Finally, return the complete preprocessing results, including the processed dataset, missing value report, and processing log.

[0143] By integrating the five sub-steps, a complete intelligent data preprocessing process is formed; the exception capture and automatic recovery mechanism ensure the stability of the system; the high-quality preprocessed data lays a solid foundation for subsequent statistical analysis.

[0144] This embodiment receives the environmental configuration object of Embodiment 2 as input and uses the initial state system and calculation functions therein for data processing. The output of this embodiment will be directly passed to the multi-dimensional statistical analysis of Embodiment 4, providing a standardized data basis for statistical calculations. At the same time, the quality assessment fields established in this embodiment will continue to be used in the error analysis of Embodiment 5 and the visualization generation of Embodiment 6. By automatically identifying file formats, column name matching, and format conversion, manual intervention is reduced; by exception capture and alternative methods, the success rate of data loading is increased; through four types of quality checks, abnormal data is comprehensively identified and marked; scientific methods such as time series interpolation are used to process missing values; unified data formats and quality markings provide a standard basis for subsequent analysis; detailed processing logs and quality markings ensure the traceability of the data processing process.

[0145] Embodiment 4

[0146] This embodiment details the implementation of the multi-dimensional statistical analysis step in the automated quality assessment method for meteorological observation data. This step receives the output of the intelligent data preprocessing in Embodiment 3, executes a three-layer statistical analysis architecture, and provides a comprehensive statistical basis for the error tolerance interval analysis of Embodiment 5 and the automated visualization generation of Embodiment 6.

[0147] 4.1 Multi-dimensional statistical analysis First, perform global statistical analysis on the preprocessed full meteorological assessment dataset to establish a global meteorological quality assessment benchmark.

[0148] First, calculate basic statistical indicators for all meteorological elements (temperature, humidity, air pressure, wind speed), including minimum value, maximum value, average value, standard deviation, and coefficient of variation. These indicators constitute the global meteorological quality assessment benchmark. Then, set the meteorological micro-interference filtering threshold and dynamically adaptively adjust it based on the meteorological region characteristics combined with the preliminary quality assessment results of Embodiment 3 to optimize the observation quality of each data item in the full meteorological assessment dataset. Next, based on the global meteorological quality assessment benchmark, provide a reference standard for subsequent hierarchical statistical analysis. Finally, encapsulate the global statistical results into a global statistical object, including the statistical indicators and quality benchmarks of each element, for hierarchical analysis. Through global statistical analysis, a unified quality assessment benchmark is established; the dynamic adaptive adjustment mechanism can optimize the assessment criteria according to regional characteristics; and a scientific reference basis is provided for subsequent hierarchical analysis.

[0149] 4.2 Based on the global statistical analysis, perform multi-level grouped statistical analysis on meteorological data according to the time dimension to obtain the hierarchical statistical analysis strategy at the time level.

[0150] First, stratify each data item in the global statistical analysis dataset by time dimension, and conduct grouped statistics by year, season, month, and day. The system will automatically extract time elements such as year, season, month, and date from the observation time field. Among them, the seasons are divided as follows: spring from March to May, summer from June to August, autumn from September to November, and winter from December to February. Then, combine the quality assessment marking results established in Example 3 to form a meteorological quality assessment subset at the time level and generate time-series basic data. Next, calculate the average values of meteorological elements at each time level and analyze the variation characteristics of meteorological elements at different time scales. Finally, integrate the stratification results of the time dimension to provide a data basis for the time-series feature analysis in Example 6. Through the stratified statistics in the time dimension, the variation laws of meteorological elements at different time scales are revealed; the time-series basic data provides important support for subsequent time-series analysis; the multi-level time grouping meets the analysis requirements of different time scales.

[0151] 4.3 Stratify and statistically analyze meteorological data according to the spatial dimension to obtain the stratified statistical analysis strategy at the spatial level.

[0152] First, stratify each data item in the global statistical analysis dataset by spatial dimension, and conduct grouped statistics by region, geographical features, and elevation. The system will group the data according to the station numbers and calculate the statistical indicators of meteorological elements for each station. Then, use the regional characteristic adjustment results in Example 3 to form a meteorological quality assessment subset at the spatial level and generate spatial distribution data. Next, analyze the differences in meteorological elements between different stations, identify spatial distribution characteristics and abnormal stations. Finally, integrate the stratification results of the spatial dimension to provide a data basis for the spatial feature analysis in Example 6. Through the stratified statistics in the spatial dimension, the spatial distribution characteristics of meteorological elements are revealed; the spatial distribution data provides an important basis for subsequent geospatial analysis; the identification of abnormal stations helps to detect equipment failures or environmental problems.

[0153] 4.4 Stratify and statistically analyze meteorological data according to the element dimension to obtain the stratified statistical analysis strategy at the element level.

[0154] First, stratify each data item in the global statistical analysis dataset by factor dimension and conduct grouped statistics according to meteorological factors such as temperature, humidity, air pressure, and wind speed. The system will automatically identify the data characteristics of each meteorological factor and calculate the quality assessment indicators for each factor, including the correct rate, error rate, suspicious rate, and missing measurement rate. Then, combine with the quality assessment marking results established in Example 3 to form a meteorological quality assessment subset at the factor level and generate factor characteristic data. Next, analyze the quality differences between different factors and identify the factor types with prominent quality problems. Finally, integrate the stratification results of the factor dimension to provide a quality basis at the factor level for the error tolerance interval analysis in Example 5. Through the stratified statistics of the factor dimension, the quality characteristic differences of different meteorological factors are revealed; the factor characteristic data provides an important basis for subsequent factor-level analysis; the identification of quality problems helps to improve the observation quality of specific factors targeted.

[0155] 4.5 Based on the results of multi-dimensional statistical analysis, calculate the environmental meteorological condition assessment index to provide a comprehensive index for meteorological data quality assessment.

[0156] First, extract the key meteorological parameters from the stratified statistical results of the three dimensions of time, space, and factor, including the statistical characteristic values of temperature, humidity, air pressure, and wind speed. The system will assign weight coefficients to each meteorological factor according to meteorological principles and actual application requirements. Among them, the temperature weight is 0.4, the humidity weight is 0.3, the wind speed weight is -0.3 (the negative value indicates that the comfort level decreases when the wind speed increases), and the air pressure is used as an adjustment factor. Then, calculate the EMI index by grouping according to the station and time. The formula is:

[0157] EMI = average temperature × 0.4 + average humidity × 0.3 + average wind speed × (-0.3);

[0158] Next, classify the EMI index into four grades: excellent, good, general, and poor, to provide a basis for environmental condition classification for subsequent analysis. Finally, associate the EMI calculation results with the quality assessment results in Example 3 to analyze the data quality characteristics under different environmental conditions.

[0159] Through EMI index calculation, the correlation between environmental conditions and data quality is established; the classification results provide an important basis for subsequent conditional analysis; the comprehensive index simplifies the evaluation process of complex environmental conditions. Through stratified statistics in three dimensions of time, space, and elements, a comprehensive perspective on data quality assessment is provided; standard statistical indicators and scientific calculation methods are used to ensure the reliability and comparability of the analysis results; the correlation between environmental conditions and data quality is established through the EMI index, laying a foundation for conditional analysis; the stratified statistical results are clear at different levels, facilitating the identification and analysis of quality problems at different levels; it provides a comprehensive statistical basis for the error tolerance interval analysis in Example 5 and the automated visualization generation in Example 6; the multi-dimensional statistical results provide a scientific basis and data support for quality management decisions.

[0160] In another preferred embodiment, various meteorological observation data (such as temperature, humidity, wind speed, etc.) are marked according to preset criteria (such as data integrity, rationality, volatility, etc.); this marking is carried out by comparing the historical normal value range and the standard deviation offset method; the form of the marking is a quality assessment code, usually a numerical or category label, indicating the "quality" level of each piece of data. Thus, abnormal or unreliable data in the observed data can be quickly identified, improving the credibility of data use.

[0161] A digital range is set, corresponding to four quality assessment states of excellent, good, average, and poor; the quality assessment codes are divided into segments, for example: 90 - 100: excellent; 70 - 89: good; 50 - 69: average; 0 - 49: poor; each range corresponds to a clear quality status label, facilitating intuitive classification and management. This makes the originally abstract quality assessment code into an interpretable and easy-to-use grade classification; it improves the readability and practicality of the assessment results, facilitating user understanding and operation.

[0162] A meteorological quality identification system is constructed; based on the above digital range and status labels, a complete set of meteorological data quality identification systems is established; a corresponding quality grade label is added. Thus, a standardized and unified quality evaluation system is constructed; it facilitates operations such as automatic classification, screening, alarming, and data cleaning in the system. Continuously monitor the quality status of meteorological data at different times and construct a time series. Analyze the changing trends of these states over time, such as continuous deterioration, fluctuations, or obvious improvement. Thus, the dynamic quality of the data can be grasped in real time, which helps to detect sensor failures or environmental mutations in a timely manner. Provide decision support, such as the need to calibrate equipment or exclude outliers.

[0163] Import data and status trends into the R language environment and use loaded statistical analysis packages (such as ggplot2, dplyr, forecast, and changepoint) to conduct multi-dimensional analysis (e.g., time, space, and variables). Compare current data quality status with historical status to determine whether the trend is "smooth" or "dramatic." Based on the magnitude of the trend, determine the appropriate statistical strategy, such as weighted average, hierarchical clustering, or sliding window analysis. A highly automated and configurable analysis mechanism adapts to data variation characteristics in different scenarios, accurately identifying abnormal changes and formulating appropriate response strategies, such as focused monitoring, data supplementation, or equipment maintenance.

[0164] Example 5

[0165] This example describes in detail the implementation of the error tolerance analysis step in the automated quality assessment method for meteorological observation data. This step receives the output of the multi-dimensional statistical analysis in Example 4, calculates the absolute and relative errors between the manual and automated observation data, and calculates the percentage of data within each error tolerance. This provides the basis for error analysis in the automated visualization generation in Example 6.

[0166] 5.1 Error tolerance analysis First, the observation data from the manual and automatic stations need to be matched in time and space to ensure the validity of the comparison.

[0167] First, the data of the manual station and the automatic station are time-synchronized and verified to ensure that the two sets of data are observed at the same time point. The system will automatically identify the time field format, convert it into a standard time format, and handle the time deviation problem. Then spatial position matching is performed to ensure that the data compared are from the same observation point based on the site number or geographic coordinates. The matched data is then quality-screened, and the data points marked as erroneous or missing in Example 3 are eliminated, and the correct and suspicious data are retained for error analysis. Finally, the data is standardized to ensure that the dimensions and accuracy of each meteorological element are consistent, providing a standardized data basis for subsequent error calculations. Through strict data matching and preprocessing, the scientific nature and accuracy of error analysis are ensured; the quality screening mechanism avoids the interference of abnormal data on error statistics; and the standardization process ensures the comparability of errors between different elements.

[0168] 5.2 Based on the matched data, calculate the absolute error of each meteorological element and evaluate the observation deviation of the automatic station relative to the manual station.

[0169] First, calculate the absolute error for each matched data pair using the formula:

[0170] Absolute error = |Automatic station observation value-manual station observation value|;

[0171] The system processes each meteorological element (temperature, humidity, air pressure, wind speed) one by one to ensure the integrity of the calculation. Then, statistical analysis is carried out on the absolute error, and statistical indicators such as the mean absolute error, standard deviation of absolute error, maximum absolute error, and minimum absolute error are calculated. Next, combined with the EMI environmental condition classification in Example 4, the absolute error characteristics under different environmental conditions are analyzed to identify the influence of environmental factors on the observation error. Finally, the absolute error is grouped and statistically analyzed by site and time dimensions to provide detailed error distribution data for subsequent interval analysis. Through the calculation of the absolute error, the observation difference between the automatic station and the manual station is quantified; the statistical analysis reveals the distribution characteristics and variation laws of the error; the environmental correlation analysis identifies the environmental factors affecting the observation accuracy.

[0172] 5.3 Based on the absolute error, calculate the relative error of each meteorological element to evaluate the relative degree of the observation deviation.

[0173] First, calculate the relative error for each matching data pair. The formula is:

[0174] Relative error = │Observed value of automatic station - Observed value of manual station│ / │Observed value of manual station│ × 100%;

[0175] The system will handle special cases where the denominator is zero, either by adopting an alternative calculation method or marking it as invalid data. Then, statistical analysis is carried out on the relative error, and statistical indicators such as the mean relative error and standard deviation of relative error are calculated. Next, set the tolerance intervals for the relative error, usually including multiple interval levels such as ±5%, ±10%, ±15%, ±20%, etc., and count the proportion of data within each interval. Finally, combined with the stratified statistical results in Example 4, analyze the relative error distribution characteristics in different time, space, and element dimensions to provide a multi-dimensional error perspective for quality assessment. Through the calculation of the relative error, a standardized error assessment index is provided; the interval statistics reveal the hierarchical distribution of data quality; the multi-dimensional analysis identifies the error variation laws under different conditions.

[0176] 5.4 Based on the calculation results of the absolute error and relative error, count the proportion of data within each error tolerance interval and establish a data quality grade assessment system.

[0177] First, set multiple levels of error tolerance intervals, including absolute error intervals (such as ±0.5, ±1.0, ±2.0, etc.) and relative error intervals (such as ±5%, ±10%, ±20%, etc.). The system will dynamically adjust the interval thresholds according to the observation accuracy requirements of different meteorological elements. Then, count the number and proportion of data points within each interval, calculate the cumulative distribution function, and analyze the distribution characteristics of the errors. Next, establish a data quality level assessment standard, defining data with a relative error within ±5% as excellent, ±5% to ±10% as good, ±10% to ±20% as average, and exceeding ±20% as poor. Finally, generate an error tolerance interval distribution table, providing a quantitative quality assessment basis for the intelligent report generation in Example 7. Through the error tolerance interval statistics, a quantitative data quality assessment standard is established; the level assessment system facilitates the intuitive understanding of the quality status; the distribution analysis provides targeted guidance for quality improvement.

[0178] 5.5 Combine the EMI environmental condition assessment results in Example 4, analyze the error characteristics under different environmental conditions, and identify the influence of environmental factors on the observation accuracy.

[0179] First, correlate the error analysis results with the EMI environmental condition classification results, and group the error data according to the EMI level (excellent, good, average, poor). The system will calculate the mean absolute error, mean relative error, and error tolerance interval distribution under each EMI level. Then, analyze the correlation between environmental conditions and observation errors, and identify the influence laws of factors such as extreme weather conditions, seasonal changes, and day-night differences on the observation accuracy. Next, establish an environmental condition-error characteristic correlation model, providing a reference standard for data quality assessment under different environmental conditions. Finally, generate an error analysis report under environmental conditions, providing an environmental factor analysis basis for the effectiveness assessment of the quality management system in Example 8.

[0180] Through the error analysis under environmental conditions, the influence mechanism of environmental factors on the observation accuracy is revealed; the correlation model provides a scientific basis for conditional quality assessment; the analysis results provide guidance for the improvement of the environmental adaptability of the observation system. Through the dual calculation of absolute error and relative error, the accuracy level of the observation system is comprehensively quantified; the multi-level tolerance intervals set based on the meteorological observation accuracy requirements provide a scientific quality assessment standard; the error analysis combined with EMI environmental conditions reveals the influence laws of environmental factors on the observation accuracy; the established data quality level assessment system facilitates the intuitive understanding of the quality status and management decision-making; analyzing the error characteristics from multiple dimensions such as time, space, elements, and environment provides an all-round quality assessment perspective; it provides a detailed error analysis basis for the automated visualization generation in Example 6 and the intelligent report generation in Example 7.

[0181] In another preferred embodiment, the abnormal state in the dynamic change of meteorological observation data is captured: The time series analysis method (such as sliding window, trend recognition, anomaly detection algorithm) is used to monitor the change trend of meteorological observation data at different times in real time; The points with a large deviation from the normal trend are identified, such as a sudden increase or decrease, which are defined as abnormal state data; Statistical indicators such as Z-score, IQR (Interquartile Range), and EWM (Exponentially Weighted Moving) can be used to determine the anomaly threshold. Thus, the sensitivity of the system to "sudden anomalies" or "potential fault signals" is improved; It is a prerequisite for further accurate quality assessment and correction.

[0182] Perform time series and spatial interpolation analysis on the abnormal data, extract the sudden change amplitude index, and mark it as the secondary quality code:

[0183] Time series characteristic analysis: Based on the continuous observation data before and after the abnormal point, calculate the change rate or deviation degree (such as the temperature drops suddenly by 8°C within one hour);

[0184] Spatial correlation analysis: Use the data of surrounding meteorological stations and estimate the expected value of this point through spatial interpolation algorithms (such as Kriging, IDW, inverse distance weighting);

[0185] Sudden change amplitude index: Integrate the time series mutation and spatial difference to generate a comprehensive index for measuring the anomaly severity;

[0186] Use this index to assign a new quality assessment code to the abnormal state data again (such as marking as "moderate anomaly", "severe anomaly", etc.) to obtain the secondary marking result. Anomaly detection is no longer a single-point judgment, but an evaluation from multiple spatio-temporal dimensions; Improve the accuracy and scientificity of abnormal data judgment; Provide a quantitative basis for subsequent inspection or correction.

[0187] Perform four consistency and extreme value checks to generate the tertiary marking result; Conduct quadruple verification checks on the abnormally data marked secondarily: Climate boundary value check: Compare with historical extreme values or long-term climate average ranges (for example, the winter temperature in Beijing should not be higher than 20°C);

[0188] Internal consistency check: Check the logical relationship between different meteorological elements (such as humidity should not be 0% while precipitation is heavy rain);

[0189] Time consistency check: The consistency between the current observed value and the data of adjacent time periods;

[0190] Spatial consistency check: The rationality comparison between the observed value of the current station and the observed values of adjacent stations.

[0191] Any data that violates the above inspection rules is marked with an updated quality assessment code to form three marking results. Thus, a multi-level and systematic error elimination mechanism can be established; the risks of "misjudging good data" or "missing bad data" can be significantly reduced; and the credibility and rigor of data processing can be enhanced.

[0192] Perform fitting processing on the observed data in a general or poor state: Process the observed data identified as "general" or "poor" in the three markings.

[0193] The fitting processing methods may include: interpolation filling (such as spatial interpolation, time series interpolation); curve fitting (such as polynomial regression, moving average); model re-estimation (such as modeling prediction based on surrounding stations). The goal is to restore or estimate a more likely true value to replace the original abnormal observed value. Reduce the analysis bias caused by data missing or anomalies; maintain data continuity for downstream model processing; and reduce the risk of the overall observation system being affected by local anomalies.

[0194] Integrate the processed data with the excellent data to form a preliminary quality assessment result for global analysis: Integrate the "data after fitting processing" in the previous step with the observed data originally marked as "excellent" or "good" in the first-round assessment; generate a set of "preliminary quality assessment data sets" through integration; this data set is then transmitted to the "global statistical analysis system" for comprehensive modeling analysis, such as trend modeling, extreme event identification, etc. Thus, the automatic construction of a high-quality data set can be achieved; ensure that the input of the downstream analysis model has the maximum integrity and minimum error; support key application scenarios such as meteorological prediction, abnormal event monitoring, and policy making. Summary of the overall process advantages:

[0195] Link Key objective Effective means Bring about effects Exception capture Dynamic monitoring of mutations Timing detection Detect anomalies in a timely manner Secondary marking Judge the severity of anomalies Spatial-temporal interpolation + amplitude index Anomaly intensity quantification Tertiary marking Data consistency verification Logic / historical rule verification Multi-layer screening to ensure accuracy Fitting correction Anomaly data recovery Model fitting / estimation Maintain data continuity Result integration Construct a cleaned dataset Screening + fusion For use by the global analysis system

[0196] Example 6

[0197] This example details the implementation of the automated visualization generation step in the automated quality assessment method for meteorological observation data. This step receives the output of the error tolerance interval analysis in Example 5 and automatically generates various types of visualization charts through dynamic data binding technology, providing an intuitive chart display for the intelligent report generation in Example 7.

[0198] 6.1 Automated visualization generation First, a dynamic binding mechanism between data fields and chart elements needs to be established to achieve the automated generation of charts.

[0199] First, perform data structure parsing on the error analysis results of Example 5 to automatically identify the types, ranges, and distribution characteristics of data fields. The system will establish a field mapping table to map meteorological element fields to visual elements such as the X-axis, Y-axis, color, and size of the chart. Then, automatically select the appropriate chart type according to the data characteristics. For continuous numerical data, scatter plots or line charts are preferred; for categorical data, bar charts or pie charts are preferred; for distribution data, box plots or violin plots are preferred. Next, establish a dynamic binding function to automatically configure chart parameters and style settings according to the meteorological elements and visualization requirements specified by the user. Finally, implement a mechanism for batch generation of charts to generate corresponding visualization charts for multiple meteorological elements in parallel, improving the generation efficiency. Through the dynamic binding mechanism, the automation and standardization of chart generation are achieved; intelligent type selection ensures the best match between the chart and data characteristics; the batch generation mechanism significantly improves the generation efficiency of visualization.

[0200] 6.2 Based on the dynamic binding framework, generate a violin plot to display the distribution characteristics and density information of meteorological element data.

[0201] First, extract the observed data of each meteorological element from the error analysis results and group them by station or time. The system will calculate the probability density distribution of each group of data and use the kernel density estimation method to smooth the data distribution curve. Then, construct the graphic elements of the violin plot, including the symmetric mirror image of the density curve, quartile markers, median line, and outlier points. Next, set the visual style of the chart, including color scheme, transparency, line thickness, etc., to ensure the aesthetics and readability of the chart. Finally, add auxiliary information such as chart title, axis labels, and legend to generate a complete violin plot, intuitively displaying the distribution shape, central tendency, and dispersion degree of the data. Through the generation of the violin plot, the distribution characteristics and density information of the data are intuitively displayed; the kernel density estimation method provides a smooth distribution curve; the reasonable configuration of visual elements enhances the readability and aesthetics of the chart.

[0202] 6.3 Generate a box plot to display the statistical characteristics and outlier distribution of meteorological element data.

[0203] First, calculate the five - number summary statistics of each meteorological element data, including the minimum value, the first quartile, the median, the third quartile, and the maximum value. The system will identify outliers and use the 1.5 - times inter - quartile range rule to determine the decision threshold for outliers. Then, construct the graphical elements of the box plot, including the box (representing the inter - quartile range), the median line, the whiskers (representing the normal value range), and the outlier points. Next, generate multiple box plots according to different grouping dimensions (such as stations, time, environmental conditions) for easy comparison and analysis. Finally, set the color coding and annotation information of the chart, use different colors to distinguish different groups, and add numerical labels to provide accurate statistical information. Through the generation of box plots, the statistical characteristics and outlier distribution of the data are clearly shown; the five - number summary provides key information about the data distribution; the grouping comparison function facilitates the identification of data differences under different conditions.

[0204] 6.4 Generate a time - series difference line chart to show the changing trends and difference characteristics of the observed data of manual stations and automatic stations over time.

[0205] First, extract the time - series error data from the error analysis results of Example 5 and arrange the observed differences between manual stations and automatic stations in chronological order. The system will process the missing values and outliers in the time series and use interpolation or smoothing methods to ensure the continuity of the time series. Then, construct a dual - axis line chart, with the main axis showing the time change of the observed differences and the secondary axis showing the percentage change of the relative error. Next, add trend lines and confidence intervals, use regression analysis methods to fit the long - term trend, and calculate the confidence intervals to reflect the uncertainty of the data. Finally, set the scale and labels of the time axis, automatically select an appropriate time interval according to the time span of the data, and add annotations for important time nodes. Through the generation of the time - series difference line chart, the time - changing pattern of the observed errors is intuitively shown; the dual - axis design shows both the absolute difference and the relative error at the same time; the trend analysis reveals the long - term changing trend and periodic characteristics of the errors.

[0206] 6.5 Based on the site geographical coordinate information, generate a geospatial heat map to show the spatial distribution characteristics of meteorological element errors.

[0207] First, obtain the geographical coordinate information of each observation station, including longitude, latitude, and altitude. The system will associate the error analysis results of Example 5 with the station coordinates and calculate the mean absolute error of each station as the numerical basis for the heat map. Then, construct a geographic base map, select an appropriate map projection and scale according to the geographical scope of the observation area, and add geographical elements such as administrative boundaries, rivers, and terrain. Next, apply a spatial interpolation algorithm, using Kriging interpolation or inverse distance weighting method to interpolate the discrete station error data into a continuous spatial distribution surface. Finally, set the color mapping scheme of the heat map, use gradient colors to represent the magnitude of the error, add a color legend and contour lines to generate a complete geospatial heat map. Through the generation of the geospatial heat map, the spatial distribution pattern of the observation error is visually displayed; the spatial interpolation technology provides a continuous error distribution surface; the geographic base map enhances the recognition and understanding of spatial positions.

[0208] 6.6 Integrate the generated various visualization charts and organize the output results according to meteorological elements and chart types.

[0209] First, establish a chart organizational structure, create classification directories according to meteorological elements (temperature, humidity, air pressure, wind speed), and each element contains various chart types such as violin plots, box plots, time series difference line charts, and geospatial heat maps. The system will generate standardized file names and metadata information for each chart, including chart type, element name, generation time, etc. Then, set the output format and quality parameters of the chart, support multiple formats such as PNG, PDF, SVG, etc., and set parameters such as resolution, size, and compression ratio to ensure the chart quality. Next, establish a chart index and directory structure, generate a chart list file, and record the paths, descriptions, and association relationships of each chart. Finally, implement the batch output function, save all the generated charts to the specified directory according to the organizational structure, and provide a complete visualization material library for the intelligent report generation of Example 7.

[0210] Through chart integration and output, a standardized visualization material library is established; the organizational structure is clear, facilitating the search and use of charts; multi-format support meets the needs of different application scenarios. Through dynamic data binding technology, a fully automated generation process from data to charts is realized; various visualization forms such as violin plots, box plots, time series line charts, and geographic heat maps are provided to comprehensively display data characteristics; through scientific color mapping, reasonable chart layout, and clear annotation information, the aesthetics and readability of the charts are ensured; the geospatial heat map visually displays the spatial distribution pattern of errors, providing a powerful tool for spatial analysis; a standardized chart organizational structure and output format are established, facilitating subsequent report generation and data management; rich visualization charts provide intuitive visual support for data quality assessment and decision-making analysis.

[0211] Example 7

[0212] This embodiment details the implementation of the intelligent report generation step in the automated quality assessment method for meteorological observation data. This step receives the output generated by the automated visualization in Embodiment 6 and the results of the error tolerance interval analysis in Embodiment 5, and automatically generates a complete meteorological quality assessment report through templated report generation technology, providing a report basis for the effectiveness assessment of the quality management system in Embodiment 8.

[0213] 7.1 Intelligent Report Generation First, it is necessary to select a suitable report template according to the evaluation purpose and data complexity, and perform personalized configuration.

[0214] First, establish a hierarchical report template library, including three purpose types: briefing type, standard type, and professional type. Each type is further divided into three complexity levels: simple, medium, and complex, for a total of nine template combinations. The system will automatically evaluate the data complexity based on characteristics such as the scale of the input data, the number of elements, and the analysis dimensions. Then, according to the evaluation purpose specified by the user or the system's intelligent recommendation, select the most suitable report template. Next, perform personalized configuration on the selected template, including setting basic information such as the report title, generation date, analysis period, and stations involved. Finally, verify the integrity and compatibility of the template to ensure that the template file exists and is compatible with the current data structure, preparing for subsequent report generation. Through intelligent template selection, the best match between the report content and the evaluation requirements is ensured; the hierarchical template library meets the report requirements for different scenarios; personalized configuration improves the pertinence and practicality of the report.

[0215] 7.2 Based on the error analysis results of Embodiment 5, automatically extract key quality problems and outliers to generate the report summary section.

[0216] First, extract key quality indicators from the error tolerance interval analysis results, including the mean absolute error of each meteorological element, the proportion within 5% of the relative error, the quality grade distribution, etc. The system will automatically judge the quality status of each element according to the preset quality assessment criteria, defining the proportion within 5% of the relative error above 90% as excellent, 70%-90% as good, 50%-70% as average, and below 50% as poor. Then calculate the overall quality status, adopt a weighted average method to comprehensively score the quality of each element, and generate an overall evaluation conclusion. Next, identify key quality problems, and through threshold comparison and anomaly detection algorithms, automatically identify the elements, stations, and time periods with poor quality. Finally, generate targeted improvement suggestions, and select appropriate improvement measures from the preset suggestion library according to the type and severity of the quality problems to form a complete report summary text. Through automatic summary generation, the core information of the evaluation results is quickly refined; intelligent problem identification improves the accuracy of quality problem discovery; targeted suggestions provide specific guidance for quality improvement.

[0217] 7.3 Based on the error analysis results and quality assessment indicators, automatically generate a standardized statistical result table.

[0218] First, design a standardized table template, including various types such as an error statistical table, an error interval distribution table, an error statistical table under EMI conditions, etc. The system extracts the corresponding data fields from the analysis results of Example 5 and organizes and formats the data according to the requirements of the table template. Then, perform data verification and cleaning, check the integrity and consistency of the data, handle missing values and outliers, and ensure the accuracy of the table data. Next, apply table styles and formatting, including header styles, data alignment, numerical precision, unit annotation, etc., to improve the readability and professionalism of the table. Finally, generate tables in multiple output formats, supporting formats such as HTML, Word, Excel, etc., to meet the usage needs of different users. Through automatic table generation, the standardized display of statistical results is ensured; the data verification mechanism guarantees the accuracy of the table content; multi-format support improves the applicability and convenience of the table.

[0219] 7.4 Based on the error statistical results and quality assessment indicators, automatically generate maintenance suggestions and optimization plans for each meteorological observation station.

[0220] First, calculate the error statistical indicators for each element at each station, including the mean absolute error, error standard deviation, outlier frequency, etc. The system compares the error levels of each station with the network average level to identify stations and elements with relatively high errors. Then, establish a maintenance priority assessment model, comprehensively considering factors such as error magnitude, data importance, and maintenance cost, and calculate the maintenance priority scores for each element at each station. Next, according to the priority scores and error characteristics, match the corresponding maintenance measures from the preset maintenance suggestion library, including different types of suggestions such as equipment calibration, environmental improvement, and equipment replacement. Finally, generate a structured maintenance suggestion table, containing information such as station number, element name, error level, priority, and specific suggestions, to provide decision support for maintenance management. Through automatic generation of maintenance suggestions, a scientific decision-making basis is provided for station management; the priority assessment model ensures the reasonable allocation of maintenance resources; the structured suggestion table facilitates the formulation and implementation of maintenance plans.

[0221] 7.5 Integrate the visual charts generated in Example 6 into the report to form a complete report with both text and pictures.

[0222] First, read various types of chart files from the visualization chart library of Example 6, including violin plots, box plots, time-series line charts, geographic heat maps, etc. The system will determine the position and size of each chart in the report according to the layout requirements of the report template. Then, perform format conversion and size adjustment on the charts to ensure the compatibility of the charts with the report format, and unify the resolution and color mode of the charts. Next, add titles, explanatory texts, and data source information to each chart to enhance the comprehensibility and traceability of the charts. Finally, establish the association relationship between the charts and the text content, insert references and analysis explanations of the charts in the report text, and form a complete report content with both pictures and texts. Through chart integration, the intuitiveness and persuasiveness of the report are enhanced; format unification ensures the professionalism and aesthetics of the report; picture-text association improves the logic and readability of the report content.

[0223] 7.6 Output the complete report content as documents in multiple formats according to user requirements.

[0224] First, integrate the various components of the report, including the abstract, statistical tables, maintenance suggestions, visualization charts, etc., to form a complete report content structure. The system will select the corresponding rendering engine and conversion tool according to the output format (Word, PDF, HTML, etc.) specified by the user. Then, perform format-specific style settings, including page layout, font style, chart embedding method, etc., to ensure the consistency and aesthetics of the report in different formats. Next, execute the report rendering and generation process, convert the structured report content into the final document format, and handle technical details such as chart embedding, page pagination, and table of contents generation. Finally, perform quality inspection and file output, verify the integrity and correctness of the generated document, save the final report to the specified location, and return the file path information.

[0225] Through multi-format output, the usage requirements of different users are met; the automated generation process greatly improves the report production efficiency; the quality inspection mechanism ensures the integrity and accuracy of the report. The whole process from template selection to content generation is intelligent, greatly reducing manual intervention; the report covers multiple aspects such as abstract, statistical analysis, maintenance suggestions, visualization charts, etc., with comprehensive content; standardized templates and professional style settings are adopted to ensure the standardization and professionalism of the report; support for personalized configuration and customization according to different evaluation purposes and data characteristics; support for multiple output formats such as Word, PDF, HTML, etc., to meet the needs of different application scenarios; through automatic problem identification and suggestion generation, strong support is provided for quality management decision-making.

[0226] Example 8

[0227] This embodiment details the implementation of the effectiveness evaluation step of the quality management system in the automated quality assessment method for meteorological observation data. This step receives the output results of the foregoing embodiments, comprehensively evaluates the operation effect of the quality management system of meteorological observation operations through a multi-dimensional index system, and forms a complete quality management closed-loop.

[0228] 8.1 By analyzing the equipment operation logs, evaluate the stable operation rate and business continuity of meteorological observation equipment.

[0229] First, extract key information from the equipment operation logs, including fields such as station number, equipment number, operation status, start time, end time, etc. The system will preprocess the log data, unify the time format, clean abnormal records, and supplement missing information. Then calculate the operation time statistics of each equipment, including total monitoring time, normal operation time, downtime, and calculate the availability rate index (availability rate = normal operation time / total monitoring time). Next, summarize the equipment availability indicators by station dimension, calculate the average availability rate, minimum availability rate, and maximum availability rate of each station, and identify stations and equipment with low availability. Finally, establish an availability evaluation standard, define an availability rate above 95% as excellent, 90% - 95% as good, 85% - 90% as average, and below 85% as poor, and generate a business availability evaluation report. Through the business availability evaluation, the stability level of equipment operation is quantified; the summary analysis at the station dimension facilitates the identification of problematic equipment and stations; the evaluation standard provides a clear quality goal for equipment management.

[0230] 8.2 By analyzing the data transmission logs, evaluate the integrity, timeliness, and reliability of meteorological data transmission.

[0231] First, extract transmission records from the data transmission logs, including information such as station number, transmission time, transmission status, data volume, delay time, etc. The system will define a transmission quality evaluation standard, and define records with a transmission status of successful and a delay time not exceeding 5 minutes as real-time transmission success. Then calculate the transmission quality indicators of each station, including total transmission times, successful transmission times, real-time transmission times, transmission success rate, real-time transmission rate, data integrity rate, etc. Next, analyze the time variation trend of transmission quality, and identify time periods with declining transmission quality and possible reasons. Finally, establish a transmission quality evaluation level, define a transmission success rate above 98% and a real-time transmission rate above 95% as excellent, and provide an evaluation basis for the optimization of the data transmission system. Through the transmission quality evaluation, a comprehensive understanding of the operation status of the data transmission system is obtained; the time trend analysis helps to identify systematic problems; the evaluation level provides a clear direction for the improvement of the transmission system.

[0232] 8.3 By analyzing the fault logs and maintenance records, evaluate the equipment failure frequency and maintenance response efficiency.

[0233] First, integrate the fault log and maintenance record data to establish the fault - maintenance correlation relationship, and match the fault occurrence time and maintenance processing time of the same device. The system will calculate key maintenance indicators, including response time (the time interval from fault occurrence to the start of maintenance) and maintenance time (the time interval from the start of maintenance to the completion of maintenance). Then, statistically analyze the fault maintenance indicators by device and site dimensions, and calculate statistics such as fault frequency, average response time, and average maintenance time. Next, analyze the distribution characteristics of fault types and severity levels, and identify high - frequency fault types and fault modes with greater impact. Finally, establish a fault maintenance evaluation criterion, defining excellent maintenance efficiency as an average response time within 2 hours and an average maintenance time within 8 hours, providing an evaluation basis for optimizing maintenance management. Through the fault maintenance evaluation, the efficiency level of maintenance management is quantified; the fault type analysis provides an important reference for preventive maintenance; and the evaluation criterion indicates the improvement direction for optimizing the maintenance process.

[0234] 8.4 Evaluate the user's satisfaction with the meteorological data quality and services by analyzing the user questionnaire data.

[0235] First, extract the evaluation data from the user questionnaire, including information such as user type, data quality score, service quality score, and overall satisfaction score. The system will standardize the score data, unify the scoring scale, and handle missing values and outliers. Then, calculate the satisfaction indicators by grouping according to user type, including average score, satisfaction ratio (the ratio of scores above 4), dissatisfaction ratio, etc. Next, analyze the influencing factors of satisfaction, and identify the key factors affecting user satisfaction, such as data timeliness, accuracy, service response speed, etc., through correlation analysis. Finally, establish a user satisfaction evaluation level, defining excellent as an average satisfaction score above 4.0 and good as 3.5 - 4.0, providing a basis for user feedback for improving service quality. Through the user satisfaction evaluation, the true feelings of users about service quality are understood; the analysis of influencing factors provides targeted guidance for service improvement; and the evaluation level provides an evaluation criterion from the user's perspective for service quality management.

[0236] 8.5 This embodiment details the implementation of the steps for evaluating the effectiveness of the quality management system in the automated quality assessment method of meteorological observation data. This step quantitatively evaluates the operation effect of the quality management system of meteorological observation operations through multi - dimensional indicators such as business availability indicators, data transmission quality indicators, fault and maintenance situation indicators, user satisfaction indicators, and external supplier evaluation indicators. Evaluate the quality of meteorological equipment, services, and system platforms provided by external suppliers by analyzing the external supplier evaluation data.

[0237] First, extract evaluation records from the supplier evaluation data, including information such as supplier ID, supplier name, supply type, product quality score, service quality score, delivery timeliness score, and overall evaluation. The system will standardize the evaluation data, unify the scoring scale, and handle missing values and outliers. Then calculate the evaluation indicators by supplier dimension, including the number of evaluations, average product quality score, average service quality score, average delivery timeliness score, average overall evaluation, and the qualified ratio of each indicator (the ratio of scores above 3.5). Next, group and count the evaluation indicators by supply type, and calculate the number of suppliers and average scores for each supply type. Finally, establish an external supplier evaluation grade, defining suppliers with an average overall evaluation above 4.0 as excellent suppliers and those with an average overall evaluation between 3.5 and 4.0 as good suppliers, providing an evaluation basis for supplier management and selection. Through the external supplier evaluation, the product and service quality of external suppliers can be comprehensively understood; the classified statistics help identify high-quality suppliers and problematic suppliers; the evaluation grade provides an objective standard for supplier management and selection.

[0238] 8.6 By integrating multi-dimensional indicators such as business availability, data transmission quality, fault repair, user satisfaction, and external supplier evaluation, a complete evaluation system for the effectiveness of the quality management system is formed.

[0239] First, establish an evaluation framework for the effectiveness of the quality management system, defining the weights and importance of each evaluation dimension. The system will sequentially execute five sub-evaluation modules, including business availability indicator evaluation, data transmission quality indicator evaluation, fault and repair situation indicator evaluation, user satisfaction indicator evaluation, and external supplier evaluation indicator evaluation. Each sub-module performs conditional execution based on the availability of input data to ensure the flexibility and adaptability of the evaluation. Then extract the key indicators of each sub-module, including business availability, data transmission success rate, data integrity rate, average response time, average repair time, overall user satisfaction, data quality satisfaction ratio, supplier overall evaluation, etc. Next, calculate the comprehensive effectiveness index of the quality management system, and integrate the indicators of each dimension into an overall effectiveness evaluation through weighted average or comprehensive scoring. Finally, generate an evaluation report on the effectiveness of the quality management system, including detailed evaluation results for each dimension and an overall effectiveness conclusion, providing comprehensive support for quality management decisions. Through the integrated evaluation, a complete evaluation framework for the quality management system is formed; the multi-dimensional indicators ensure the comprehensiveness and objectivity of the evaluation; the comprehensive index facilitates the management to quickly understand the overall situation of the system; the evaluation report provides a clear direction and focus for quality improvement.

[0240] 8.7 Establish a data flow and result association mechanism between each evaluation module to ensure the consistency and traceability of evaluation results.

[0241] First, establish a unified data interface standard, define the format requirements and field specifications for various types of input data, and ensure the standardization and consistency of data. The system will establish a data preprocessing pipeline to clean, verify, and convert the format of various types of input logs and evaluation data, unify the time format, handle missing values, and standardize the encoding. Then, establish an evaluation result correlation mechanism to perform correlation analysis on the results of different evaluation modules through key fields such as site number, equipment number, and timestamp, and identify the correlations between influencing factors. Next, establish an evaluation result consistency verification mechanism to ensure the logical consistency of the evaluation results of each module through cross-validation and logical verification, such as sites with low business availability corresponding to higher failure frequencies. Finally, establish an evaluation result traceability mechanism to record the source of input data, processing process, and result output of each evaluation, and support the auditing and retrospective analysis of evaluation results. Through the data flow mechanism, the standardization of the evaluation process and the reliability of the results are ensured; the correlation analysis reveals the internal relationships between different quality indicators; the consistency verification improves the accuracy of the evaluation results; and the traceability mechanism enhances the transparency and credibility of the evaluation process.

[0242] 8.8 Establish a dynamic monitoring and early warning mechanism for the effectiveness of the quality management system to achieve real-time tracking of the quality status and early warning of abnormalities.

[0243] First, establish a quality indicator monitoring threshold system, and set the normal range, early warning threshold, and alarm threshold for each quality indicator based on historical data and industry standards. The system will establish a real-time data collection mechanism to regularly collect data such as operation logs, transmission logs, and fault records from each business system to ensure the timeliness and integrity of monitoring data. Then, establish a quality indicator calculation engine to automatically calculate the quality indicators of each dimension according to the preset calculation rules and cycles, including daily indicators, weekly indicators, monthly indicators, and other different time granularities. Next, establish an anomaly detection and early warning mechanism to timely detect abnormal changes in quality indicators through methods such as threshold comparison, trend analysis, and anomaly pattern recognition, and automatically trigger early warning notifications. Finally, establish a visual display of the quality status, and intuitively display the operation status and change trend of the quality management system through dashboards, trend charts, heat maps, etc., to support management decision-making. Through the dynamic monitoring mechanism, real-time control of the quality status is achieved; the early warning mechanism ensures the timely discovery and handling of problems; the visual display improves the understandability of quality information; and the dynamic tracking provides continuous data support for quality improvement.

[0244] 8.9 Establish a continuous improvement mechanism for the quality management system based on evaluation results to form a closed-loop management of evaluation - analysis - improvement - verification.

[0245] First, establish an analysis framework for evaluation results to deeply analyze the evaluation results of the effectiveness of the quality management system, identify the root causes of quality problems and improvement opportunities. The system will establish a problem classification and priority ranking mechanism, and classify and rank the identified quality problems according to the severity, impact scope and improvement difficulty of the problems to determine the priority order of improvement. Then formulate targeted improvement measures, including different types of improvement actions such as process optimization, technology upgrade, personnel training, and system improvement, to form a specific improvement plan and schedule. Next, establish an improvement effect tracking mechanism, and monitor the implementation effect of improvement measures and verify the achievement of improvement goals through regular evaluation and comparative analysis. Finally, establish an improvement experience summary and promotion mechanism, form standardized best practices from successful improvement experiences, and promote and apply them throughout the quality management system to achieve the spiral rise of continuous improvement. Through the continuous improvement mechanism, a closed-loop system for quality management is established; problem-oriented improvement ensures the effective allocation of resources; effect tracking guarantees the effectiveness of improvement measures; and experience promotion realizes the overall improvement of the quality management level.

[0246] 8.10 In this embodiment, through the establishment of a complete evaluation system for the effectiveness of the quality management system, the overall improvement and continuous optimization of the quality management of meteorological observation data are achieved.

[0247] Through the evaluation of business availability indicators, the quantitative monitoring of the stability of meteorological observation equipment is realized, and the equipment availability rate is increased from the original 85% to over 95%, significantly improving the continuity of observation operations. Through the evaluation of data transmission quality indicators, the accurate control of the integrity and timeliness of data transmission is realized, the data transmission success rate reaches over 99%, and the real-time transmission rate reaches over 98%, ensuring the high-quality transmission of data. Through the evaluation of indicators for faults and maintenance situations, the effective management of equipment faults and maintenance efficiency is realized, the average fault response time is shortened to within 1 hour, and the average maintenance time is controlled within 4 hours, greatly improving the maintenance efficiency. Through the evaluation of user satisfaction indicators, the user perspective evaluation of service quality is realized, the overall user satisfaction reaches 4.2 points (out of 5), and the proportion of users satisfied with data quality reaches over 90%, significantly improving the user experience. Through the evaluation of external supplier evaluation indicators, the objective evaluation of supplier quality is realized, the proportion of high-quality suppliers reaches over 80%, and the overall service level of suppliers is significantly improved.

[0248] The implementation of the evaluation of the effectiveness of the quality management system has enabled the quality management of meteorological observation data to transform from qualitative management to quantitative management, from passive response to active prevention, and from single-point improvement to system optimization. The overall quality management level has been significantly improved, the stability of data quality has been enhanced, and user satisfaction has been continuously improved, laying a solid foundation for the overall improvement of meteorological service quality.

[0249] The effectiveness evaluation method of the quality management system in this embodiment has the advantages of comprehensive evaluation dimensions, scientific index design, strong adaptability, intuitive and clear results, powerful decision-making support, and complete management closed-loop. Through multi-dimensional quantitative evaluation, it realizes the refinement and scientificization of quality management, providing strong support for the high-quality development of meteorological observation business.

[0250] Evaluate the effectiveness of the quality management system from multiple dimensions such as business availability, data transmission quality, faults and repairs, user satisfaction, and external supplier evaluation to comprehensively understand the system operation status; design scientific and reasonable indicators for each evaluation dimension, such as availability rate, transmission success rate, average response time, user satisfaction, etc., to quantify the evaluation results for easy comparison and analysis; the function design supports flexible input, and the evaluation dimensions and indicators can be selected according to the actual situation to adapt to different evaluation needs; the evaluation results are organized by dimension and indicator, with clear levels for easy understanding and use; by quantitatively evaluating the effectiveness of the quality management system, it provides an objective basis for management decision-making to guide system optimization and improvement; the management closed-loop is complete: covering all aspects of the quality management system, from equipment operation, data transmission, fault repair to user feedback and supplier management, forming a complete management closed-loop.

[0251] Such as Figure 2 As shown, an automated quality evaluation system for meteorological observation data includes:

[0252] Data initialization module: used to collect multiple meteorological observation data and perform initialization configuration to obtain a meteorological evaluation parameter set;

[0253] Data preprocessing module: used to perform preprocessing operations on each parameter in the meteorological evaluation parameter set to obtain a complete meteorological evaluation data set;

[0254] Statistical analysis module: used to perform global statistical analysis on the complete meteorological evaluation data set to obtain a global statistical analysis data set; divide the global statistical analysis data set into meteorological quality evaluation subsets by region and depth level; and further analyze the meteorological quality evaluation subsets to obtain hierarchical statistical analysis results;

[0255] Error analysis module: used to obtain the absolute error and relative error between the meteorological evaluation manual station and automatic station data according to the hierarchical statistical analysis results, and set the proportion of data within each error tolerance interval, and obtain meteorological quality characteristic data through iterative verification;

[0256] Report generation module: used to generate meteorological quality distribution characteristics in different regions and levels according to the meteorological quality characteristic data to obtain the time series difference between the meteorological evaluation manual station and automatic station data, and save the generated data and the obtained data to folders respectively, and output the final meteorological quality evaluation report by matching the output formats set by different users.

[0257] The embodiments of the present invention have been described above. However, these embodiments are not limited to the specific implementation manners described above. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of these embodiments, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of these embodiments.

Claims

1. An automated quality assessment method for meteorological observation data, characterized in that, Including the following steps: Collect multiple meteorological observation data and perform initialization configuration to obtain a meteorological evaluation parameter set; Perform preprocessing operations on each parameter in the meteorological evaluation parameter set to obtain a complete meteorological evaluation data set; Perform global statistical analysis on the complete meteorological evaluation data set to obtain a global statistical analysis data set; Divide the global statistical analysis data set into meteorological quality evaluation subsets according to regions and depth levels; and further analyze the meteorological quality evaluation subsets; Among them, further analyzing the meteorological quality evaluation subset includes: performing grouped statistical analysis through three levels of stratification in the time dimension, spatial dimension, and element dimension, conducting comprehensive analysis in combination with the EMI environmental meteorological condition evaluation index, and performing quantitative separation analysis by quantitatively characterizing the impact of meteorological conditions on the quality evaluation of each subset data item, and finally integrating different analysis results to obtain a hierarchical statistical analysis result; Obtain the absolute error and relative error between the meteorological evaluation manual station and automatic station data according to the hierarchical statistical analysis result, set the proportion of data within each error tolerance interval, and obtain meteorological quality characteristic data through iterative verification; Generate meteorological quality distribution characteristics for different regions and levels according to the meteorological quality characteristic data to obtain the time series difference between the meteorological evaluation manual station and automatic station data, save the generated data and the obtained data to folders respectively, and output the final meteorological quality evaluation report by matching the output formats set by different users.

2. The automated quality assessment method for meteorological observation data according to claim 1, characterized in that The initialization configuration includes: Create and define multiple initial states, and mark the meteorological observation data as correct, suspicious, incorrect, and missing respectively to obtain the marked data of the initial state of the meteorological observation data; Construct meteorological quality evaluation indicators according to the marked data, and the meteorological quality evaluation indicators are the actual rate, correct rate, error rate, and suspicious rate; Detect whether the initialization configuration of the meteorological observation data is completed in the current meteorological observation environment according to the marked data; Load data allocation resource configuration for the meteorological observation data that has not completed the initialization configuration according to the meteorological quality evaluation indicators, and the data allocation resource configuration dynamically adjusts the memory usage and thread allocation according to the data scale and resources to complete the initialization configuration of the meteorological quality evaluation and obtain the initialized meteorological evaluation parameter set.

3. The automated quality assessment method for meteorological observation data according to claim 1, characterized in that Data preprocessing includes: Receive the configuration path parameter of the environment adaptive initialization configuration, and the configuration path parameter is used to cooperate with the meteorological observation data for initialization configuration operations; Automatically verify the configuration path parameter to judge whether the configuration path parameter exists and is accessible; When the configuration path parameter exists and is accessible, load the meteorological observation data and read and check whether the meteorological observation data is complete; When the meteorological observation data is incomplete, perform format conversion adjustment and missing value supplementation processing on the meteorological observation data, and at the same time perform normalization processing to obtain the processed meteorological evaluation parameter set, which is the complete meteorological evaluation data set.

4. The automated quality assessment method for meteorological observation data according to claim 3, characterized in that Create multiple quality evaluation codes, and the quality evaluation codes include 0-3 evaluation thresholds as quality evaluation digital codes; Mark the current multiple meteorological observation data with the quality evaluation data respectively to obtain the marking result; Set the digital range for the marking results, and each digital range represents four quality assessment states: excellent, good, average, and poor; Construct a meteorological quality identification system based on the digital range set by the quality assessment code and the quality assessment state; Based on the meteorological quality identification system, identify the dynamic changes of current multiple meteorological observation data at different times to obtain the change trend of the meteorological observation state; Conduct multi-dimensional statistical analysis according to the change trend of the meteorological observation state to obtain the quality state of the meteorological observation data under multi-dimensional analysis, and judge whether the change trend amplitude between the current meteorological observation data and the meteorological observation data after dynamic change is in a gentle or rapid state, so as to configure a hierarchical statistical analysis strategy.

5. The automated quality assessment method for meteorological observation data according to claim 4, wherein The hierarchical statistical analysis strategy includes: Capture the abnormal state of the dynamic change of meteorological observation data at different times to obtain abnormal state data; Interpolate the meteorological time series characteristics and spatial correlation of the abnormal state data to obtain the sudden drop or sudden rise amplitude index, and use the quality assessment code to mark the abnormal state data according to the amplitude index to obtain a secondary marking result; Further conduct climate boundary value check, internal consistency check, time consistency check and spatial consistency check on the abnormal state data in the secondary marking result, and use the quality assessment code to mark the data that does not conform to the meteorological inspection rules to obtain a tertiary marking result; Perform fitting processing on the meteorological observation data with general or poor quality assessment states in the three-round marking results to obtain a processing result; Integrate the processed meteorological observation data with the original meteorological observation data with excellent or good quality assessment states to generate a preliminary quality assessment result and transmit it to the global statistical analysis for collaborative analysis and processing.

6. The automated quality assessment method for meteorological observation data according to claim 1, characterized in that Perform global statistical analysis on the entire meteorological assessment data set to obtain a global statistical analysis data set, including: Set a meteorological micro-interference filtering threshold, and dynamically adjust it according to the meteorological regional characteristics combined with the preliminary quality assessment result to optimize the observation quality of each data item in the entire meteorological assessment data set; Obtain the basic statistical indicators such as the extreme value, mean value, standard deviation and coefficient of variation of the meteorological elements of each data item in the entire meteorological assessment data set, and establish a global meteorological quality assessment benchmark; Based on the global meteorological quality assessment benchmark, conduct multi-level grouping statistical analysis on meteorological data according to the time dimension, space dimension and element dimension to obtain a hierarchical statistical analysis strategy to form an environmental quality assessment data set; Combine the hierarchical statistical analysis strategy with the meteorological quality identification system and the quality assessment indicators such as the actual occupancy rate, correct rate, error rate and suspicious rate at each level, and introduce the EMI environmental meteorological condition assessment index for comprehensive quantitative analysis; The EMI environmental meteorological condition assessment index represents a comprehensive index characterizing the processes of aerosol emission, sedimentation, transport and diffusion under the influence of meteorological conditions. The larger the EMI value, the more unfavorable the meteorological conditions are for the diffusion of air pollutants; Through iterative statistical verification and threshold determination, fuse the global statistical results and the hierarchical statistical results, and finally generate a complete global statistical analysis data set including multi-dimensional statistical features, quality assessment identification and EMI environmental meteorological condition assessment index.

7. The automated quality assessment method for meteorological observation data according to claim 6, characterized in that Divide the global statistical analysis dataset into meteorological quality assessment subsets by region and depth level; and further analyze the meteorological quality assessment subsets to obtain hierarchical statistical analysis results, including: Stratify each data item in the global statistical analysis dataset in the time dimension, including grouping and statistics by year, season, month, and day, combine the quality assessment marking results to form a meteorological quality assessment subset at the time level, and generate time-series basic data; Stratify each data item in the global statistical analysis dataset in the space dimension, including grouping and statistics by region, geographical features, and elevation, use the regional characteristics to adjust the results to form a meteorological quality assessment subset at the space level, and generate spatial distribution data; Stratify each data item in the global statistical analysis dataset in the element dimension, calculate the quality assessment indicators of meteorological elements such as temperature, humidity, air pressure, and wind speed respectively, obtain the hierarchical statistical analysis results at the element level, and generate element distribution characteristic data; Adopt the EMI environmental meteorological condition assessment index for comprehensive analysis to obtain hierarchical statistical analysis results; Conduct quantitative separation analysis by quantitatively characterizing the impact of meteorological conditions on the quality assessment of data items in each meteorological quality assessment subset to obtain quantitative separation analysis results; The said quantitative separation analysis includes: based on the EMI environmental meteorological condition assessment index, obtain the emission change rate RE and the meteorological condition change rate RW through RE=(R1 / R0) / (E1 / E0)-1 and RW=(E1 / E0)-1; where, R0 and R1 are the measured meteorological data parameters in the comparison period respectively, and E0 and E1 are the EMI environmental meteorological condition assessment indexes in the comparison period respectively; Analyze through the EMI environmental meteorological condition assessment index to obtain the contribution rate of meteorological condition changes to data quality, to obtain the correlation between the EMI environmental meteorological condition assessment index and the meteorological observation data quality index, and quantitatively evaluate the impact degree of meteorological conditions on the quality of observation data; Based on the spatio-temporal distribution characteristics of EMI, analyze the change trend of meteorological conditions in different regions and the impact degree on the quality of meteorological observation data; Integrate the two impact degrees to obtain quantitative separation analysis results.

8. The automated quality assessment method for meteorological observation data according to claim 1, characterized in that Obtain the absolute error and relative error between the meteorological assessment manual station and automatic station data according to the hierarchical statistical analysis results, and set the proportion of data within each error tolerance interval, and obtain meteorological quality characteristic data through iterative verification, including: The error tolerance interval analysis is based on the hierarchical statistical analysis results and the EMI environmental meteorological condition assessment index, and is realized by constructing an error statistical function. The error statistical function receives the meteorological assessment manual station observation data and automatic station observation data as input parameters and performs the following operations: Calculate the absolute error value and relative error percentage between the two groups of meteorological observation data in combination with the hierarchical statistical analysis results at the element level; And use the meteorological quality assessment subsets at the time and space levels to count the proportion of data with relative error within the ±5% threshold range and absolute error value within the ±10% threshold range; Calculate the standard deviation based on the EMI environmental meteorological condition evaluation index, and finally return it in the statistical result list including the proportion of the error interval containing the absolute error value and relative error and the standard deviation. Generate an error interval ratio distribution table according to the calculation results and pass it to the meteorological quality characteristic data as a key quantitative index.

9. The automated quality assessment method for meteorological observation data according to claim 8, characterized in that Dynamically bind data fields based on the element distribution characteristic data and the error interval ratio distribution table to obtain the probability density distribution characteristics of meteorological elements in different regions and levels. Combine the meteorological quality identification system to label the data quality level to support the time series feature analysis with the basic data. Use the probability density distribution characteristics and the time dimension stratified data, combine with the error statistical results to generate time series differences, so as to obtain the time change trend of the meteorological evaluation manual station and automatic station data. Distinguish the data points within different error tolerance intervals through error coding and pass the time series characteristic data to the spatial feature analysis. Based on the data characteristics obtained from the time series feature analysis and the spatial dimension stratified data, combine with the error interval ratio distribution to obtain the geospatial thermal effect, so as to obtain the spatial distribution characteristics of the meteorological quality evaluation index. Dynamically adjust the error mapping range according to the EMI environmental meteorological condition evaluation index, and store the spatial distribution characteristics at the same time. Obtain the meteorological quality characteristic data through different dimensional analyses such as probability density distribution, time series feature analysis and spatial feature analysis, generate a comprehensive report on the meteorological quality distribution characteristics in different regions and levels, integrate the time series difference analysis results of the meteorological evaluation manual station and automatic station data, and save the generated data and the obtained data to the corresponding folders according to the analysis dimensions respectively. That is, output the final meteorological quality evaluation report by matching the output formats set by different users.

10. A meteorological observation data automatic quality assessment system for executing a meteorological observation data automatic quality assessment method as described in any one of claims 1-9, characterized in that, Including: Data initialization module: used to collect multiple meteorological observation data and perform initialization configuration to obtain a set of meteorological evaluation parameters. Data preprocessing module: used to perform preprocessing operations on each parameter in the set of meteorological evaluation parameters to obtain the full set of meteorological evaluation data. Statistical analysis module: used to perform global statistical analysis on the full set of meteorological evaluation data to obtain the global statistical analysis data set. Divide the global statistical analysis data set into meteorological quality evaluation subsets according to regions and depth levels; and further analyze the meteorological quality evaluation subsets. Among them, the further analysis of the meteorological quality evaluation subsets includes: performing grouped statistical analysis through three levels of stratification in the time dimension, spatial dimension and element dimension, combining with the EMI environmental meteorological condition evaluation index for comprehensive analysis, and performing quantitative separation analysis by quantitatively characterizing the impact of meteorological conditions on the quality evaluation of each subset data item. Integrate different analysis results to finally obtain the hierarchical statistical analysis results. Error analysis module: used to obtain the absolute error and relative error of the meteorological evaluation manual station and automatic station data according to the hierarchical statistical analysis results, set the proportion of data within each error tolerance interval, and obtain the meteorological quality characteristic data through iterative verification. Report generation module: It is used to generate the meteorological quality distribution characteristics of different regions and levels based on the meteorological quality characteristic data, so as to obtain the temporal sequence differences between the meteorological assessment manual stations and the automatic station data, and save the generated data and the obtained data into folders respectively. By matching the output formats set by different users, the final meteorological quality assessment report is output.

Citation Information

Patent Citations

  • Weather forecast data quality detection method

    CN113742927A

  • Technology for monitoring and evaluating low-atmosphere self-purification capacity process

    CN115081845A

  • Meteorological space normalization method and system based on random forest

    CN115840793A

  • Multi-source meteorological data integration and quality control system

    CN117333046A

  • Meteorological unmanned aerial vehicle ground measurement and control platform

    CN118068448A

Cited By

  • Ocean apparent optical measurement data integration method, device, equipment and medium

    CN120950840A

  • PeakFit data automatic processing method and system

    CN121455956A

  • Method, system and equipment for verifying data coupling integrity in flood simulation

    CN121902461A

  • Quality control methods, devices, equipment, media, and procedures for meteorological observation data

    CN122573282A