A method and system for automatic quality assessment of meteorological observation data

Through environmental adaptive initialization and intelligent data preprocessing, combined with multi-dimensional statistical analysis and automated report generation, the problems of low efficiency and poor accuracy of traditional meteorological observation data quality assessment are solved, and the automated quality assessment and standardized report generation of meteorological data are realized.

CN120410340BActive Publication Date: 2025-08-26辽宁省生态气象和卫星遥感中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510915995.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-26
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Traditional artificial meteorological observation data quality assessment methods are time-consuming and error-prone, and cannot automatically identify missing components, cannot automatically generate standardized reports, and lack path verification mechanisms, resulting in program interruptions.

Method used

Using environmental adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis and automated report generation methods, automated quality assessment is realized through the R software package, including data initialization, preprocessing, global statistical analysis, hierarchical statistical analysis, error tolerance interval analysis and report generation.

Benefits of technology

It realizes automation of meteorological data quality evaluation, improves efficiency and accuracy, reduces manual operation errors, provides multi-dimensional comprehensive analysis capabilities, generates standardized reports, and enhances the stability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410340B_ABST
    Figure CN120410340B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of meteorological quality assessment, and discloses a method and system for automated quality assessment of meteorological observation data. The method comprises: collecting observation data, obtaining a meteorological assessment parameter set through initialization configuration; performing preprocessing operations to obtain a full meteorological assessment data set; performing global statistical analysis at the same time to obtain a global statistical analysis data set; dividing the global statistical analysis data set into meteorological quality assessment subsets according to regions and depth levels; obtaining hierarchical statistical analysis results; setting the data proportion within each error tolerance interval by obtaining absolute error and relative error; generating meteorological quality distribution characteristics of different regions and levels to obtain time series differences, and finally outputting a final meteorological quality assessment report; thereby realizing environmental adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis, error tolerance interval analysis, and automatic report generation, so as to improve the efficiency of meteorological data quality assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of meteorological quality assessment, and more particularly, to a method and system for automatic quality assessment of meteorological observation data. Background Art

[0002] Meteorological observation data is a crucial foundation for weather forecasting, climate research, and environmental monitoring. Its quality directly impacts the accuracy and reliability of meteorological services. With the continuous expansion of meteorological observation networks and increased automation, the volume of meteorological observation data is growing exponentially. Traditional manual quality assessment methods are no longer sufficient for large-scale data processing.

[0003] Currently, traditional meteorological observation data quality assessment and comparison rely on manual operations, requiring the calculation of statistical indicators item by item, which is time-consuming and prone to errors. Existing tools require manual installation of dependency packages and configuration of the environment, and are unable to automatically identify missing components and complete configuration deployment; nor can they automatically generate standardized reports. There is a lack of a path verification mechanism when loading data, which can easily lead to program interruptions due to missing files.

[0004] Therefore, how to achieve environmental adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis, error tolerance analysis, and automatic report generation to improve the efficiency and accuracy of meteorological data quality assessment has become an urgent problem to be solved. Summary of the Invention

[0005] The present invention provides a method and system for automated quality assessment of meteorological observation data, which solves the technical problem in the prior art that the ultra-early stage risk of autoimmune diseases cannot be effectively predicted.

[0006] The present invention provides a method and system for automatically evaluating the quality of meteorological observation data, comprising:

[0007] In a first aspect, a method for automatically evaluating the quality of meteorological observation data comprises the following steps:

[0008] Collect multiple meteorological observation data and perform initial configuration to obtain a set of meteorological assessment parameters;

[0009] Perform preprocessing operations on various parameters in the meteorological assessment parameter set to obtain the full meteorological assessment data set;

[0010] Performing global statistical analysis on the entire meteorological assessment data set to obtain a global statistical analysis data set;

[0011] The global statistical analysis dataset is divided into meteorological quality assessment subsets according to regional and depth levels; and the meteorological quality assessment subsets are further analyzed to obtain hierarchical statistical analysis results;

[0012] Based on the results of hierarchical statistical analysis, the absolute and relative errors of the meteorological assessment data from manual and automatic stations are obtained, and the proportion of data within each error tolerance range is set. After iterative verification, meteorological quality characteristic data is obtained.

[0013] Based on the meteorological quality characteristic data, meteorological quality distribution characteristics of different regions and levels are generated to obtain the temporal differences between the meteorological assessment manual station and automatic station data, and the generated data and acquired data are saved in folders respectively. By matching the output formats set by different users, the final meteorological quality assessment report is output.

[0014] Furthermore, the initialization configuration includes:

[0015] Create and define multiple initial states, and mark meteorological observation data as correct, suspicious, erroneous, and missing, respectively, to obtain marked data of the initial states of meteorological observation data;

[0016] Constructing meteorological quality assessment indicators based on the labeled data, wherein the meteorological quality assessment indicators are actual rate, correct rate, error rate and suspicious rate;

[0017] Detecting whether the initialization configuration of meteorological observation data is completed in the current meteorological observation environment according to the marked data;

[0018] According to the meteorological quality assessment indicators, the meteorological observation data that has not completed the initialization configuration is loaded into the data allocation resource configuration. The data allocation resource configuration dynamically adjusts the memory usage and thread allocation according to the data scale and resources to complete the operation meteorological quality assessment initialization configuration and obtain the initialized meteorological assessment parameter set.

[0019] Furthermore, data preprocessing includes:

[0020] By receiving the configuration path parameters of the environment adaptive initialization configuration, the configuration path parameters are used to perform the initialization configuration operation in conjunction with the meteorological observation data;

[0021] Automatically check the configuration path parameters to determine whether they exist and are accessible;

[0022] When the configuration path parameters exist and are accessible, load the meteorological observation data, read and verify whether the meteorological observation data is complete;

[0023] When the meteorological observation data is incomplete, the format conversion and missing value supplement processing are performed on the meteorological observation data, and normalization processing is performed at the same time to obtain the processed meteorological assessment parameter set, which is the full meteorological assessment data set.

[0024] Furthermore, the operating environment quality assessment configuration includes:

[0025] Creating a plurality of quality assessment codes, wherein the quality assessment codes include assessment thresholds of 0 to 3 as quality assessment digits;

[0026] The quality assessment data respectively marks the current multiple meteorological observation data to obtain the marking results;

[0027] The marking results are set into a digital range, each digital range represents four quality assessment states: excellent, good, average and poor;

[0028] Establish a meteorological quality identification system based on the digital range set by the quality assessment code and quality assessment status;

[0029] Based on the meteorological quality identification system, the dynamic changes of multiple meteorological observation data at different times are identified to obtain the changing trend of meteorological observation status;

[0030] According to the changing trend of meteorological observation status, multi-dimensional statistical analysis is performed through the loaded R software package to obtain the quality assessment status corresponding to the quality assessment code of meteorological observation data in the multi-dimensional analysis, which is the quality status of observation data under the multi-dimensional analysis;

[0031] As multi-dimensional analysis progresses, the data status will be re-evaluated; each status (such as normal, slightly abnormal, and strongly abnormal) will be mapped to a quality assessment code. The quality assessment code is not fixed at one time, but can change as the analysis dimension deepens. For example: original observation data (temperature, humidity, wind speed, etc.) → preliminary quality assessment (based on single-point inspection) → initial assessment code (such as: good) → multi-dimensional analysis (temporal trend, spatial comparison, climate constraints, variable relationship) → status change identification (such as deviation from the norm / inconsistency) → quality assessment code update (good → abnormal) → quality assessment status result generation (abnormal → "needs correction");

[0032] And judge whether the change trend between the current meteorological observation data and the meteorological observation data after dynamic changes is in a gentle or rapid state, so as to configure the hierarchical statistical analysis strategy.

[0033] Furthermore, the stratified statistical analysis strategy includes:

[0034] Capture the abnormal state of meteorological observation data that changes dynamically at different times to obtain abnormal state data;

[0035] The meteorological time series characteristics and spatial correlation of the abnormal state data are interpolated to obtain the sudden drop or sudden rise amplitude index, and the abnormal state data are marked according to the amplitude index using the quality assessment code to obtain the secondary marking result;

[0036] The abnormal state data in the secondary marking results are further checked for climate limit values, internal consistency, temporal consistency, and spatial consistency. The data that does not meet the meteorological inspection rules are marked using quality assessment codes to obtain the tertiary marking results.

[0037] Fitting is performed on the meteorological observation data with general or poor quality assessment status in the three rounds of marking results to obtain processing results;

[0038] The processed meteorological observation data and the original meteorological observation data with excellent or good quality assessment status are integrated to generate preliminary quality assessment results and transmitted to the global statistical analysis for coordinated analysis and processing.

[0039] Furthermore, a global statistical analysis is performed on the entire meteorological assessment data set to obtain a global statistical analysis data set, including:

[0040] Set the meteorological micro-disturbance filtering threshold and dynamically adjust it according to the characteristics of the meteorological region and the preliminary quality assessment results to optimize the observation quality of each data item in the full meteorological assessment dataset;

[0041] Obtain basic statistical indicators of extreme values, mean values, standard deviations, and coefficients of variation of meteorological elements for each data item in the full meteorological assessment dataset, and establish a global meteorological quality assessment benchmark;

[0042] Based on the global meteorological quality assessment benchmark, a multi-level grouping statistical analysis of meteorological data is performed according to the time dimension, space dimension and element dimension to obtain a hierarchical statistical analysis strategy to form an environmental quality assessment data set;

[0043] The hierarchical statistical analysis strategy is combined with the meteorological quality identification system and the quality assessment indicators of actual rate, correct rate, error rate and suspicious rate at each level, and the EMI environmental meteorological condition assessment index is introduced for comprehensive quantitative analysis;

[0044] The EMI environmental meteorological condition assessment index represents a comprehensive index that characterizes the processes of aerosol emission, deposition, transmission, and diffusion under the influence of meteorological conditions. A larger EMI value indicates that the meteorological conditions are more unfavorable for the diffusion of atmospheric pollutants.

[0045] Through iterative statistical verification and threshold determination, the global statistical results are fused with the hierarchical statistical results, and finally a complete global statistical analysis data set is generated, which includes multi-dimensional statistical features, quality assessment identification and EMI environmental meteorological condition assessment index.

[0046] Furthermore, the global statistical analysis dataset is divided into meteorological quality assessment subsets by region and depth level; and the meteorological quality assessment subsets are further analyzed to obtain hierarchical statistical analysis results, including:

[0047] The data items in the global statistical analysis data set are layered in the time dimension, including grouping statistics by year, season, month, and day. Combined with the quality assessment marking results, a time-level meteorological quality assessment subset is formed to generate time series basic data.

[0048] The spatial dimension of each data item in the global statistical analysis data set is stratified, including grouping statistics by region, geographical features and elevation. The results are adjusted using regional characteristics to form spatial-level meteorological quality assessment subsets and generate spatial distribution data.

[0049] The data items in the global statistical analysis data set are layered by element dimension, and the quality assessment indicators of the meteorological elements of temperature, humidity, air pressure, and wind speed are calculated respectively. The hierarchical statistical analysis results at the element level are obtained, and the element distribution characteristic data is generated;

[0050] Comprehensive analysis is conducted using the EMI environmental meteorological condition assessment index to obtain hierarchical statistical analysis results;

[0051] By quantitatively characterizing the impact of meteorological conditions on the quality assessment of each meteorological quality assessment subset data item, a quantitative separation analysis is performed to obtain the quantitative separation analysis results;

[0052] The quantitative separation analysis comprises: based on the EMI environmental meteorological condition evaluation index, obtaining the emission change rate RE and the meteorological condition change rate RW by RE=(R1 / R0) / (E1 / E0)-1 and RW=(E1 / E0)-1; wherein R0 and R1 are respectively the measured meteorological data parameters of the comparison period, and E0 and E1 are respectively the EMI environmental meteorological condition evaluation index of the comparison period;

[0053] The contribution rate of meteorological condition changes to data quality is obtained through EMI environmental meteorological condition evaluation index analysis, so as to obtain the correlation between the EMI environmental meteorological condition evaluation index and meteorological observation data quality indicators, and quantitatively evaluate the impact of meteorological conditions on observation data quality;

[0054] Based on the spatiotemporal distribution characteristics of EMI, the changing trends of meteorological conditions in different regions and their impact on the quality of meteorological observation data are analyzed;

[0055] The two influence levels were integrated to obtain the quantitative separation analysis results.

[0056] Furthermore, the absolute and relative errors of the meteorological assessment data from manual and automatic stations are obtained based on the results of hierarchical statistical analysis, and the proportion of data within each error tolerance range is set. After iterative verification, meteorological quality characteristic data is obtained, including:

[0057] The error tolerance analysis is based on the hierarchical statistical analysis results and the EMI environmental meteorological condition evaluation index, and is implemented by constructing an error statistical function. The error statistical function receives the meteorological evaluation manual station observation data and the automatic station observation data as input parameters and performs the following operations:

[0058] The absolute error value and relative error percentage between the two sets of meteorological observation data were calculated based on the hierarchical statistical analysis results at the factor level;

[0059] The meteorological quality assessment subsets at the time and space levels were used to calculate the proportion of data within the relative error threshold of ±5% and the absolute error within the threshold of ±10%.

[0060] The standard deviation is calculated based on the EMI environmental meteorological condition evaluation index, and finally a statistical result list containing the error interval percentage and standard deviation of the absolute error value and relative error is returned. Based on the calculation results, an error interval ratio distribution table is generated and passed to the meteorological quality characteristic data as a key quantitative indicator.

[0061] Furthermore, based on the element distribution characteristic data and the error interval ratio distribution table, the ggplot2 package is used to dynamically bind data fields to obtain the probability density distribution characteristics of meteorological elements in different regions and levels. The data quality level is annotated in combination with the meteorological quality identification system to support the analysis of basic data and time series characteristics.

[0062] Using probability density distribution characteristics and time dimension layered data, combined with error statistics results, time series differences are generated to obtain the temporal variation trend of meteorological assessment data from manual and automatic stations. Error coding is used to distinguish data points within different error tolerance intervals, and the time series feature data is passed to spatial feature analysis.

[0063] Based on the data features obtained from time series feature analysis and spatial dimension layered data, combined with the error interval ratio distribution, the geospatial thermal effect is obtained to obtain the spatial distribution characteristics of the meteorological quality assessment index. The error mapping range is dynamically adjusted according to the EMI environmental meteorological condition assessment index, and the spatial distribution characteristics are stored in the data.

[0064] Meteorological quality characteristic data is obtained through analysis of different dimensions including probability density distribution, time series characteristic analysis and spatial characteristic analysis, and a comprehensive report on meteorological quality distribution characteristics in different regions and levels is generated. The time series difference analysis results of meteorological assessment manual station and automatic station data are integrated, and the generated data and acquired data are saved in corresponding folders according to the analysis dimensions; that is, the final meteorological quality assessment report is output by matching the output formats set by different users.

[0065] In a second aspect, a system for automatically evaluating the quality of meteorological observation data is provided, which is used to perform an automated quality evaluation method for meteorological observation data, comprising:

[0066] Data initialization module: used to collect multiple meteorological observation data and perform initial configuration to obtain meteorological assessment parameter sets;

[0067] Data preprocessing module: used to perform preprocessing operations on various parameters in the meteorological assessment parameter set to obtain the full meteorological assessment data set;

[0068] Statistical analysis module: used to perform global statistical analysis on the full meteorological assessment data set to obtain a global statistical analysis data set; divide the global statistical analysis data set into meteorological quality assessment subsets by region and depth level; and further analyze the meteorological quality assessment subsets to obtain hierarchical statistical analysis results;

[0069] Error analysis module: used to obtain the absolute and relative errors of meteorological assessment data from manual and automatic stations based on the results of hierarchical statistical analysis, set the proportion of data within each error tolerance range, and obtain meteorological quality characteristic data through iterative verification;

[0070] Report generation module: used to generate meteorological quality distribution characteristics of different regions and levels based on meteorological quality characteristic data, so as to obtain the time series difference between meteorological assessment manual station and automatic station data, and save the generated data and acquired data into folders respectively, and output the final meteorological quality assessment report by matching the output format set by different users.

[0071] The beneficial effects of the present invention are as follows: the present invention realizes the automation of the whole process from environment configuration to data processing through environment adaptive initialization configuration and intelligent dependency package management, thereby greatly improving work efficiency; by establishing a meteorological quality assessment identification system for multiple rounds of labeling and multi-dimensional consistency checks (climate limit value check, internal consistency check, time consistency check, spatial consistency check), it effectively reduces the problem of easy errors in manual operation and enhances the accuracy and reliability of data processing; the introduction of the EMI environmental meteorological condition assessment index realizes multi-level grouping statistical analysis of time dimension, space dimension and element dimension, which can quantitatively separate the impact of meteorological condition changes on data quality and provide more comprehensive and accurate multi-dimensional comprehensive analysis capabilities; through the path verification mechanism, format conversion adjustment, missing value supplementation processing and abnormal state capture, the problem of program interruption caused by file missing in traditional methods is solved. The system solves the problem of automatic fault tolerance and exception handling, and improves the stability and robustness of the system. It realizes data visualization based on the ggplot2 package, automatically generates standardized reports including probability density distribution, time series feature analysis and spatial distribution features, supports multiple output formats, and solves the technical problem of being unable to automatically generate standardized reports. By dynamically adjusting the computing resource allocation strategy of memory usage and thread allocation, adaptive optimization is performed according to the data scale and system resources, resource optimization and performance improvement are achieved, and the efficiency of large-scale meteorological data processing is improved. By setting the allowable interval of ±5% relative error threshold and ±10% absolute error threshold, combined with the iterative verification mechanism, the accurate quantification of the data difference between the meteorological assessment manual station and the automatic station is achieved, and finally the final meteorological quality assessment report is obtained, thereby improving the error quantification and quality assessment accuracy, and providing a scientific basis for meteorological data quality assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is a flow chart of a method for automated quality assessment of meteorological observation data provided in an embodiment of the present invention;

[0073] Figure 2 It is a schematic diagram of a module of an automatic quality assessment system for meteorological observation data provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0074] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0075] At least one embodiment of the present invention discloses a method and system for automated quality assessment of meteorological observation data, comprising:

[0076] like Figure 1 As shown, a method for automatically evaluating the quality of meteorological observation data includes the following steps:

[0077] Step 1: Collect multiple meteorological observation data and perform initial configuration to obtain a meteorological assessment parameter set;

[0078] Step 2: Preprocess the parameters in the meteorological assessment parameter set to obtain the full meteorological assessment data set;

[0079] Step 3: Perform global statistical analysis on the entire meteorological assessment dataset to obtain a global statistical analysis dataset;

[0080] Step 4: Divide the global statistical analysis dataset into meteorological quality assessment subsets by region and depth level; and further analyze the meteorological quality assessment subsets to obtain hierarchical statistical analysis results;

[0081] Step 5: Obtain the absolute and relative errors of the meteorological assessment data from the manual and automatic stations based on the hierarchical statistical analysis results, set the percentage of data within each error tolerance range, and obtain meteorological quality characteristic data through iterative verification.

[0082] Step 6: Generate meteorological quality distribution characteristics for different regions and levels based on meteorological quality characteristic data to obtain the temporal differences between meteorological assessment manual station and automatic station data, and save the generated data and acquired data into folders respectively. Output the final meteorological quality assessment report by matching the output formats set by different users.

[0083] Specifically: data initialization completes the operating environment configuration by automatically detecting and installing missing R packages (such as ggplot2 and readxl);

[0084] Data preprocessing: load xlsx format data, verify file integrity, perform missing value processing and format conversion;

[0085] Statistical analysis engine: Global statistics: calculate extreme values, mean, standard deviation, and coefficient of variation of the entire data set;

[0086] Hierarchical analysis: Divide data by region and depth level, and calculate the absolute error and the percentage of the tolerance interval (±5% / 10%);

[0087] Time series comparison: using the monitoring date as the horizontal axis, a line graph of the soil moisture difference between manual and automatic stations is generated;

[0088] Visualization: Use violin plots to show data distribution density and line charts to show time series differences;

[0089] Report generation: Automatically generate text / Word format reports based on templates, including statistical tables and chart references.

[0090] The above operation performs the following algorithm:

[0091] Automatic chart generation algorithm: Based on the ggplot2 package, dynamically bind data fields and batch output regional layered charts. Process control: users only need to set the input path, statistical parameters, and output format, and the main function (main.R) automatically schedules the execution of each module;

[0092] The output results include: statistical table (output_table.xlsx), charts (plots folder), and analysis report (.txt / .docx).

[0093] Example 1

[0094] This example provides a complete automated quality assessment method for meteorological observation data. This method, implemented in the R language, includes six core steps: environmental adaptive initialization, intelligent data preprocessing, multi-dimensional statistical analysis, error tolerance analysis, automated visualization, and intelligent report generation.

[0095] 1.1 The automated quality assessment method for meteorological observation data in this embodiment employs a modular design. This architecture primarily comprises five functional modules, corresponding to five core steps: data initialization module, data preprocessing module, statistical analysis module, error analysis module, and report generation module. Information transfer and functional collaboration between these modules are achieved through data flows, forming a complete data processing chain.

[0096] 1.2 Environment Adaptation Initialization Step: First, check whether the required packages are installed in the current R environment. The main packages that the system needs to check include: dplyr, tidyr, and data.table for data processing; readxl and openxlsx for data reading; ggplot2 and plotly for visualization; and rmarkdown and knitr for report generation.

[0097] By creating a list of required software packages and then checking each package one by one to see if it is installed, the system automatically installs the packages that are not installed. At the same time, the system creates an initial state system, including four status codes: correct (0), suspicious (1), error (2), and missing (3). Then, quality assessment indicators are constructed to calculate indicators such as the actual rate, correct rate, suspicious rate, and error rate. Finally, the computing resource allocation strategy is configured to dynamically adjust resource usage based on the system memory and the number of CPU cores. Through automated environment configuration, the problem of the traditional method requiring manual installation of dependent packages one by one is solved, greatly improving work efficiency; a standardized initial state system is established, providing a unified standard for subsequent data quality assessment; and the dynamic resource allocation strategy improves the efficiency of large-scale data processing.

[0098] 1.3 The intelligent data preprocessing step receives the output configuration of the environment adaptive initialization in Example 1.2, performs path legitimacy verification and file accessibility detection, reads meteorological observation data in Excel format, and performs format standardization conversion and missing value processing on the data.

[0099] First, the data path parameters set by the user are verified to check whether the file exists, whether it is readable, and whether the file format is supported. If the file does not exist, the system will record an error message and prompt the user to check the path settings, and try to search for files with the same name in adjacent directories; if the file is unreadable, the system will check the file permissions and try to modify the access permissions. If access is still not possible, the user will be prompted to contact the system administrator; if the file format is not supported, the system will provide format conversion suggestions and try to use general text reading methods for data parsing.

[0100] It then automatically selects the appropriate loading method based on the file format, supporting multiple formats such as xlsx, xls, and csv. For files in unknown formats, the system will attempt to read them in text format and perform intelligent format recognition. The loaded data is then format-checked to ensure that necessary fields such as site number, observation time, temperature, humidity, air pressure, and wind speed are included. When necessary fields are missing, the system generates field mapping suggestions and allows the user to manually specify the field correspondence. Format standardization conversion is then performed, converting the observation time to a date-time type and the meteorological elements to a numeric type. For data that fails to convert, the system records the exception information and marks it as an error state. Finally, missing value detection and processing are performed. Missing values ​​for observation time are directly deleted, and missing values ​​for meteorological elements are filled using a time series interpolation method. The interpolated data is marked as suspicious using a quality control code. When interpolation fails, the system maintains the original missing state and records the processing results in the log.

[0101] Through the path verification mechanism and exception handling, the problem of program interruption caused by file missing in traditional methods is solved; intelligent format conversion and missing value processing improve the degree of automation of data processing; four quality checks (climate limit value check, internal consistency check, time consistency check, spatial consistency check) effectively identify abnormal data and improve data quality.

[0102] 1.4 Multi-dimensional statistical analysis step: Receive the full meteorological assessment data set after preprocessing in Example 1.3 and execute a three-layer statistical analysis architecture: first, perform a global statistical analysis, then divide the data into subsets by region and depth level, and finally perform hierarchical statistical analysis on each subset.

[0103] First, a global statistical analysis is performed to calculate basic statistical indicators such as extreme values, means, standard deviations, and coefficients of variation for each meteorological element, establishing a global meteorological quality assessment benchmark. Then, a temporal stratification is performed, with statistical analysis grouped by year, season, month, and day. Combined with the quality assessment labeling results, a time-based meteorological quality assessment subset is formed. Next, a spatial stratification is performed, with statistical analysis grouped by site, forming a spatial-based meteorological quality assessment subset. Subsequently, a factor-based stratification is performed, with quality assessment indicators for each meteorological element calculated, including accuracy, error rate, and suspicious rate. Finally, an EMI environmental meteorological condition assessment index is calculated, which comprehensively represents the processes of aerosol emission, deposition, transmission, and diffusion under the influence of meteorological conditions.

[0104] Through multi-level group statistical analysis, a comprehensive analysis of the time dimension, space dimension and factor dimension is achieved; the introduction of the EMI environmental meteorological condition assessment index can quantitatively separate the impact of changes in meteorological conditions on data quality, providing a more comprehensive and accurate multi-dimensional comprehensive analysis capability.

[0105] 1.5 Error tolerance analysis step: Receive the hierarchical statistical analysis results in Example 1.4, calculate the absolute error and relative error of the meteorological assessment manual station data and the automatic station data, and count the proportion of data within each error tolerance interval.

[0106] First, an error statistics function was constructed, accepting observation data from both a manual meteorological assessment station and an automatic meteorological assessment station as input parameters. The absolute error and relative error percentage between the two sets of meteorological observations were then calculated using the results of a hierarchical statistical analysis at the element level. The percentage of data with relative errors within the ±5% threshold and absolute errors within the ±10% threshold was then calculated using the meteorological quality assessment subsets at the temporal and spatial levels. The standard deviation was then calculated based on the EMI environmental meteorological condition assessment index to analyze the error distribution characteristics under different EMI conditions. Finally, an iterative verification mechanism was used to verify and optimize the error statistics, generating a table of error interval ratio distributions.

[0107] By setting the allowable interval of ±5% relative error threshold and ±10% absolute error threshold, combined with an iterative verification mechanism, accurate quantification of the data differences between manual and automatic meteorological assessment stations is achieved; error analysis based on the EMI index can identify the degree of impact of meteorological conditions on data quality, thereby improving the accuracy of error quantification and quality assessment.

[0108] 1.6 Automated visualization generation step receives the error statistics results in Example 1.5 and the statistical analysis results in Example 1.4, dynamically binds data fields through the ggplot2 package, and automatically generates various types of visualization charts.

[0109] First, based on the element distribution characteristic data and the error interval ratio distribution table, the ggplot2 package dynamically binds data fields to generate probability density distribution characteristic charts for meteorological elements in different regions and hierarchies. The data quality levels are then labeled using the meteorological quality identification system. Then, using the probability density distribution characteristics and the temporal dimension-stratified data, combined with error statistics, a time series difference chart is generated to show the temporal trends of the meteorological assessment data from manual and automatic stations. Error coding is used to distinguish data points within different error tolerance intervals. Next, based on the data characteristics obtained from the time series feature analysis and the spatial dimension-stratified data, combined with the error interval ratio distribution, a geospatial thermal effect chart is generated to illustrate the spatial distribution characteristics of meteorological quality assessment indicators. The error mapping range is dynamically adjusted based on the EMI environmental meteorological condition assessment index. Finally, the generated visualization charts are saved in the corresponding folders according to the analysis dimension.

[0110] Data visualization is achieved based on the ggplot2 package, which automatically generates standardized charts containing probability density distribution, time series feature analysis, and spatial distribution features. Through visual elements such as color, size, and shape, data distribution, change trends, and spatial characteristics are intuitively displayed, solving the technical problem of being unable to automatically generate standardized visualizations.

[0111] 1.7 Intelligent report generation step integrates the visualization chart in Example 1.6 and the error analysis results in Example 1.5, and generates a complete meteorological quality assessment report according to the output format set by the user.

[0112] First, based on the assessment purpose and data complexity, an appropriate template is selected from a pre-set report template library. The template library includes three purpose types: brief, standard, and professional, and three complexity levels: simple, medium, and complex. Then, based on the error analysis results, key quality issues are automatically extracted, the quality status of each element is determined, and a report summary is generated, including the overall assessment results, the results of each element, and improvement suggestions. Next, based on the error analysis results and quality assessment indicators, standardized statistical result tables are generated, including error statistics, error interval distribution tables, and error statistics under EMI conditions. Maintenance recommendations and optimization plans are then provided for meteorological observation sites based on the error statistics and quality assessment indicators, and maintenance priorities are determined based on the error status of each site. Finally, the report summary, statistical tables, visualization charts, and maintenance recommendations are integrated into the selected template and rendered using the rmarkdown package according to the user's output format (Word, HTML, or PDF).

[0113] By automatically extracting key quality issues and outliers and proposing targeted improvement suggestions, the technical problem of being unable to automatically generate standardized reports is solved; multiple output formats are supported to meet the needs of different users; and targeted maintenance suggestions are provided for sites based on error statistics, providing a scientific basis for meteorological data quality assessment.

[0114] The six steps of this embodiment form a complete data processing chain: the environment adaptive initialization of embodiment 1.2 provides the basic environment and configuration for all subsequent steps; the intelligent data preprocessing of embodiment 1.3 receives the environment configuration, and outputs the preprocessed data for statistical analysis; the multi-dimensional statistical analysis of embodiment 1.4 receives the preprocessed data, and outputs the statistical results for error analysis and visualization; the error tolerance interval analysis of embodiment 1.5 receives the statistical results, and outputs the error analysis results for visualization and report generation; the automated visualization generation of embodiment 1.6 receives the statistical and error analysis results, and outputs charts for report generation; the intelligent report generation of embodiment 1.7 integrates the outputs of all previous steps to generate a final evaluation report.

[0115] The entire process from environment configuration to report generation is automated, eliminating the need for human intervention and significantly improving work efficiency. Exception capture and automatic recovery mechanisms ensure the system's stable operation in the face of various abnormal situations. The system can dynamically adjust processing strategies based on different data characteristics and system resources to adapt to various operating environments. Hierarchical statistical analysis across time, space, and elements provides a comprehensive perspective on data quality assessment. Multiple types of visual charts are automatically generated to intuitively display data quality distribution characteristics. Computing resource allocation strategies optimize memory usage and thread allocation to improve large-scale data processing efficiency.

[0116] Example 2

[0117] This embodiment describes in detail the specific implementation of the environmental adaptation initialization step in the method for automated quality assessment of meteorological observation data. This step is the foundation of the entire method, providing the necessary software environment and configuration parameters for subsequent data processing, and forms a continuation relationship with Example 1.

[0118] 2.1 Environment Adaptive Initialization First, check whether the required software packages are installed in the current R environment. This embodiment adopts an intelligent detection mechanism, which not only checks whether the software packages are installed, but also checks whether the versions meet the requirements.

[0119] First, define a list of required packages and their minimum version requirements. This includes data processing packages such as dplyr, tidyr, and data.table; data reading packages such as readxl and openxlsx; visualization packages such as ggplot2 and plotly; and report generation packages such as rmarkdown and knitr. Then, check each package to see if it is installed. If not, automatically install it. For installed packages, check whether their version meets the minimum requirements. If not, automatically update them. Next, load all required packages and record the installation and loading results. Finally, check to see if all packages have been successfully installed. If the installation fails, give the corresponding error message and return to the previous step to repeat the automatic installation process by updating the installation path.

[0120] 2.2 Environmental Adaptation Initialization also includes creating and defining quality assessment codes to mark the quality status of meteorological observation data. This embodiment defines four basic quality assessment codes and encapsulates them in a structured object.

[0121] First, the basic quality assessment codes are defined, including correct (0), suspicious (1), wrong (2), and missing (3). Then, the quality assessment status is defined, including excellent (only correct data), good (correct data and a small amount of suspicious data), general (correct, suspicious, and a small amount of wrong data), and poor (all types of data, including a large amount of missing and wrong data). Next, the quality assessment thresholds are set, including the accuracy threshold for good quality (0.9), the accuracy threshold for general quality (0.7), and the accuracy threshold for poor quality (0.5). Finally, the quality assessment code, assessment status, and control threshold are encapsulated into a complete quality assessment system object. A quality assessment code system with 0-3 assessment thresholds and four-level quality assessment status (excellent, good, general, and poor) is established, providing a unified standard for subsequent data quality assessment; the structured quality assessment system is easy to expand and maintain.

[0122] 2.3 Environmental Adaptation Initialization also includes constructing a quality assessment index calculation function for calculating the quality assessment index of meteorological observation data. This embodiment defines four basic quality assessment indicators: actual rate, correct rate, error rate, and suspicious rate.

[0123] First, create a function to calculate quality assessment indicators. This function accepts a dataset containing quality assessment fields as input. It then validates the input parameters to ensure the presence of the specified quality assessment fields in the data. Next, it calculates basic statistics, including the number of expected observations, the number of actual observations, the number of correct observations, the number of suspect observations, the number of erroneous observations, and the number of missing observations. Quality assessment indicators are then calculated, including the actual observation rate (number of actual observations divided by the number of expected observations), the correctness rate (number of correct observations divided by the number of actual observations), the suspectness rate (number of suspect observations divided by the number of actual observations), the error rate (number of erroneous observations divided by the number of actual observations), and the missing observation rate (number of missing observations divided by the number of expected observations). Finally, the quality assessment status of the dataset is determined based on the correctness rate, and the complete calculation results are returned. This function provides a standardized method for calculating quality assessment indicators, ensuring consistency and comparability of assessment results. Automated quality status determination reduces the risk of manual errors. The functional design facilitates reuse in subsequent steps.

[0124] 2.4 The last step of the environment adaptation initialization is to configure the computing resource allocation strategy to dynamically adjust memory usage and thread allocation according to the data scale and system resources.

[0125] First, system resource information, including available memory and the number of CPU cores, is obtained using appropriate methods for different operating systems. Resource allocation strategies are then dynamically adjusted based on the data size and system resources, setting a memory limit of 80% of available memory and a thread count equal to the number of CPU cores minus one (reserving one core for the system). The data block size is then set for batch processing of large-scale data. A parallel computing environment is then configured. If the system supports multithreading, a parallel computing cluster is created and the necessary variables and functions are exported. Finally, all resource configuration information is encapsulated into a resource policy object. By dynamically adjusting the computing resource allocation strategy for memory usage and thread allocation, adaptive optimization is performed based on data size and system resources, achieving resource optimization and performance improvement. The configuration of the parallel computing environment improves the efficiency of large-scale meteorological data processing.

[0126] First, R package detection and installation are performed to ensure that all required packages are correctly installed and loaded. Next, the initial state system is constructed to create a standardized quality assessment framework. Next, functions for calculating quality assessment metrics are constructed to provide a unified quality assessment methodology. Computing resource allocation strategies are then configured to optimize system performance. Finally, all initialization results are consolidated into an environment configuration object, containing package states, quality assessment systems, quality assessment functions, and resource strategies, for use in subsequent steps.

[0127] By integrating these four sub-steps, a complete environment-adaptive initialization process is formed, providing a unified foundation for all subsequent steps. The modular design facilitates maintenance and expansion, and the adaptive mechanism can adapt to various operating environments based on different operating systems and hardware configurations. Finally, the four steps are integrated into a complete environment-adaptive initialization function, providing the foundation for the intelligent data preprocessing in Example 3.

[0128] This embodiment serves as the basis of the entire system, and the environment configuration object it outputs will be directly used by the intelligent data preprocessing of Example 3, providing the necessary tools and standards for data loading, format conversion, and quality inspection. At the same time, the initial state system established by this embodiment will run through the multi-dimensional statistical analysis of Example 4, the error tolerance interval analysis of Example 5, and the intelligent report generation of Example 7 to ensure the consistency of the entire evaluation process. This makes it possible to automatically detect and install the required R software packages without manual intervention; ensure the normal operation of the system; create a complete initial state system and quality assessment index calculation function; dynamically adjust memory usage and thread allocation according to system resources to improve data processing efficiency; be able to adapt to various operating environments according to different operating systems and hardware configurations; the initial state system and quality assessment index calculation function are designed to be extensible, and new control codes and evaluation indicators can be added as needed.

[0129] Example 3

[0130] This example describes in detail the implementation of the intelligent data preprocessing step in the automated quality assessment method for meteorological observation data. This step receives the output of the adaptive initialization configuration in Example 2 and performs operations such as path verification, data loading, format conversion, and missing value processing, providing a standardized data foundation for the multi-dimensional statistical analysis in Example 4.

[0131] 3.1 Intelligent Data Preprocessing First, the data path parameters set by the user are verified to ensure that the data file exists and is accessible. At the same time, a try-catch mechanism is used to capture exceptions and automatically recover, improving the robustness of the system.

[0132] First, ensure the path is in a valid string format to ensure the correctness of the input parameters. Then, check whether the file in the specified path exists to avoid file-not-existing errors during subsequent processing. Next, attempt to open the file connection to verify that the file is readable and that file access permissions are correct. Then, check the file extension and verify the supported file formats (xlsx, xls, csv) to ensure that the data can be loaded correctly.

[0133] When a file doesn't exist or is inaccessible, the system automatically initiates a response strategy: first, it searches for files with the same name in the current directory and common data directories, automatically updating the path if found. Secondly, it provides a user interface to allow users to reselect the correct file path. Thirdly, it checks the network path's connection status and attempts to reestablish connections for network-stored data files. Finally, it logs detailed error information and suggested solutions to help users quickly locate and resolve path issues. Finally, it encapsulates the verification results into a path information object, containing the file path, format type, and verification status, for use in subsequent data loading steps.

[0134] 3.2 After the data path verification is passed, the next step is to load the data and detect its format. This embodiment supports multiple data formats (xlsx, xls, csv) and can automatically select the appropriate loading method based on the file format.

[0135] First, based on the file format verified in Example 3.1, the corresponding data loading method is automatically selected. For Excel files, the readxl package is used, and for CSV files, the read.csv function is used. After loading the data, it is verified that the data package is not empty and contains valid records. Next, the data format is checked to include the necessary columns (station number, observation time, temperature, humidity, air pressure, and wind speed) to ensure that the data structure meets the requirements for subsequent processing. An attempt is then made to handle column name mismatches, using an intelligent matching mechanism to automatically identify and convert common column name variations. If the primary loading method fails, the system attempts to reload using an alternative method to increase the success rate of data loading. Finally, the successfully loaded raw data is returned for use in the format standardization step. The intelligent data loading function can automatically select the appropriate loading method based on the file format; automatic column name matching reduces loading failures caused by data format differences; and the alternative loading mechanism improves the success rate of data loading and the system's fault tolerance.

[0136] 3.3 After the data is loaded successfully, it is necessary to perform format standardization conversion on the data to ensure that the data type of each field is correct and provide a standardized data format for the statistical analysis of Example 4.

[0137] First, the observation time field is converted to a standard date and time type. The system automatically tries various common date format patterns, including year-month-day hour:minute:second, year / month / day hour:minute:second, and other formats, to ensure that various date and time formats can be correctly identified and converted. Then, the meteorological element fields (temperature, humidity, air pressure, wind speed) are converted to numeric types to ensure the correctness of subsequent numerical calculations. Next, the site number is converted to a string type to maintain the consistency of the site identification. For any exceptions that occur during the conversion process, the system will log a warning message and attempt to retain the original value to avoid data loss due to format conversion failure. Finally, the standardized data set is returned, and all fields have the correct data type. Automated format conversion reduces the workload of manual processing; automatic recognition of multiple date formats improves the success rate of conversion; and standardized data format provides a reliable data foundation for subsequent statistical analysis.

[0138] 3.4 After the data format is standardized, missing values ​​need to be detected and processed to ensure the integrity of the data and provide high-quality data for the statistical analysis of Example 4.

[0139] First, the number of missing values ​​in each field is detected, the number of missing values ​​and the missing rate for each field are counted, and a missing value report is generated. Then, the missing values ​​of the observation time are processed. Since the observation time is a key time identifier and cannot be reasonably supplemented, records containing missing values ​​for the observation time are directly deleted. Next, the missing values ​​of the meteorological elements are processed. A corresponding quality assessment field is created for each meteorological element and marked using the initial state system established in Example 2. Subsequently, the meteorological element data for each site is grouped by site, and time series interpolation is performed on the meteorological element data. Missing values ​​are filled using the linear interpolation method, and the interpolated data is marked as suspicious. For missing values ​​that cannot be interpolated, the missing state mark is retained. Finally, the processed dataset and missing value report are returned, providing a basis for subsequent quality checks. The time series interpolation method can reasonably fill missing values ​​for meteorological elements; the initial state mark ensures the traceability of the interpolated data; and the missing value report provides an important reference for data quality assessment. Intelligent data preprocessing also includes four quality checks on meteorological data: climate limit value check, internal consistency check, temporal consistency check, and spatial consistency check. These checks can identify abnormal data and improve data quality.

[0140] First, a climate limit check is performed, setting reasonable ranges for each meteorological element (temperature -40°C to 50°C, humidity 0% to 100%, pressure 850hPa to 1050hPa, wind speed 0m / s to 60m / s). Data outside these limits is marked as erroneous, and data close to these limits is marked as suspicious. An internal consistency check is then performed to verify the physical relationships between different meteorological elements. For example, when the temperature is below 0°C, the humidity should not be 100%. Data that does not conform to physical laws is marked as suspicious. A temporal consistency check is then performed, calculating the magnitude of change in meteorological elements between adjacent moments and setting reasonable change thresholds (temperature 5°C, humidity 20%, pressure 10hPa, wind speed 10m / s). Data with anomalous changes is marked as suspicious. A spatial consistency check is then performed, comparing meteorological elements at different sites at the same time, calculating the mean and standard deviation. Data that deviates by more than three times the standard deviation is marked as suspicious. Finally, a dataset that has passed all four quality checks is returned, with all anomalous data appropriately flagged.

[0141] Through multiple rounds of labeling and multi-dimensional consistency checks, we effectively reduced the potential for manual errors and enhanced the accuracy and reliability of data processing. Four quality checks covered key aspects of meteorological data quality assessment, providing comprehensive data quality assurance. Finally, we integrated these five steps into a complete data preprocessing function, providing high-quality preprocessed data for the multi-dimensional statistical analysis in Example 4.

[0142] First, data path validation and exception handling are performed to ensure that data files are accessible. Next, data loading and format detection are performed to obtain raw data and verify the data structure. Next, data format standardization conversion is performed to ensure that all fields have the correct data type. Missing value detection and handling are then performed to improve data integrity. Finally, four quality checks are performed to identify and flag anomalous data. Exception catching mechanisms are implemented throughout the entire process to ensure clear error messages are provided when problems arise. Finally, complete preprocessing results are returned, including the processed dataset, a missing value report, and a processing log.

[0143] By integrating five sub-steps, a complete intelligent data preprocessing process is formed; the exception capture and automatic recovery mechanism ensures the stability of the system; and the high-quality preprocessed data lays a solid foundation for subsequent statistical analysis.

[0144] This embodiment receives the environment configuration object of Example 2 as input, and uses the initial state system and calculation function therein for data processing. The output of this embodiment will be directly passed to the multi-dimensional statistical analysis of Example 4, providing a standardized data basis for statistical calculations. At the same time, the quality assessment fields established in this embodiment will continue to be used in the error analysis of Example 5 and the visualization generation of Example 6. By automatically identifying file formats, column name matching, and format conversion, manual intervention is reduced; the success rate of data loading is improved through exception capture and backup methods; four quality checks are performed to comprehensively identify and mark abnormal data; scientific methods such as time series interpolation are used to handle missing values; unified data formats and quality tags provide a standard basis for subsequent analysis; detailed processing logs and quality tags ensure the traceability of the data processing process.

[0145] Example 4

[0146] This example describes in detail the implementation of the multi-dimensional statistical analysis step in the automated quality assessment method for meteorological observation data. This step receives the output of the intelligent data preprocessing in Example 3 and executes a three-layer statistical analysis architecture, providing a comprehensive statistical foundation for the error tolerance analysis in Example 5 and the automated visualization generation in Example 6.

[0147] 4.1 Multidimensional statistical analysis First, a global statistical analysis is performed on the preprocessed meteorological assessment full dataset to establish a global meteorological quality assessment benchmark.

[0148] First, basic statistical indicators are calculated for all meteorological elements (temperature, humidity, air pressure, and wind speed), including minimum, maximum, mean, standard deviation, and coefficient of variation. These indicators constitute a global meteorological quality assessment benchmark. A meteorological micro-disturbance filtering threshold is then set and dynamically and adaptively adjusted based on the meteorological regional characteristics and the preliminary quality assessment results of Example 3 to optimize the observation quality of each data item in the full meteorological assessment dataset. Based on the global meteorological quality assessment benchmark, a reference standard is then provided for subsequent hierarchical statistical analysis. Finally, the global statistical results are encapsulated into a global statistical object, containing the statistical indicators and quality benchmarks for each element, for use in hierarchical analysis. A unified quality assessment benchmark is established through global statistical analysis; the dynamic adaptive adjustment mechanism can optimize the assessment criteria based on regional characteristics, providing a scientific reference basis for subsequent hierarchical analysis.

[0149] 4.2 Based on the global statistical analysis, multi-level grouping statistical analysis of meteorological data is carried out according to the time dimension to obtain a hierarchical statistical analysis strategy at the time level.

[0150] First, the time dimension of each data item in the global statistical analysis data set is layered, and grouped and counted by year, season, month, and day. The system will automatically extract time elements such as year, season, month, and date from the observation time field, where the seasons are divided into spring from March to May, summer from June to August, autumn from September to November, and winter from December to February. Then, the quality assessment marking results established in Example 3 are combined to form a meteorological quality assessment subset of the time level, and generate time series basic data. Then, the average value of the meteorological elements at each time level is calculated, and the change characteristics of the meteorological elements at different time scales are analyzed. Finally, the time dimension layering results are integrated to provide a data basis for the time series feature analysis of Example 6. Through time dimension layered statistics, the changing laws of meteorological elements at different time scales are revealed; the time series basic data provides important support for subsequent time series analysis; and the multi-level time grouping meets the analysis requirements of different time scales.

[0151] 4.3 Conduct hierarchical statistical analysis on meteorological data according to spatial dimensions to obtain hierarchical statistical analysis strategies at the spatial level.

[0152] First, the data items in the global statistical analysis data set are spatially layered, and grouped and counted by region, geographic features, and elevation. The system groups the data according to the site number and calculates the statistical indicators of meteorological elements for each site. Then, the regional characteristic adjustment results of Example 3 are used to form a spatial-level meteorological quality assessment subset to generate spatial distribution data. Then, the differences in meteorological elements between different sites are analyzed to identify spatial distribution characteristics and abnormal sites. Finally, the spatial dimension layering results are integrated to provide a data basis for the spatial feature analysis of Example 6. Through spatial dimension layered statistics, the spatial distribution characteristics of meteorological elements are revealed; spatial distribution data provides an important basis for subsequent geographic spatial analysis; abnormal site identification helps to discover equipment failures or environmental problems.

[0153] 4.4 Conduct hierarchical statistical analysis on meteorological data according to element dimensions to obtain hierarchical statistical analysis strategies at the element level.

[0154] First, the data items in the global statistical analysis data set are layered according to the element dimension, and grouped and counted according to meteorological elements such as temperature, humidity, air pressure, and wind speed. The system will automatically identify the data characteristics of each meteorological element and calculate the quality assessment indicators of each element, including accuracy, error rate, suspicion rate, and missing rate. Then, combined with the quality assessment marking results established in Example 3, a meteorological quality assessment subset at the element level is formed to generate element feature data. Then, the quality differences between different elements are analyzed to identify the element types with prominent quality problems. Finally, the element dimension stratification results are integrated to provide an element-level quality foundation for the error tolerance interval analysis of Example 5. Through element dimension stratification statistics, the quality feature differences of different meteorological elements are revealed; the element feature data provides an important basis for subsequent element-level analysis; and the identification of quality problems helps to improve the observation quality of specific elements in a targeted manner.

[0155] 4.5 Based on the results of multi-dimensional statistical analysis, the environmental meteorological condition assessment index is calculated to provide a comprehensive indicator for meteorological data quality assessment.

[0156] First, key meteorological parameters, including the statistical characteristic values ​​of temperature, humidity, air pressure, and wind speed, are extracted from the hierarchical statistical results of the three dimensions of time, space, and elements. Based on meteorological principles and actual application requirements, the system assigns weight coefficients to each meteorological factor, with temperature weighted at 0.4, humidity weighted at 0.3, and wind speed weighted at -0.3 (a negative value indicates decreased comfort with increasing wind speed). Air pressure serves as a moderating factor. The EMI index is then calculated by site and time grouping using the following formula:

[0157] EMI = average temperature × 0.4 + average humidity × 0.3 + average wind speed × (-0.3);

[0158] Then, the EMI index is classified into four levels: excellent, good, fair, and poor, providing a basis for environmental condition classification for subsequent analysis. Finally, the EMI calculation results are correlated with the quality assessment results of Example 3 to analyze the data quality characteristics under different environmental conditions.

[0159] Through the calculation of the EMI index, the correlation between environmental conditions and data quality is established; the classification results provide an important basis for subsequent conditional analysis; the comprehensive indicators simplify the evaluation process of complex environmental conditions. Through hierarchical statistics in the three dimensions of time, space, and elements, a comprehensive data quality evaluation perspective is provided; standard statistical indicators and scientific calculation methods are used to ensure the reliability and comparability of the analysis results; the correlation between environmental conditions and data quality is established through the EMI index, laying the foundation for conditional analysis; the hierarchical statistical results are clearly layered, which facilitates the identification and analysis of quality problems at different levels; it provides a comprehensive statistical basis for the error tolerance interval analysis of Example 5 and the automated visualization generation of Example 6; the multi-dimensional statistical results provide a scientific basis and data support for quality management decisions.

[0160] In another preferred embodiment, various types of meteorological observation data (such as temperature, humidity, and wind speed) are labeled according to pre-set criteria (e.g., data integrity, rationality, and volatility). This labeling is performed by comparing historical normal value ranges and standard deviation offsets. The labeling takes the form of a quality assessment code, typically a numerical value or categorical label, indicating the "quality" of each piece of data. This allows for rapid identification of anomalies or unreliable data within observations, enhancing the trustworthiness of data usage.

[0161] Set numerical ranges corresponding to the four quality assessment statuses of Excellent, Good, Fair, and Poor; divide the quality assessment code into segments, for example: 90-100: Excellent; 70-89: Good; 50-69: Fair; 0-49: Poor. Each range corresponds to a clear quality status label, facilitating intuitive classification and management. This transforms the previously abstract quality assessment code into an interpretable and easy-to-use grading system, improving the readability and practicality of the assessment results and making them easier for users to understand and operate.

[0162] Establish a meteorological quality identification system. Based on the aforementioned digital ranges and status labels, establish a comprehensive meteorological data quality identification system and assign corresponding quality grade labels. This will create a standardized and unified quality evaluation system, facilitating automated classification, screening, alerting, and data cleansing within the system. Continuously monitor the quality status of meteorological data at different times and construct a time series. Analyze trends in these statuses over time, such as continued deterioration, fluctuations, and significant improvements. This provides real-time visibility into data quality dynamics, helping to promptly detect sensor failures or sudden environmental changes. This also provides decision support, such as when to calibrate equipment or remove outliers.

[0163] Import data and status trends into the R language environment and use loaded statistical analysis packages (such as ggplot2, dplyr, forecast, and changepoint) to conduct multi-dimensional analysis (e.g., time, space, and variables). Compare current data quality status with historical status to determine whether the trend is "smooth" or "dramatic." Based on the magnitude of the trend, determine the appropriate statistical strategy, such as weighted average, hierarchical clustering, or sliding window analysis. A highly automated and configurable analysis mechanism adapts to data variation characteristics in different scenarios, accurately identifying abnormal changes and formulating appropriate response strategies, such as focused monitoring, data supplementation, or equipment maintenance.

[0164] Example 5

[0165] This example describes in detail the implementation of the error tolerance analysis step in the automated quality assessment method for meteorological observation data. This step receives the output of the multi-dimensional statistical analysis in Example 4, calculates the absolute and relative errors between the manual and automated observation data, and calculates the percentage of data within each error tolerance. This provides the basis for error analysis in the automated visualization generation in Example 6.

[0166] 5.1 Error tolerance analysis First, the observation data from the manual and automatic stations need to be matched in time and space to ensure the validity of the comparison.

[0167] First, the data of the manual station and the automatic station are time-synchronized and verified to ensure that the two sets of data are observed at the same time point. The system will automatically identify the time field format, convert it into a standard time format, and handle the time deviation problem. Then spatial position matching is performed to ensure that the data compared are from the same observation point based on the site number or geographic coordinates. The matched data is then quality-screened, and the data points marked as erroneous or missing in Example 3 are eliminated, and the correct and suspicious data are retained for error analysis. Finally, the data is standardized to ensure that the dimensions and accuracy of each meteorological element are consistent, providing a standardized data basis for subsequent error calculations. Through strict data matching and preprocessing, the scientific nature and accuracy of error analysis are ensured; the quality screening mechanism avoids the interference of abnormal data on error statistics; and the standardization process ensures the comparability of errors between different elements.

[0168] 5.2 Based on the matched data, calculate the absolute error of each meteorological element and evaluate the observation deviation of the automatic station relative to the manual station.

[0169] First, calculate the absolute error for each matched data pair using the formula:

[0170] Absolute error = |Automatic station observation value-manual station observation value|;

[0171] The system processes each meteorological element (temperature, humidity, air pressure, and wind speed) individually to ensure calculation integrity. It then performs a statistical analysis of the absolute error, calculating statistical indicators such as mean absolute error, absolute error standard deviation, maximum absolute error, and minimum absolute error. Combining the EMI environmental condition classification described in Example 4, the absolute error characteristics under different environmental conditions are analyzed to identify the impact of environmental factors on observation errors. Finally, the absolute errors are grouped and statistically analyzed by station and time, providing detailed error distribution data for subsequent interval analysis. Absolute error calculation quantifies the observation differences between automatic and manual stations; statistical analysis reveals the distribution characteristics and variation patterns of the errors; and environmental correlation analysis identifies environmental factors that affect observation accuracy.

[0172] 5.3 Based on the absolute error, calculate the relative error of each meteorological element and evaluate the relative degree of observation deviation.

[0173] First, the relative error is calculated for each matched data pair. The formula is:

[0174] Relative error = │Automatic station observation value - Manual station observation value│ / │Manual station observation value│×100%;

[0175] The system will handle the special case where the denominator is zero, using an alternative calculation method or marking it as invalid data. The relative error is then statistically analyzed to calculate statistical indicators such as the average relative error and the relative error standard deviation. The allowable interval of the relative error is then set, usually including multiple interval levels such as ±5%, ±10%, ±15%, ±20%, and the proportion of data in each interval is counted. Finally, combined with the hierarchical statistical results of Example 4, the relative error distribution characteristics under different time, space, and factor dimensions are analyzed to provide a multi-dimensional error perspective for quality assessment. Through relative error calculation, a standardized error assessment index is provided; interval statistics reveal the grade distribution of data quality; and multi-dimensional analysis identifies the error change law under different conditions.

[0176] 5.4 Based on the calculation results of absolute error and relative error, the proportion of data within each error tolerance range is counted and a data quality level assessment system is established.

[0177] First, a multi-level error tolerance interval is set, including an absolute error interval (such as ±0.5, ±1.0, ±2.0, etc.) and a relative error interval (such as ±5%, ±10%, ±20%, etc.). The system will dynamically adjust the interval threshold according to the observation accuracy requirements of different meteorological elements. Then, the number and proportion of data points in each interval are counted, the cumulative distribution function is calculated, and the distribution characteristics of the error are analyzed. Next, a data quality grade assessment standard is established, and data with a relative error within the range of ±5% is defined as excellent, ±5% to ±10% as good, ±10% to ±20% as average, and more than ±20% as poor. Finally, an error tolerance interval distribution table is generated to provide a quantitative quality assessment basis for the intelligent report generation of Example 7. Through error tolerance interval statistics, a quantitative data quality assessment standard is established; the grade assessment system facilitates intuitive understanding of the quality status; and the distribution analysis provides targeted guidance for quality improvement.

[0178] 5.5 Combined with the EMI environmental condition evaluation results of Example 4, the error characteristics under different environmental conditions are analyzed to identify the impact of environmental factors on observation accuracy.

[0179] First, the error analysis results are associated with the EMI environmental condition classification results, and the error data are grouped according to the EMI level (excellent, good, general, poor). The system calculates the average absolute error, average relative error, and error tolerance distribution under each EMI level. Then, the correlation between environmental conditions and observation errors is analyzed to identify the influence of factors such as extreme weather conditions, seasonal changes, and day and night differences on observation accuracy. Then, an environmental condition-error feature association model is established to provide a reference standard for data quality assessment under different environmental conditions. Finally, an error analysis report under environmental conditions is generated to provide an environmental factor analysis basis for the quality management system effectiveness assessment of Example 8.

[0180] Error analysis under environmental conditions reveals the mechanism by which environmental factors influence observation accuracy. The correlation model provides a scientific basis for conditional quality assessment. The analysis results provide guidance for improving the environmental adaptability of the observation system. Through the dual calculation of absolute and relative errors, the accuracy level of the observation system is comprehensively quantified. A scientific quality assessment standard is provided based on the multi-level tolerance intervals set based on the meteorological observation accuracy requirements. Error analysis combined with EMI environmental conditions reveals the influence of environmental factors on observation accuracy. The established data quality rating assessment system facilitates intuitive understanding of quality status and management decisions. Error characteristics are analyzed from multiple dimensions, including time, space, elements, and environment, providing a comprehensive quality assessment perspective. This provides a detailed error analysis foundation for the automated visualization generation of Example 6 and the intelligent report generation of Example 7.

[0181] Another preferred embodiment captures abnormalities in the dynamic changes of meteorological observation data: Time series analysis methods (such as sliding windows, trend identification, and anomaly detection algorithms) are used to monitor the changing trends of meteorological observation data at different times in real time. Points that deviate significantly from the normal trend, such as sudden increases or decreases, are identified as abnormal data. Statistical indicators such as the Z-score, IQR (interquartile range), and EWM (exponentially weighted moving average) can be used to determine the abnormality threshold. This improves the system's sensitivity to "sudden anomalies" or "potential fault signals," providing a prerequisite for further precise quality assessment and correction.

[0182] Perform temporal and spatial interpolation analysis on abnormal data, extract the sudden change amplitude index, and mark it as a secondary quality code:

[0183] Time series characteristic analysis: Calculate the rate of change or degree of deviation based on continuous observations before and after an anomaly (e.g., a sudden drop of 8°C in temperature within an hour);

[0184] Spatial correlation analysis: Using data from surrounding meteorological stations, estimate the expected value of the point using spatial interpolation algorithms (such as Kriging, IDW, and inverse distance weighted).

[0185] Sudden change amplitude index: combines temporal mutations and spatial differences to generate a comprehensive index to measure the severity of anomalies;

[0186] This index is used to assign a new quality assessment code to abnormal data (e.g., "moderate abnormality," "serious abnormality," etc.), resulting in a secondary labeling result. Anomaly detection is no longer a single-point judgment, but rather an assessment across multiple dimensions of time and space. This improves the accuracy and scientific nature of abnormal data judgments and provides a quantitative basis for subsequent inspections or corrections.

[0187] Four consistency and extreme value checks are performed to generate three-fold labeling results; four verification checks are performed on the secondary labeled abnormal data: climate limit value check: comparison with historical extreme values ​​or long-term climate average range (for example, the winter temperature in Beijing should not be higher than 20°C);

[0188] Internal consistency check: check the logical relationship between different meteorological elements (e.g. humidity should not be 0% and precipitation should be heavy rain at the same time);

[0189] Time consistency check: consistency between the current observation value and the adjacent time data;

[0190] Spatial consistency check: reasonableness comparison between the observations of the current site and those of neighboring sites.

[0191] Any data that violates the above inspection rules will be marked with an updated quality assessment code, resulting in a three-fold marking result. This will enable a multi-level, systematic error elimination mechanism, significantly reducing the risk of "misjudging good data" or "omitting bad data," and enhancing the credibility and rigor of data processing.

[0192] Fitting the observation data with normal or bad status: Process the observation data identified as "normal" or "bad" in the three marks;

[0193] Fitting methods may include: interpolation (e.g., spatial interpolation, time series interpolation); curve fitting (e.g., polynomial regression, sliding average); and model re-estimation (e.g., modeling and prediction based on surrounding stations). The goal is to restore or estimate a more likely true value to replace the original anomalous observation. This reduces analytical bias caused by missing or anomaly data, maintains data continuity to facilitate downstream model processing, and reduces the risk of local anomalies affecting the overall observation system.

[0194] After integration, the processed and excellent data are combined to form preliminary quality assessment results for global analysis: the "fitted data" from the previous step are merged with the observation data originally marked as "excellent" or "good" in the previous round of evaluation; a set of "preliminary quality assessment data sets" is generated; this data set is then transmitted to the "global statistical analysis system" for comprehensive modeling analysis, such as trend modeling and extreme event identification. This enables the automatic construction of high-quality data sets, ensures that the input of downstream analysis models has maximum integrity and minimum error, and supports key application scenarios such as weather forecasting, abnormal event monitoring, and policy formulation. Summary of the overall process advantages:

[0195] Link Key Objectives Effective means Bring results Exception Capture Dynamic monitoring of mutations Timing Detection Detect abnormalities promptly Secondary marking Determine the severity of the abnormality Spatiotemporal interpolation + amplitude index Anomaly intensity quantification Three marks Data consistency verification Logical / historical rule verification Multi-layer screening ensures accuracy Fitting correction Abnormal data recovery Model fitting / estimation Maintain data continuity Results integration Building a clean dataset Screening + Fusion For use by the global analysis system

[0196] Example 6

[0197] This example describes in detail the implementation of the automated visualization generation step in the automated quality assessment method for meteorological observation data. This step receives the output of the error tolerance analysis in Example 5 and automatically generates various types of visualization charts through dynamic data binding technology, providing intuitive chart presentation for the intelligent report generation in Example 7.

[0198] 6.1 Automated visualization generation first requires establishing a dynamic binding mechanism between data fields and chart elements to achieve automated chart generation.

[0199] First, the error analysis results of Example 5 are parsed for data structure, and the type, range and distribution characteristics of the data field are automatically identified. The system will establish a field mapping table to map the meteorological element field to the X-axis, Y-axis, color, size and other visual elements of the chart. Then, the appropriate chart type is automatically selected according to the data characteristics. Continuous numerical data prefers scatter plots or line charts, categorical data prefers bar charts or pie charts, and distribution data prefers box plots or violin charts. Then, a dynamic binding function is established to automatically configure chart parameters and style settings based on the meteorological elements and visualization requirements specified by the user. Finally, a batch generation mechanism for charts is implemented to generate corresponding visualization charts for multiple meteorological elements in parallel, thereby improving generation efficiency. Through the dynamic binding mechanism, automation and standardization of chart generation are achieved; intelligent type selection ensures the best match between charts and data features; and the batch generation mechanism greatly improves the generation efficiency of visualization.

[0200] 6.2 Based on the dynamic binding framework, a violin plot is generated to display the distribution characteristics and density information of meteorological element data.

[0201] First, the observation data of each meteorological element is extracted from the error analysis results and grouped by station or time. The system calculates the probability density distribution of each set of data and uses the kernel density estimation method to smooth the data distribution curve. Then, the graphical elements of the violin plot are constructed, including the symmetrical mirror image of the density curve, quartile marks, median line, and outlier points. Next, the visual style of the chart is set, including color scheme, transparency, line thickness, etc., to ensure the aesthetics and readability of the chart. Finally, auxiliary information such as chart title, coordinate axis labels, legend, etc. is added to generate a complete violin plot, which intuitively displays the distribution shape, central tendency, and degree of dispersion of the data. Through the generation of the violin plot, the distribution characteristics and density information of the data are intuitively displayed; the kernel density estimation method provides a smooth distribution curve; and the reasonable configuration of visual elements enhances the readability and aesthetics of the chart.

[0202] 6.3 Generate box plots to display the statistical characteristics and outlier distribution of meteorological element data.

[0203] First, the five-number summary statistics of each meteorological element data are calculated, including the minimum value, first quartile, median, third quartile, and maximum value. The system will identify outliers and use the 1.5 times interquartile range rule to determine the threshold for outlier judgment. Then, the graphical elements of the box plot are constructed, including the box (representing the interquartile range), median line, whiskers (representing the normal value range), and outlier points. Then, multiple box plots are generated according to different grouping dimensions (such as site, time, and environmental conditions) to facilitate comparative analysis. Finally, the color coding and annotation information of the chart are set, using different colors to distinguish different groups, and adding numerical labels to provide accurate statistical information. Through the box plot generation, the statistical characteristics of the data and the distribution of outliers are clearly displayed; the five-number summary provides key information on the data distribution; and the group comparison function makes it easy to identify data differences under different conditions.

[0204] 6.4 Generate a time series difference line chart to show the changing trends and difference characteristics of the observation data of the manual station and the automatic station over time.

[0205] First, extract the time series error data from the error analysis results of Example 5, and arrange the observation differences between the manual station and the automatic station in chronological order. The system will handle missing values ​​and outliers in the time series, and use interpolation or smoothing methods to ensure the continuity of the time series. Then construct a dual-axis line chart, with the primary axis showing the time change of the observation difference and the secondary axis showing the percentage change of the relative error. Then add trend lines and confidence intervals, use regression analysis methods to fit long-term trends, and calculate confidence intervals to reflect the uncertainty of the data. Finally, set the scale and label of the time axis, automatically select the appropriate time interval according to the time span of the data, and add annotations for important time nodes. The time series difference line chart is generated to intuitively show the time change pattern of the observation error; the dual-axis design simultaneously displays the absolute difference and relative error; and the trend analysis reveals the long-term change trend and periodic characteristics of the error.

[0206] 6.5 Based on the geographic coordinate information of the site, generate a geospatial heat map to show the spatial distribution characteristics of meteorological element errors.

[0207] First, obtain the geographic coordinate information of each observation site, including longitude, latitude and altitude. The system will associate the error analysis results of Example 5 with the site coordinates, and calculate the average absolute error of each site as the numerical basis of the heat map. Then construct a geographic base map, select a suitable map projection and scale according to the geographic scope of the observation area, and add geographic elements such as administrative boundaries, rivers, terrain, etc. Then apply the spatial interpolation algorithm, and use Kriging interpolation or inverse distance weighted method to interpolate the discrete site error data into a continuous spatial distribution surface. Finally, set the color mapping scheme of the heat map, use gradient colors to represent the size of the error, add color legends and contour lines, and generate a complete geospatial heat map. Through the generation of geospatial heat maps, the spatial distribution pattern of observation errors is intuitively displayed; spatial interpolation technology provides a continuous error distribution surface; and the geographic base map enhances the recognition and understanding of spatial locations.

[0208] 6.6 Integrate the various visual charts generated and organize the output results according to meteorological elements and chart types.

[0209] First, a chart organization structure is established, creating a classification directory based on meteorological elements (temperature, humidity, air pressure, and wind speed). Each element includes a variety of chart types, such as violin plots, box plots, time series difference line charts, and geospatial heat maps. The system generates standardized file names and metadata for each chart, including chart type, element name, generation time, etc. Then, the output format and quality parameters of the chart are set, supporting multiple formats such as PNG, PDF, and SVG. Parameters such as resolution, size, and compression ratio are set to ensure chart quality. Next, a chart index and directory structure are established, and a chart list file is generated to record the path, description, and association relationship of each chart. Finally, a batch output function is implemented, saving all generated charts to a specified directory according to the organizational structure, providing a complete visualization material library for the intelligent report generation of Example 7.

[0210] Through chart integration and output, a standardized visualization library has been established; the clear organizational structure facilitates the search and use of charts; and multi-format support meets the needs of different application scenarios. Through dynamic data binding technology, a fully automated generation process from data to charts is achieved; violin plots, box plots, time series line charts, geographic heat maps and other visualization forms are provided to fully display data characteristics; through scientific color mapping, reasonable chart layout and clear annotation information, the aesthetics and readability of the charts are ensured; geographic spatial heat maps intuitively display the spatial distribution pattern of errors, providing a powerful tool for spatial analysis; a standardized chart organization structure and output format have been established to facilitate subsequent report generation and data management; a rich set of visualization charts provide intuitive visual support for data quality assessment and decision analysis.

[0211] Example 7

[0212] This example describes in detail the implementation of the intelligent report generation step in the automated quality assessment method for meteorological observation data. This step receives the output of the automated visualization generated in Example 6 and the results of the error tolerance analysis in Example 5, and automatically generates a complete meteorological quality assessment report using templated report generation technology. This provides the reporting basis for the quality management system effectiveness assessment in Example 8.

[0213] 7.1 Intelligent report generation first requires selecting an appropriate report template based on the assessment purpose and data complexity, and performing personalized configuration.

[0214] First, a hierarchical report template library is established, including three purpose types: briefing type, standard type, and professional type. Each type is further divided into three complexity levels: simple, medium, and complex, with a total of nine template combinations. The system will automatically evaluate the complexity of the data based on the scale of the input data, the number of elements, the analysis dimension and other characteristics. Then, according to the user-specified evaluation purpose or the system's intelligent recommendation, the most suitable report template is selected. Then, the selected template is personalized, including the setting of basic information such as the report title, generation date, analysis cycle, and involved sites. Finally, the integrity and compatibility of the template are verified to ensure that the template file exists and is compatible with the current data structure, in preparation for subsequent report generation. Through intelligent template selection, the best match between report content and evaluation needs is ensured; the hierarchical template library meets the reporting needs of different scenarios; and personalized configuration improves the pertinence and practicality of the report.

[0215] 7.2 Based on the error analysis results of Example 5, key quality issues and outliers are automatically extracted and a report summary is generated.

[0216] First, key quality indicators are extracted from the error tolerance analysis results, including the mean absolute error of each meteorological element, the percentage within a 5% relative error range, and the quality grade distribution. The system automatically determines the quality status of each element based on pre-set quality assessment criteria, defining a percentage within a 5% relative error range of 90% or higher as excellent, 70%-90% as good, 50%-70% as fair, and below 50% as poor. The overall quality status is then calculated, using a weighted average method to combine the quality scores of each element to generate an overall assessment conclusion. Key quality issues are then identified, using threshold comparison and anomaly detection algorithms to automatically identify elements, sites, and time periods with poor quality. Finally, targeted improvement recommendations are generated, selecting appropriate improvement measures from a pre-set library based on the type and severity of the quality issue, to form a comprehensive report summary. Automatic summary generation quickly extracts the core information from the assessment results; intelligent problem identification improves the accuracy of quality issue detection; and targeted recommendations provide specific guidance for quality improvement.

[0217] 7.3 Based on the error analysis results and quality assessment indicators, a standardized statistical result table is automatically generated.

[0218] First, a standardized table template is designed, including various types such as error statistics table, error interval distribution table, and error statistics table under EMI conditions. The system will extract the corresponding data fields from the analysis results of Example 5 and organize and format the data according to the requirements of the table template. Then data verification and cleaning are performed to check the integrity and consistency of the data, handle missing values ​​and outliers, and ensure the accuracy of the table data. Then, table style and format settings are applied, including header style, data alignment, numerical accuracy, unit annotation, etc., to improve the readability and professionalism of the table. Finally, tables in multiple output formats are generated, supporting formats such as HTML, Word, Excel, etc. to meet the needs of different users. Through automatic table generation, the standardized display of statistical results is ensured; the data verification mechanism ensures the accuracy of the table content; and multi-format support improves the applicability and convenience of the table.

[0219] 7.4 Based on error statistics and quality assessment indicators, maintenance recommendations and optimization plans are automatically generated for each meteorological observation station.

[0220] First, error statistics are calculated for each element at each site, including mean absolute error, standard deviation of error, and frequency of outliers. The system then compares each site's error level with the network-wide average to identify sites and elements with elevated errors. A maintenance priority assessment model is then established, comprehensively considering factors such as error size, data importance, and maintenance cost to calculate a maintenance priority score for each element at each site. Based on the priority score and error characteristics, corresponding maintenance measures are matched from a pre-set maintenance recommendation library, including different types of recommendations such as equipment calibration, environmental improvement, and equipment replacement. Finally, a structured maintenance recommendation table is generated, containing information such as site number, element name, error level, priority, and specific recommendations, providing decision support for maintenance management. Automatic maintenance recommendation generation provides a scientific basis for site management decisions; the priority assessment model ensures the rational allocation of maintenance resources; and the structured recommendation table facilitates the development and execution of maintenance plans.

[0221] 7.5 Integrate the visualization chart generated in Example 6 into the report to form a complete report with both pictures and text.

[0222] First, various chart files are read from the visualization chart library of Example 6, including violin plots, box plots, time series line charts, geographic heat maps, etc. The system will determine the position and size of each chart in the report according to the layout requirements of the report template. Then the format conversion and size adjustment of the chart are performed to ensure the compatibility of the chart with the report format, and unify the resolution and color mode of the chart. Then, a title, explanatory text and data source information are added to each chart to enhance the comprehensibility and traceability of the chart. Finally, an association relationship is established between the chart and the text content, and a reference and analysis description of the chart is inserted into the report text to form a complete report content with both pictures and text. Through chart integration, the intuitiveness and persuasiveness of the report are enhanced; the unified format ensures the professionalism and aesthetics of the report; and the association between pictures and text improves the logic and readability of the report content.

[0223] 7.6 According to user needs, the complete report content will be output into documents in various formats.

[0224] First, the various components of the report are integrated, including the summary, statistical tables, maintenance recommendations, and visual charts, to form a complete report content structure. The system will select the appropriate rendering engine and conversion tool based on the user-specified output format (Word, PDF, HTML, etc.). Format-specific style settings are then performed, including page layout, font style, chart embedding method, etc., to ensure the consistency and aesthetics of reports in different formats. The report rendering and generation process is then executed to convert the structured report content into the final document format, handling technical details such as chart embedding, page pagination, and directory generation. Finally, quality inspection and file output are performed to verify the integrity and correctness of the generated document, save the final report to the specified location, and return the file path information.

[0225] Multi-format output meets the needs of diverse users; automated generation significantly improves report production efficiency; and quality inspection mechanisms ensure report integrity and accuracy. The entire process, from template selection to content generation, is intelligent, significantly reducing manual intervention. Comprehensive reports cover summaries, statistical analysis, maintenance recommendations, and visual charts. Standardized templates and professional style settings ensure the standardization and professionalism of reports. Individual configuration and customization are supported based on different assessment objectives and data characteristics. Multiple output formats, including Word, PDF, and HTML, are supported to meet the needs of diverse application scenarios. Automatic problem identification and suggestion generation provide strong support for quality management decisions.

[0226] Example 8

[0227] This example describes in detail the implementation of the quality management system effectiveness assessment step within the automated quality assessment method for meteorological observation data. This step receives the output from the previous examples and comprehensively evaluates the effectiveness of the quality management system for meteorological observation services using a multi-dimensional indicator system, forming a complete quality management closed loop.

[0228] 8.1 Evaluate the stable operation rate and business continuity of meteorological observation equipment by analyzing equipment operation logs.

[0229] First, key information is extracted from the device operation log, including fields such as site number, device number, operating status, start time, and end time. The system preprocesses the log data, standardizing the time format, cleaning abnormal records, and supplementing missing information. Next, the system calculates the operating time statistics for each device, including total monitoring time, uptime, and downtime, and calculates the availability index (availability = uptime / total monitoring time). The device availability index is then aggregated by site, calculating the average, minimum, and maximum availability for each site to identify sites and devices with low availability. Finally, an availability evaluation standard is established, defining an availability above 95% as excellent, 90%-95% as good, 85%-90% as fair, and below 85% as poor. A service availability evaluation report is then generated. This service availability evaluation quantifies the stability of device operations. The site-level summary analysis facilitates the identification of problematic devices and sites. The evaluation standard provides clear quality targets for device management.

[0230] 8.2 Evaluate the integrity, timeliness, and reliability of meteorological data transmission by analyzing data transmission logs.

[0231] First, transmission records are extracted from the data transmission log, including information such as site number, transmission time, transmission status, data volume, and delay time. The system defines transmission quality assessment criteria, defining records with a successful transmission status and a delay time of no more than 5 minutes as successful real-time transmission. Transmission quality indicators are then calculated for each site, including total transmission times, successful transmission times, real-time transmission times, transmission success rate, real-time transmission rate, and data integrity rate. The temporal trends in transmission quality are then analyzed to identify time periods and possible causes of degraded transmission quality. Finally, a transmission quality assessment grade is established, defining a transmission success rate of over 98% and a real-time transmission rate of over 95% as excellent, providing an assessment basis for optimizing the data transmission system. Through transmission quality assessment, a comprehensive understanding of the operating status of the data transmission system is obtained; temporal trend analysis helps identify systemic problems; and the assessment grade provides a clear direction for improving the transmission system.

[0232] 8.3 Evaluate equipment failure frequency and maintenance response efficiency by analyzing fault logs and maintenance records.

[0233] First, fault log and maintenance record data are integrated to establish a fault-repair relationship, matching the fault occurrence time and repair processing time for the same equipment. The system then calculates key maintenance metrics, including response time (the time interval from fault occurrence to repair start) and repair time (the time interval from repair start to repair completion). Fault repair metrics are then aggregated by equipment and site, calculating statistics such as fault frequency, average response time, and average repair time. Next, the distribution characteristics of fault type and severity are analyzed to identify high-frequency fault types and high-impact fault modes. Finally, fault repair evaluation criteria are established, defining excellent maintenance efficiency as an average response time of less than 2 hours and an average repair time of less than 8 hours, providing an evaluation basis for optimizing maintenance management. Through fault repair evaluation, the efficiency level of maintenance management is quantified; fault type analysis provides important reference for preventive maintenance; and the evaluation criteria point the way to improvement in maintenance processes.

[0234] 8.4 Evaluate user satisfaction with meteorological data quality and services by analyzing user survey data.

[0235] First, evaluation data is extracted from user surveys, including information such as user type, data quality score, service quality score, and overall satisfaction score. The system standardizes the rating data, unifies the rating scale, and handles missing and outliers. Satisfaction indicators are then calculated by user type, including average score, satisfaction rate (percentage of users scoring 4 or above), and dissatisfaction rate. Next, factors influencing satisfaction are analyzed, and correlation analysis is used to identify key factors influencing user satisfaction, such as data timeliness, accuracy, and service response speed. Finally, a user satisfaction rating scale is established, defining an average satisfaction score of 4.0 or above as excellent and 3.5-4.0 as good. This provides a basis for user feedback to improve service quality. Through user satisfaction assessments, we understand users' true perceptions of service quality; the analysis of influencing factors provides targeted guidance for service improvements; and the rating scale provides evaluation criteria from a user perspective for service quality management.

[0236] 8.5 This example describes in detail the implementation of the quality management system effectiveness assessment step within the automated quality assessment method for meteorological observation data. This step quantitatively evaluates the effectiveness of the quality management system for meteorological observation services using multiple metrics, including service availability, data transmission quality, fault and repair status, user satisfaction, and external supplier evaluation. By analyzing external supplier evaluation data, the quality of meteorological equipment, services, and system platforms provided by external suppliers is assessed.

[0237] First, evaluation records are extracted from supplier evaluation data, including supplier ID, supplier name, supply type, product quality score, service quality score, delivery timeliness score, and overall evaluation. The system standardizes the evaluation data, unifies the rating scale, and handles missing and outliers. Evaluation metrics are then calculated for each supplier dimension, including the number of evaluations, average product quality score, average service quality score, average delivery timeliness score, average overall rating, and the pass rate for each metric (the percentage of suppliers with a score of 3.5 or above). Evaluation metrics are then grouped and statistically analyzed by supply type, calculating the number of suppliers and the average rating for each supply type. Finally, an external supplier evaluation rating system is established, defining suppliers with an average overall rating of 4.0 or above as excellent and those with an average rating of 3.5-4.0 as good, providing an evaluation basis for supplier management and selection. External supplier evaluations provide a comprehensive understanding of the product and service quality of external suppliers. Categorized statistics help identify high-quality and problematic suppliers. The evaluation rating system provides objective criteria for supplier management and selection.

[0238] 8.6 A complete quality management system effectiveness evaluation system is formed by integrating multi-dimensional indicators such as business availability, data transmission quality, fault repair, user satisfaction and external supplier evaluation.

[0239] First, a quality management system effectiveness assessment framework is established, defining the weights and importance of each assessment dimension. The system then sequentially executes five sub-assessment modules, including service availability, data transmission quality, fault and repair status, user satisfaction, and supplier evaluation. Each sub-module is conditionally executed based on the availability of input data, ensuring flexibility and adaptability. Key indicators for each sub-module are then extracted, including service availability, data transmission success rate, data integrity rate, average response time, average repair time, overall user satisfaction, data quality satisfaction rate, and overall supplier evaluation. A comprehensive quality management system effectiveness index is then calculated, combining the indicators from each dimension into an overall effectiveness evaluation using a weighted average or composite score. Finally, a quality management system effectiveness assessment report is generated, containing detailed assessment results for each dimension and an overall effectiveness conclusion, providing comprehensive support for quality management decision-making. This integrated assessment creates a complete quality management system evaluation framework. Multi-dimensional indicators ensure comprehensiveness and objectivity; the comprehensive index allows management to quickly understand the overall status of the system; and the assessment report provides clear direction and focus for quality improvement.

[0240] 8.7 Establish a data flow and result association mechanism between assessment modules to ensure the consistency and traceability of assessment results.

[0241] First, a unified data interface standard is established, defining the format requirements and field specifications for all types of input data to ensure data standardization and consistency. The system then establishes a data preprocessing pipeline to clean, validate, and convert all types of input logs and evaluation data, standardizing time formats, handling missing values, and standardizing coding. Next, a correlation mechanism for evaluation results is established. Using key fields such as site number, device number, and timestamp, the results of different evaluation modules are analyzed for correlation, identifying correlations between influencing factors. Next, a consistency verification mechanism for evaluation results is established. Through cross-validation and logical checks, the logical consistency of evaluation results across modules is ensured. For example, sites with low service availability are associated with higher failure rates. Finally, a traceability mechanism for evaluation results is established, recording the input data source, processing process, and output results for each evaluation, supporting auditing and retrospective analysis of evaluation results. This data flow mechanism ensures the standardization of the evaluation process and the reliability of the results. Correlation analysis reveals the inherent connections between different quality indicators. Consistency verification improves the accuracy of evaluation results. The traceability mechanism enhances the transparency and credibility of the evaluation process.

[0242] 8.8 Establish a dynamic monitoring and early warning mechanism for the effectiveness of the quality management system to achieve real-time tracking of quality status and abnormal early warning.

[0243] First, a quality indicator monitoring threshold system was established. Based on historical data and industry standards, normal ranges, warning thresholds, and alarm thresholds were set for each quality indicator. A real-time data collection mechanism was implemented to regularly collect data such as operation logs, transmission logs, and fault records from various business systems to ensure the timeliness and integrity of monitoring data. A quality indicator calculation engine was then established to automatically calculate quality indicators for various dimensions, including daily, weekly, and monthly timescales, according to pre-set calculation rules and cycles. An anomaly detection and early warning mechanism was then established. Through threshold comparison, trend analysis, and abnormal pattern recognition, abnormal changes in quality indicators were promptly identified and automatically triggered. Finally, a visual display of quality status was established. Through dashboards, trend charts, and heat maps, the operating status and changing trends of the quality management system were intuitively displayed to support management decision-making. This dynamic monitoring mechanism enabled real-time control of quality status; the early warning mechanism ensured the timely detection and resolution of issues; the visual display improved the comprehensibility of quality information; and dynamic tracking provided continuous data support for quality improvement.

[0244] 8.9 Establish a continuous improvement mechanism for the quality management system based on the assessment results, forming a closed-loop management of assessment-analysis-improvement-verification.

[0245] First, a framework for analyzing assessment results is established to conduct an in-depth analysis of the quality management system effectiveness assessment results, identifying the root causes of quality issues and improvement opportunities. The system then establishes a problem classification and prioritization mechanism, categorizing and prioritizing identified quality issues based on severity, scope, and difficulty of improvement, thereby prioritizing improvements. Targeted improvement measures are then developed, including various types of actions, such as process optimization, technology upgrades, personnel training, and system improvements, forming a detailed improvement plan and timeline. Next, a mechanism for tracking improvement results is established. Through regular evaluation and comparative analysis, the effectiveness of improvement measures is monitored and the achievement of improvement goals is verified. Finally, a mechanism for summarizing and promoting improvement experiences is established, transforming successful improvement experiences into standardized best practices for application throughout the quality management system, achieving a continuous improvement spiral. Through this continuous improvement mechanism, a closed-loop quality management system is established; problem-oriented improvement ensures efficient resource allocation; tracking results ensures the effectiveness of improvement measures; and the promotion of experience enhances the overall quality management level.

[0246] 8.10 This embodiment achieves comprehensive improvement and continuous optimization of meteorological observation data quality management by establishing a complete quality management system effectiveness evaluation system.

[0247] Through the evaluation of service availability indicators, we have achieved quantitative monitoring of the stability of meteorological observation equipment, increasing equipment availability from 85% to over 95%, significantly improving the continuity of observation services. Through the evaluation of data transmission quality indicators, we have achieved precise control over the integrity and timeliness of data transmission, with a data transmission success rate exceeding 99% and a real-time transmission rate exceeding 98%, ensuring high-quality data transmission. Through the evaluation of fault and maintenance indicators, we have achieved effective management of equipment faults and repair efficiency, reducing the average fault response time to less than 1 hour and the average repair time to less than 4 hours, significantly improving repair efficiency. Through the evaluation of user satisfaction indicators, we have achieved a user-perspective evaluation of service quality, with overall user satisfaction reaching 4.2 points (out of 5) and a data quality satisfaction rate exceeding 90%, significantly improving the user experience. Through the evaluation of external supplier evaluation indicators, we have achieved an objective evaluation of supplier quality, with the proportion of high-quality suppliers exceeding 80%, significantly improving the overall service level of suppliers.

[0248] The implementation of the quality management system effectiveness assessment has enabled a shift in meteorological observation data quality management from qualitative to quantitative management, from passive response to proactive prevention, and from single-point improvement to system optimization. This has significantly improved overall quality management, enhanced data quality stability, and continuously improved user satisfaction, laying a solid foundation for the comprehensive improvement of meteorological service quality.

[0249] The quality management system effectiveness assessment method of this embodiment has the advantages of comprehensive assessment dimensions, scientific indicator design, strong adaptability, intuitive and clear results, strong decision-making support, and a complete management closed loop. Through multi-dimensional quantitative assessment, it achieves refinement and scientificization of quality management, and provides strong support for the high-quality development of meteorological observation services.

[0250] Evaluate the effectiveness of the quality management system from multiple dimensions, including business availability, data transmission quality, fault and repair conditions, user satisfaction, and external supplier evaluation, to fully understand the system operation status; design scientific and reasonable indicators for each evaluation dimension, such as availability rate, transmission success rate, average response time, user satisfaction, etc., to quantify the evaluation results for easy comparison and analysis; function design supports flexible input, and the evaluation dimensions and indicators can be selected according to actual conditions to adapt to different evaluation needs; the evaluation results are organized by dimensions and indicators with clear levels for easy understanding and use; through quantitative evaluation of the effectiveness of the quality management system, provide an objective basis for management decision-making, and guide system optimization and improvement; the management closed loop is complete: it covers all aspects of the quality management system, from equipment operation, data transmission, fault repair to user feedback and supplier management, forming a complete management closed loop.

[0251] like Figure 2 As shown, a meteorological observation data automatic quality assessment system includes:

[0252] Data initialization module: used to collect multiple meteorological observation data and perform initial configuration to obtain meteorological assessment parameter sets;

[0253] Data preprocessing module: used to perform preprocessing operations on various parameters in the meteorological assessment parameter set to obtain the full meteorological assessment data set;

[0254] Statistical analysis module: used to perform global statistical analysis on the full meteorological assessment data set to obtain a global statistical analysis data set; divide the global statistical analysis data set into meteorological quality assessment subsets by region and depth level; and further analyze the meteorological quality assessment subsets to obtain hierarchical statistical analysis results;

[0255] Error analysis module: used to obtain the absolute and relative errors of meteorological assessment data from manual and automatic stations based on the results of hierarchical statistical analysis, set the proportion of data within each error tolerance range, and obtain meteorological quality characteristic data through iterative verification;

[0256] Report generation module: used to generate meteorological quality distribution characteristics of different regions and levels based on meteorological quality characteristic data, so as to obtain the time series difference between meteorological assessment manual station and automatic station data, and save the generated data and acquired data into folders respectively, and output the final meteorological quality assessment report by matching the output format set by different users.

[0257] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A method for automated quality assessment of meteorological observation data, characterized in that: The following steps are involved: Collect multiple meteorological observation data and perform initial configuration to obtain a set of meteorological assessment parameters; Perform preprocessing operations on various parameters in the meteorological assessment parameter set to obtain the full meteorological assessment data set; Performing global statistical analysis on the entire meteorological assessment data set to obtain a global statistical analysis data set; The global statistical analysis dataset is divided into meteorological quality assessment subsets by region and depth level; and the meteorological quality assessment subsets are further analyzed; Further analysis of the meteorological quality assessment subsets includes: group statistical analysis based on time dimension stratification, spatial dimension stratification, and factor dimension stratification, comprehensive analysis combined with the EMI environmental meteorological condition assessment index, and quantitative separation analysis of the impact of meteorological conditions on the quality assessment of each subset data item by quantitatively characterizing them. The different analysis results are then integrated to ultimately obtain hierarchical statistical analysis results. The quantitative separation analysis includes: based on the EMI environmental meteorological condition evaluation index, and Obtain the emission change rate RE and meteorological condition change rate RW; where R0 and R1 are the measured meteorological data parameters during the comparison period, and E0 and E1 are the EMI environmental meteorological condition evaluation indexes during the comparison period; The contribution rate of meteorological condition changes to data quality is obtained through EMI environmental meteorological condition evaluation index analysis, so as to obtain the correlation between the EMI environmental meteorological condition evaluation index and meteorological observation data quality indicators, and quantitatively evaluate the impact of meteorological conditions on observation data quality; Based on the spatiotemporal distribution characteristics of EMI, the changing trends of meteorological conditions in different regions and their impact on the quality of meteorological observation data are analyzed; The two influence levels were integrated to obtain the quantitative separation analysis results; Based on the results of hierarchical statistical analysis, the absolute and relative errors of the meteorological assessment data from manual and automatic stations are obtained, and the proportion of data within each error tolerance range is set. After iterative verification, meteorological quality characteristic data is obtained. Based on the meteorological quality characteristic data, meteorological quality distribution characteristics of different regions and levels are generated to obtain the temporal differences between the meteorological assessment manual station and automatic station data, and the generated data and acquired data are saved in folders respectively. By matching the output formats set by different users, the final meteorological quality assessment report is output.

2. The method for automatic quality assessment of meteorological observation data according to claim 1, characterized in that: The initialization configuration includes: Create and define multiple initial states, and mark meteorological observation data as correct, suspicious, erroneous, and missing, respectively, to obtain marked data of the initial states of meteorological observation data; Constructing meteorological quality assessment indicators based on the labeled data, wherein the meteorological quality assessment indicators are actual rate, correct rate, error rate and suspicious rate; Detecting whether the initialization configuration of meteorological observation data is completed in the current meteorological observation environment according to the marked data; According to the meteorological quality assessment indicators, the meteorological observation data that has not completed the initialization configuration is loaded into the data allocation resource configuration. The data allocation resource configuration dynamically adjusts the memory usage and thread allocation according to the data scale and resources to complete the operation meteorological quality assessment initialization configuration and obtain the initialized meteorological assessment parameter set.

3. The method for automatic quality assessment of meteorological observation data according to claim 1, characterized in that: Data preprocessing includes: By receiving the configuration path parameters of the environment adaptive initialization configuration, the configuration path parameters are used to perform the initialization configuration operation in conjunction with the meteorological observation data; Automatically check the configuration path parameters to determine whether they exist and are accessible; When the configuration path parameters exist and are accessible, load the meteorological observation data, read and verify whether the meteorological observation data is complete; When the meteorological observation data is incomplete, the format conversion and missing value supplement processing are performed on the meteorological observation data, and normalization processing is performed at the same time to obtain the processed meteorological assessment parameter set, which is the full meteorological assessment data set.

4. The method for automatic quality assessment of meteorological observation data according to claim 3, characterized in that: Creating a plurality of quality assessment codes, wherein the quality assessment codes include assessment thresholds of 0 to 3 as quality assessment digits; The quality assessment data respectively marks the current multiple meteorological observation data to obtain the marking results; The marking results are set into a digital range, each digital range represents four quality assessment states: excellent, good, average and poor; Establish a meteorological quality identification system based on the digital range set by the quality assessment code and quality assessment status; Based on the meteorological quality identification system, the dynamic changes of multiple meteorological observation data at different times are identified to obtain the changing trend of meteorological observation status; Multi-dimensional statistical analysis is performed based on the changing trend of meteorological observation status to obtain the quality status of meteorological observation data under multi-dimensional analysis, and to determine whether the changing trend amplitude between the current meteorological observation data and the meteorological observation data after dynamic changes is in a smooth or rapid state, so as to configure a hierarchical statistical analysis strategy.

5. The method for automatic quality assessment of meteorological observation data according to claim 4, characterized in that: The stratified statistical analysis strategy includes: Capture the abnormal state of meteorological observation data that changes dynamically at different times to obtain abnormal state data; The meteorological time series characteristics and spatial correlation of the abnormal state data are interpolated to obtain the sudden drop or sudden rise amplitude index, and the abnormal state data are marked according to the amplitude index using the quality assessment code to obtain the secondary marking result; The abnormal state data in the secondary marking results are further checked for climate limit values, internal consistency, temporal consistency, and spatial consistency. The data that does not meet the meteorological inspection rules are marked using quality assessment codes to obtain the tertiary marking results. Fitting is performed on the meteorological observation data with general or poor quality assessment status in the three rounds of marking results to obtain processing results; The processed meteorological observation data and the original meteorological observation data with excellent or good quality assessment status are integrated to generate preliminary quality assessment results and transmitted to the global statistical analysis for coordinated analysis and processing.

6. The method for automatic quality assessment of meteorological observation data according to claim 1, characterized in that: Perform global statistical analysis on the full meteorological assessment dataset to obtain a global statistical analysis dataset, including: Set the meteorological micro-disturbance filtering threshold and dynamically adjust it according to the characteristics of the meteorological region and the preliminary quality assessment results to optimize the observation quality of each data item in the full meteorological assessment dataset; Obtain basic statistical indicators of extreme values, mean values, standard deviations, and coefficients of variation of meteorological elements for each data item in the full meteorological assessment dataset, and establish a global meteorological quality assessment benchmark; Based on the global meteorological quality assessment benchmark, a multi-level grouping statistical analysis of meteorological data is performed according to the time dimension, space dimension and element dimension to obtain a hierarchical statistical analysis strategy to form an environmental quality assessment data set; The hierarchical statistical analysis strategy is combined with the meteorological quality identification system and the quality assessment indicators of actual rate, correct rate, error rate and suspicious rate at each level, and the EMI environmental meteorological condition assessment index is introduced for comprehensive quantitative analysis; The EMI environmental meteorological condition assessment index represents a comprehensive index that characterizes the processes of aerosol emission, deposition, transmission, and diffusion under the influence of meteorological conditions. A larger EMI value indicates that the meteorological conditions are more unfavorable for the diffusion of atmospheric pollutants. Through iterative statistical verification and threshold determination, the global statistical results are fused with the hierarchical statistical results, and finally a complete global statistical analysis data set is generated, which includes multi-dimensional statistical features, quality assessment identification and EMI environmental meteorological condition assessment index.

7. The method for automatic quality assessment of meteorological observation data according to claim 6, characterized in that: The global statistical analysis dataset is divided into meteorological quality assessment subsets by region and depth level. The meteorological quality assessment subsets are further analyzed to obtain hierarchical statistical analysis results, including: The data items in the global statistical analysis data set are layered in the time dimension, including grouping statistics by year, season, month, and day. Combined with the quality assessment marking results, a time-level meteorological quality assessment subset is formed to generate time series basic data. The spatial dimension of each data item in the global statistical analysis data set is stratified, including grouping statistics by region, geographical features and elevation. The results are adjusted using regional characteristics to form spatial-level meteorological quality assessment subsets and generate spatial distribution data. The data items in the global statistical analysis data set are layered by element dimension, and the quality assessment indicators of the meteorological elements of temperature, humidity, air pressure, and wind speed are calculated respectively. The hierarchical statistical analysis results at the element level are obtained, and the element distribution characteristic data is generated; Comprehensive analysis is conducted using the EMI environmental meteorological condition assessment index to obtain hierarchical statistical analysis results; By quantitatively characterizing the impact of meteorological conditions on the quality assessment of each meteorological quality assessment subset data item, a quantitative separation analysis is performed to obtain the quantitative separation analysis results.

8. The method for automatic quality assessment of meteorological observation data according to claim 1, characterized in that: Based on the results of hierarchical statistical analysis, the absolute and relative errors of the meteorological assessment data from manual and automatic stations are obtained, and the proportion of data within each error tolerance range is set. After iterative verification, meteorological quality characteristic data is obtained, including: The error tolerance analysis is based on the hierarchical statistical analysis results and the EMI environmental meteorological condition evaluation index, and is implemented by constructing an error statistical function. The error statistical function receives the meteorological evaluation manual station observation data and the automatic station observation data as input parameters and performs the following operations: The absolute error value and relative error percentage between the two sets of meteorological observation data were calculated based on the hierarchical statistical analysis results at the factor level; The meteorological quality assessment subsets at the time and space levels were used to calculate the proportion of data within the relative error threshold of ±5% and the absolute error within the threshold of ±10%. The standard deviation is calculated based on the EMI environmental meteorological condition evaluation index, and finally a statistical result list containing the error interval percentage and standard deviation of the absolute error value and relative error is returned. Based on the calculation results, an error interval ratio distribution table is generated and passed to the meteorological quality characteristic data as a key quantitative indicator.

9. The method for automatic quality assessment of meteorological observation data according to claim 8, characterized in that: Dynamically bind data fields based on element distribution feature data and error interval ratio distribution tables to obtain the probability density distribution characteristics of meteorological elements in different regions and levels, and mark data quality levels in combination with the meteorological quality identification system to support basic data and time series feature analysis; Using probability density distribution characteristics and time dimension layered data, combined with error statistics results, time series differences are generated to obtain the temporal variation trend of meteorological assessment data from manual and automatic stations. Error coding is used to distinguish data points within different error tolerance intervals, and the time series feature data is passed to spatial feature analysis. Based on the data features obtained from time series feature analysis and spatial dimension layered data, combined with the error interval ratio distribution, the geospatial thermal effect is obtained to obtain the spatial distribution characteristics of the meteorological quality assessment index. The error mapping range is dynamically adjusted according to the EMI environmental meteorological condition assessment index, and the spatial distribution characteristics are stored in the data. Obtain meteorological quality characteristic data based on probability density distribution, time series characteristic analysis, and spatial characteristic analysis. Generate a comprehensive report on meteorological quality distribution characteristics in different regions and levels. Integrate the time series difference analysis results of meteorological assessment manual station and automatic station data. Save the generated data and acquired data into corresponding folders according to the analysis dimensions. That is, by matching the output formats set by different users, the final meteorological quality assessment report is output.

10. A meteorological observation data automated quality assessment system, configured to execute a meteorological observation data automated quality assessment method according to any one of claims 1 to 9, characterized in that: include: Data initialization module: used to collect multiple meteorological observation data and perform initial configuration to obtain a set of meteorological assessment parameters; Data preprocessing module: used to perform preprocessing operations on various parameters in the meteorological assessment parameter set to obtain the full meteorological assessment data set; Statistical analysis module: used to perform global statistical analysis on the entire meteorological assessment data set to obtain a global statistical analysis data set; The global statistical analysis dataset is divided into meteorological quality assessment subsets by region and depth; and the meteorological quality assessment subsets are further analyzed. The further analysis of the meteorological quality assessment subsets includes: group statistical analysis by time dimension stratification, spatial dimension stratification, and factor dimension stratification, comprehensive analysis combined with the EMI environmental meteorological condition assessment index, and quantitative separation analysis of the impact of meteorological conditions on the quality assessment of each subset data item by quantitatively characterizing them. The different analysis results are finally integrated to obtain the hierarchical statistical analysis results. Error analysis module: used to obtain the absolute and relative errors of meteorological assessment data from manual and automatic stations based on the results of hierarchical statistical analysis, set the proportion of data within each error tolerance range, and obtain meteorological quality characteristic data through iterative verification; Report generation module: used to generate meteorological quality distribution characteristics of different regions and levels based on meteorological quality characteristic data, so as to obtain the time series difference between meteorological assessment manual station and automatic station data, and save the generated data and acquired data into folders respectively, and output the final meteorological quality assessment report by matching the output format set by different users.

Citation Information

Patent Citations

  • Weather forecast data quality detection method

    CN113742927A

  • Technology for monitoring and evaluating low-atmosphere self-purification capacity process

    CN115081845A