Multi-source data cross quality control and instrument fault intelligent diagnosis method, system, medium and equipment
Patent Information
- Application Number
- CN202610923760.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-29
AI Technical Summary
1、本发明从常规参数扩展到颗粒物微物理、化学组分和VOCs等多维参数,实现了从“总量监控”到“组分解析”级别的质控。
Smart Images

Figure CN122839192A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of environmental monitoring technology and data analysis, and in particular to a method, system, medium, and equipment for multi-source data quality control, abnormal data identification, and intelligent diagnosis of instrument faults for multi-source data cross-quality control and instrument faults in an automatic ambient air monitoring superstation. Background Technology
[0002] Current technology primarily relies on monitoring the status parameters of sensors within a single instrument (such as internal temperature, pressure, flow rate, and light source energy). When these direct parameters exceed thresholds, the instrument triggers an alarm. This method has significant limitations: (1) Single quality control dimension: It still relies heavily on the internal status parameters (flow rate, temperature, pressure, etc.) of a single instrument for judgment, and cannot cope with data distortion caused by the failure of multiple components, systematic deviation of the sampling system, or complex external interference.
[0003] (2) Data "island" phenomenon: The massive amount of multi-source data generated within the super station is not effectively correlated. For example, although there is a clear physicochemical correlation between the chemical component (sulfate, nitrate, organic matter, elemental carbon, etc.) data measured by PM2.5 online mass spectrometers (SPAMS, ACSM) and the PM2.5 mass concentration data based on β-ray method or oscillating balance method (the sum of the mass of chemical components should be approximately equal to the total mass concentration), there is a lack of automated cross-validation mechanism.
[0004] (3) Insufficient ability to diagnose complex faults: Simple threshold judgment is completely ineffective for complex problems such as slow drift of instrument detectors, adsorption / desorption of reactive gases (such as HONO) on the sampling tube wall, and abnormal VOC species spectrum.
[0005] (4) Relying on fixed rules and lacking intelligent learning: Traditional quality control rules are based on fixed thresholds and linear relationships, which cannot adaptively learn complex nonlinear patterns in the data. They have weak identification ability for sudden changes (such as sudden pollution emissions) that do not conform to conventional experience data but are not instrument failures. The misjudgment rate is high. Summary of the Invention
[0006] To address the aforementioned problems, the purpose of this invention is to provide a method, system, medium, and device for cross-source data quality control and intelligent instrument fault diagnosis, which can accurately trace the source of data anomalies, not only distinguishing between "real contamination" and "equipment failure," but also locating suspected fault links.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for cross-quality control of multi-source data and intelligent diagnosis of instrument faults, comprising: real-time acquisition and fusion of multi-source monitoring data from an environmental monitoring superstation, the multi-source monitoring data including conventional pollutant concentration data, particulate matter physicochemical property data, volatile organic compound data, and auxiliary parameter data; parallel input of the fused multi-source monitoring data into an intelligent data health assessment model, the model including a rule base verification module based on physicochemical correlation and an anomaly pattern recognition module based on machine learning, to perform multi-dimensional cross-validation and anomaly detection on the multi-source monitoring data respectively, and output rule verification results and anomaly identification results; combining the rule verification results and anomaly identification results, using a fusion decision mechanism to determine whether the current data anomaly is caused by a real pollution event or by instrument fault, and tracing and locating the type of instrument fault to generate a preliminary diagnostic conclusion; generating a visual diagnostic report from the preliminary diagnostic conclusion and related data information, and outputting an alarm.
[0008] Furthermore, the rule base verification module based on physicochemical correlation includes the following multiple cross-validation sub-modules that work in parallel: The particulate matter chemical quality closure verification submodule is used to calculate in real time the ratio of the sum of the mass concentrations of the chemical components of particulate matter to the monitored value of the total mass concentration of particulate matter, and to determine whether the ratio is within a preset reasonable range, so as to verify the integrity of the particulate matter quality data. The characteristic element and volatile organic compound species comparison value verification submodule is used to monitor the normal deviation of the ratio of specific inorganic elements and the stability of the ratio of homologous volatile organic compound characteristic species pairs, so as to diagnose the anomalies of the sampling or detection system. The chemical and physical property correlation verification submodule is used to verify the stability of mass scattering efficiency between particulate matter scattering coefficient and particulate matter mass concentration and particle size distribution, as well as to verify the correlation between particulate matter concentration and relative humidity and visibility. The photochemical coupling verification submodule is used to incorporate free radical precursors into the photochemical system, verify the correlation between their diurnal concentration variation and solar radiation intensity, and analyze the spatiotemporal coupling relationship between the ozone generation potential calculated from volatile organic compounds and the measured ozone concentration.
[0009] Furthermore, the specific processing steps of the machine learning-based anomaly pattern recognition module include: Use multi-source monitoring data under normal historical operating conditions to train time series prediction models or unsupervised learning models to learn the normal change patterns of each monitoring parameter and their interrelationships. The real-time collected data is input into the trained model to calculate the reconstruction error or prediction deviation. When the error or deviation exceeds the preset abnormal threshold, the corresponding data is marked as abnormal. When an anomaly is detected, contribution analysis is used to identify one or more variables that contribute the most to the anomaly, providing clues for tracing the source of the fault.
[0010] Furthermore, a fusion decision-making mechanism is employed to determine whether the current data anomaly is caused by a genuine contamination event or by instrument malfunction, including: When the multi-dimensional data shows consistency, conforms to the laws of atmospheric physicochemical processes, and the anomalies in the machine learning model are widespread and corroborated by meteorological conditions and data from surrounding stations, it is diagnosed as a real pollution event. When the total mass of particulate matter is normal but chemical components are missing, the comparison values of multiple volatile organic compounds show a systematic shift, or the concentration of reactive gases is consistently lower than the model prediction value and its correlation with other parameters is less than at least one of the set values, the corresponding instrument malfunction type is diagnosed.
[0011] Furthermore, after the alarm output, there is also a knowledge accumulation and model self-learning process: the diagnosis results confirmed by the operation and maintenance personnel are used as labeled data and fed back to the machine learning model and rule base for model retraining and dynamic adjustment of rule thresholds.
[0012] Furthermore, after real-time collection and fusion of multi-source monitoring data from environmental monitoring superstations, and before data health assessment, a data preprocessing step is also included: time alignment of multi-source monitoring data from different sampling periods with a unified time benchmark, and interpolation or aggregation to fill in missing values in the data.
[0013] Furthermore, the auxiliary parameter data includes: five meteorological parameters of the station, environmental parameters of the station building, and internal status parameters of each instrument; and data from different sources and formats are converted into standard formats through a unified data interface.
[0014] A multi-source data cross-quality control and instrument fault intelligent diagnosis system includes: a data acquisition and fusion module, which collects and fuses multi-source monitoring data from an environmental monitoring superstation in real time. The multi-source monitoring data includes conventional pollutant concentration data, particulate matter physicochemical property data, volatile organic compound data, and auxiliary parameter data; a data health assessment module, which inputs the fused multi-source monitoring data in parallel into an intelligent data health assessment model. This model includes a rule base verification module based on physicochemical correlation and an anomaly pattern recognition module based on machine learning, to perform multi-dimensional cross-validation and anomaly detection on the multi-source monitoring data, and outputs rule verification results and anomaly identification results; and an intelligent diagnosis and output module, which integrates the rule verification results and anomaly identification results, uses a fusion decision mechanism to determine whether the current data anomaly is caused by a real pollution event or an instrument fault, traces and locates the type of instrument fault, and generates a preliminary diagnostic conclusion; and generates a visualized diagnostic report from the preliminary diagnostic conclusion and related data information, and outputs an alarm.
[0015] A computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform any of the methods described above.
[0016] A computing device includes: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods described above.
[0017] The present invention has the following advantages due to the adoption of the above technical solutions: 1. This invention extends from conventional parameters to multi-dimensional parameters such as particulate matter microphysical and chemical composition and VOCs, realizing quality control from "total monitoring" to "component analysis".
[0018] 2. This invention enables the system to learn from data and discover unknown abnormal patterns through machine learning, reducing the absolute dependence on prior knowledge and enhancing its adaptability to new faults and complex contamination.
[0019] 3. This invention utilizes strong physicochemical constraints such as mass closure and element ratios to make diagnostic conclusions more reliable and source tracing more accurate.
[0020] 4. This invention enables the entire quality control system to be continuously optimized as operational experience is accumulated through a human-computer interaction feedback loop, thus becoming a "vitality expert system". Attached Figure Description
[0021] Figure 1 This is a flowchart of the multi-source data cross-quality control and instrument fault intelligent diagnosis method in this embodiment of the invention. Detailed Implementation
[0022] To overcome the shortcomings of existing technologies mentioned above, this invention provides a method, system, medium, and equipment for multi-source data cross-quality control and intelligent instrument fault diagnosis. It deeply mines and utilizes the inherent physical, chemical, and statistical correlations between different observation parameters within a superstation to construct a multi-level, three-dimensional "data health" assessment system. Machine learning technology is introduced to learn data patterns under normal operating conditions from massive historical data, enhancing the ability to identify abrupt and slowly drifting faults that do not conform to conventional experience data but are not instrument malfunctions. This enables precise tracing of data anomalies, not only distinguishing between "real contamination" and "equipment failure," but also further locating suspected fault links (such as sampling, detection, calibration, etc.) and automatically generating diagnostic reports to guide efficient operation and maintenance.
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.
[0024] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0025] In one embodiment of the present invention, a method for cross-source data quality control and intelligent instrument fault diagnosis is provided. In this embodiment, as shown... Figure 1 The diagram illustrates the overall architecture and operational logic of the multi-source data quality control and fault self-diagnosis system for atmospheric observation superstations. The system begins with the collection and fusion of comprehensive monitoring data from the superstation, and, relying on a dynamic database and analysis platform, inputs the data in parallel to the quality control modules. These modules perform multi-dimensional and intelligent cross-checks on data health through multi-parameter quality closure verification, characteristic element ratio analysis, VOCs species pair verification, optical-physical relationship verification, and machine learning-based anomaly pattern recognition. The verification results converge at the intelligent diagnosis and source tracing center, where the system distinguishes between real pollution events and instrument malfunctions. Finally, the diagnostic conclusions are pushed to maintenance personnel through a visual alarm interface, while experience is fed back to the knowledge base, forming a continuously optimized, closed-loop intelligent operation and maintenance system.
[0026] Specifically, the multi-source data cross-quality control and instrument fault intelligent diagnosis method of the present invention includes the following steps: 1) Collect and integrate multi-source monitoring data from environmental monitoring superstations in real time. Multi-source monitoring data includes conventional pollutant concentration data, particulate matter physicochemical properties data, volatile organic compound data, and auxiliary parameter data.
[0027] 2) The fused multi-source monitoring data is input in parallel into the intelligent data health assessment model. The model includes a rule base verification module based on physical and chemical correlation and an anomaly pattern recognition module based on machine learning, so as to perform multi-dimensional cross-validation and anomaly detection on the multi-source monitoring data respectively, and output the rule verification results and anomaly recognition results.
[0028] 3) Based on the combined rule verification results and anomaly identification results, a fusion decision-making mechanism is used to determine whether the current data anomaly is caused by a real contamination event or by an instrument malfunction, and the instrument malfunction type is traced and located to generate a preliminary diagnostic conclusion.
[0029] 4) Generate a visual diagnostic report from the preliminary diagnostic conclusions and related data information, and output alarms.
[0030] In step 1) above, the concentration data for conventional pollutants are SO2, NOx, O3, CO, PM2.5, and PM2.5. 10 The physicochemical properties data for particulate matter include particle size distribution, chemical composition (such as the concentrations of sulfates, nitrates, ammonium salts, organic matter, elemental carbon, and various inorganic elements), optical properties (scattering coefficient, absorption coefficient), and number concentration. Volatile organic compound (VOC) data includes various hydrocarbons and aldehydes / ketones, along with their calculated ozone formation potential (OFP).
[0031] The auxiliary parameter data includes: five meteorological parameters of the station, environmental parameters of the station building, and internal status parameters of each instrument; and data from different sources and formats are converted into standard formats through a unified data interface.
[0032] In this embodiment, real-time data acquisition is achieved by building the following hardware and data layers: 1.1) Equipment Interconnection: Ensure that all instruments in the super station (conventional six-parameter monitor, particulate matter chemical composition monitor, VOCs online chromatograph, particle size analyzer, HONO analyzer, etc.) and the data output ports of the weather station are connected to a central data acquisition server via a local area network.
[0033] 1.2) Unified data interface: Deploy data acquisition software on the acquisition server to parse and convert data from different manufacturers and different communication protocols (such as Modbus, TCP / IP, SDI-12) into a unified JSON or Avro format.
[0034] 1.3) Establish a database: Deploy a high-performance time-series database to store all monitoring data, status parameters, and meteorological data reported by all instruments. Each data entry includes a precise timestamp, device ID, and parameter name.
[0035] In step 2) above, a rule base is established based on expert experience. The quality closure rule is: input formula (sulfate + nitrate + ammonium salt + organic matter + elemental carbon) / PM 2.5 Mass concentration. Set a reasonable range, such as 0.7 to 1.3. Exceeding this range will trigger a quality control event.
[0036] Species comparison rule: Calculate the ratio of ethylbenzene to m-xylene in VOCs and set its normal fluctuation range (e.g., ±30%).
[0037] Physical association rules: Establish scattering coefficient / PM 2.5 The daily range of mass concentration is called "mass scattering efficiency," and its stability is monitored.
[0038] In this embodiment, the rule base verification module based on physicochemical correlation includes the following multiple cross-validation sub-modules that operate in parallel: (1) Particulate matter chemical quality closure verification submodule, which is used to calculate the ratio of the sum of the mass concentrations of particulate matter chemical components to the monitored value of the total mass concentration of particulate matter in real time, and to determine whether the ratio is within the preset reasonable range, so as to verify the integrity of particulate matter quality data.
[0039] Specifically, the sum of the mass concentrations of PM2.5 chemical components (sulfate + nitrate + ammonium salt + organic matter + elemental carbon + crustal material) is calculated in real time and compared with the data from a PM2.5 mass concentration monitor. Under conditions of no significant dust or moisture interference, the two should exhibit a good linear relationship and a reasonable mass balance range. If the sum of the component concentrations deviates significantly from the total mass, it suggests that the data from one or both may be distorted.
[0040] (2) Characteristic element and volatile organic compound species comparison value verification submodule, used to monitor the normal deviation of the ratio of specific inorganic elements and the stability of the ratio of homologous volatile organic compound characteristic species, so as to diagnose the abnormality of the sampling or detection system.
[0041] Specifically, characteristic element ratio verification: quality control is performed using specific inorganic element ratios. For example, the [Si] / [Ca] ratio can be used to identify and differentiate dust sources; the [Zn] / [Pb] ratio may point to a specific industrial source. If this ratio deviates significantly from the local normal range during a certain period, it may indicate that the instrument is contaminated or that there are abnormal emissions.
[0042] VOCs characteristic species comparison value verification: This method utilizes the stability of the ratio of VOCs species pairs with homology and similar chemical lifetimes (such as ethylbenzene / m-xylene) to diagnose whether the VOCs monitoring system is functioning correctly. Sudden changes in the ratio may originate from changes in column performance or calibration issues.
[0043] (3) Chemical and physical property correlation verification submodule, used to verify the stability of mass scattering efficiency between particulate matter scattering coefficient and particulate matter mass concentration and particle size distribution, as well as to verify the correlation between particulate matter concentration and relative humidity and visibility.
[0044] Specifically, the relationship between the particulate scattering coefficient and the particulate mass concentration and size distribution is used to verify (mass scattering efficiency). In the long run, this relationship should be relatively stable. If the scattering coefficient remains constant while the particulate concentration increases sharply, or vice versa, then instrument drift may be occurring on one of the two sides.
[0045] (4) Photochemical coupling verification submodule, which is used to incorporate free radical precursors into the photochemical system, verify the correlation between their daily concentration variation and solar radiation intensity, and analyze the spatiotemporal coupling relationship between the ozone generation potential calculated from volatile organic compounds and the measured ozone concentration.
[0046] Specifically, the photochemical coupling verification submodule is an enhanced physicochemical correlation cross-validation module. The coupling verification of free radical measurements with their precursors and solar radiation intensity involves incorporating free radical precursors such as HONO into the photochemical system. For example, HONO photolysis is an important early morning source of OH free radicals, and its diurnal concentration variation should have a specific correlation with solar radiation intensity.
[0047] Coupling verification of ozone formation potential and measured O3 concentration: Analysis of the spatiotemporal coupling relationship between ozone formation potential (OFP) of VOCs and measured O3 concentration.
[0048] In step 2) above, the specific processing procedure of the machine learning-based anomaly pattern recognition module includes the following steps: 2.1) Use multi-source monitoring data under historical normal operating conditions to train time series prediction models or unsupervised learning models to learn the normal change patterns of each monitoring parameter and their interrelationships.
[0049] In this context, normal pattern learning involves using normal historical data from long-term series to train a time-series prediction model (such as LSTM) or an unsupervised learning model (such as an autoencoder) to learn the normal variation patterns of various parameters and their interrelationships.
[0050] The specific training process is as follows: Collect "clean" data from the past year that has been manually verified as error-free as the training set.
[0051] Model selection: For a single parameter (such as PM) 2.5 (Concentration), using an LSTM (Long Short-Term Memory) model. Let it learn PM. 2.5 The normal pattern of change within a day, a week, or even a year.
[0052] Start training: Train the LSTM model using historical data, allowing it to predict the PM of the next hour based on data from the previous 24 hours. 2.5 The training objective is to minimize the error between the predicted and actual values.
[0053] Set a threshold: After training, run the model on "clean" data and statistically analyze the distribution of prediction errors. Set the 95th percentile of the errors as the outlier threshold.
[0054] 2.2) Input the real-time collected data into the trained model, calculate the reconstruction error or prediction bias, and mark the corresponding data as abnormal when the error or bias exceeds the preset anomaly threshold. This step can detect sudden changes and slow, complex drifts that do not conform to conventional experience data but are not due to instrument malfunctions. These are difficult to capture with fixed rules.
[0055] 2.3) When an anomaly is detected, one or more variables that contribute the most to the anomaly are identified through contribution analysis, providing clues for fault tracing.
[0056] In this embodiment, the real-time diagnostic process is as follows: The latest hourly data, after preprocessing, will be simultaneously fed into the rule base and the machine learning model.
[0057] Rule base check: Automatically calculates all preset rules. For example, it found this morning's "chemical composition and / or PM2.5" rules. 2.5 The "mass concentration" ratio suddenly dropped to 0.5, triggering the rule and generating a "PM" rule. 2.5 An abnormal record of "severe imbalance in chemical mass closure".
[0058] Machine learning model inference: Meanwhile, the latest PM 2.5 The data is also fed into the pre-trained LSTM model. The model predicts the current PM based on data from the previous 24 hours. 2.5 The value should be 50 μg / m³, but the measured value is 10 μg / m³. The prediction error far exceeds the threshold, and the model generates a "PM" error. 2.5 An abnormal record is generated when the concentration shows a sudden change that does not conform to conventional experience data, and a high anomaly score is output.
[0059] In step 3) above, the fusion decision-making mechanism is used to determine whether the current data anomaly is caused by a real contamination event or an instrument malfunction, including the following steps: 3.1) When the multi-dimensional data (mass concentration, chemical composition, VOCs spectrum, optical properties) show consistent changes that conform to the laws of atmospheric physicochemical processes, and the anomalies in the machine learning model are widespread (i.e., multiple parameters are abnormal at the same time) and are consistent with meteorological conditions and data from surrounding stations, it is diagnosed as a real pollution event.
[0060] 3.2) When the total particulate matter mass is normal but chemical component items are missing, multiple volatile organic compound species comparison values show systematic shifts, or the concentration of reactive gases is consistently lower than the model's predicted value and its correlation with other parameters is less than at least one of the set values (for example, during air pollution, NOx and CO in the atmosphere usually rise synchronously because they share the same emission source. Correlation coefficients can be quantified using algorithms, such as correlation or ratio. If they exceed the range, an alarm is triggered indicating a potential problem), the instrument is diagnosed as having the corresponding possible fault type, providing clues for troubleshooting.
[0061] Systematic bias should refer not only to volatile organic compounds (VOCs) but also to the ratios between the chemical components of PM2.5. Common ranges for these ratios among representative components can usually be determined based on atmospheric chemistry theories and long-term observational databases. If these ratios exceed this range, it typically indicates a problem with the sampling pipeline, instrument measurement, or calibration procedures, or it could be due to a nearby emission source affecting the measurement data. This ratio range exhibits regional and seasonal variations. For example, a significantly lower measured PM2.5 concentration than the sum of the concentrations of its chemical components constitutes a systematic bias. Furthermore, in typical air pollution episodes, the concentration of primary PM2.5 components is often much higher than that of secondary components, which is also frequently cited as evidence of data bias.
[0062] In this embodiment, abnormal data from the chemical component analyzer: if the total PM2.5 mass is normal, but one or more chemical components are suddenly missing, it indicates a malfunction in the sample introduction, ionization, or detection system of the component analyzer.
[0063] VOCs monitor calibration drift: Multiple VOCs species comparison values simultaneously exhibit systematic shifts, and the machine learning model detects slow distribution changes.
[0064] Adsorption in the sampling system: For reactive gases such as HONO, if the concentration value is consistently lower than the value predicted by the machine learning model based on historical data and meteorological conditions, and the correlation with other related parameters (such as NO2) is lower than the set value, it suggests that there may be adsorption or chemical reaction losses in the sampling pipeline.
[0065] In step 4) above, a diagnostic report is generated, including: Time: the start time of the abnormality. Abnormal parameter: PM 2.5 Mass concentration. Abnormal manifestation: The concentration value drops sharply and does not match the chemical composition. Preliminary diagnosis: PM2.5 2.5 The mass concentration monitor is suspected of being malfunctioning. Recommended steps: Check the instrument's sampling flow rate and whether the cutting head is clogged, and perform a manual calibration.
[0066] In this embodiment, the diagnostic report is a visualized report that includes a data time series diagram, a correlation diagram, a model anomaly score, and diagnostic conclusions. The diagnostic report is pushed to the operations engineer responsible for the site via the operations and maintenance platform's message center, SMS, or App.
[0067] In the above embodiments, after the alarm output in step 4), there is also a knowledge accumulation and model self-learning process: the diagnosis results confirmed by the operation and maintenance personnel are used as labeled data and fed back to the machine learning model and rule base for model retraining and dynamic adjustment of rule thresholds.
[0068] Specifically, the engineer marked the alarm as "handled, diagnosis correct" on the operations and maintenance platform, and filled in the actual cause as "sampling pipeline blockage". This confirmation feedback from the engineer was recorded in the knowledge base. These confirmed cases can serve as valuable data for future optimization of machine learning models or adjustment of rule thresholds, making the system increasingly "intelligent".
[0069] In the above embodiments, after step 1) real-time acquisition and fusion of multi-source monitoring data from the environmental monitoring superstation, and before step 2) data health assessment, a data preprocessing step is also included: time alignment of multi-source monitoring data from different sampling periods with a unified time benchmark, and interpolation or aggregation filling of missing values in the data.
[0070] In this embodiment, the latest data is retrieved from the time-series database at a set time (e.g., every minute).
[0071] Data alignment: Since different instruments have different response times and sampling periods, data are interpolated or aggregated using a unified time base (such as 1 hour) to ensure that all data are aligned at the same point in time.
[0072] Missing value handling: For data that is temporarily missing, simple methods such as linear interpolation are used to fill in the missing values.
[0073] In one embodiment of the present invention, a multi-source data cross-quality control and instrument fault intelligent diagnosis system is provided, comprising: The data acquisition and fusion module collects and fuses multi-source monitoring data from the environmental monitoring superstation in real time. The multi-source monitoring data includes conventional pollutant concentration data, particulate matter physicochemical properties data, volatile organic compound data, and auxiliary parameter data. The data health assessment module inputs the fused multi-source monitoring data into the intelligent data health assessment model in parallel. This model includes a rule base verification module based on physical and chemical correlations and an anomaly pattern recognition module based on machine learning, so as to perform multi-dimensional cross-validation and anomaly detection on the multi-source monitoring data respectively, and output the rule verification results and anomaly recognition results. The intelligent diagnosis and output module integrates rule verification results and anomaly identification results, and uses a fusion decision-making mechanism to determine whether the current data anomaly is caused by a real contamination event or an instrument malfunction, and traces and locates the type of instrument malfunction to generate a preliminary diagnostic conclusion. The preliminary diagnostic conclusions and related data information will be used to generate a visual diagnostic report and generate alarms.
[0074] In the above embodiments, the rule base verification module based on physicochemical correlation includes the following multiple cross-validation sub-modules that operate in parallel: The particulate matter chemical quality closure verification submodule is used to calculate in real time the ratio of the sum of the mass concentrations of the chemical components of particulate matter to the monitored value of the total mass concentration of particulate matter, and to determine whether the ratio is within a preset reasonable range, so as to verify the integrity of the particulate matter quality data. The characteristic element and volatile organic compound species comparison value verification submodule is used to monitor the normal deviation of the ratio of specific inorganic elements and the stability of the ratio of homologous volatile organic compound characteristic species pairs, so as to diagnose the anomalies of the sampling or detection system. The chemical and physical property correlation verification submodule is used to verify the stability of mass scattering efficiency between particulate matter scattering coefficient and particulate matter mass concentration and particle size distribution, as well as to verify the correlation between particulate matter concentration and relative humidity and visibility. The photochemical coupling verification submodule is used to incorporate free radical precursors into the photochemical system, verify the correlation between their diurnal concentration variation and solar radiation intensity, and analyze the spatiotemporal coupling relationship between the ozone generation potential calculated from volatile organic compounds and the measured ozone concentration.
[0075] In the above embodiments, the specific processing procedure of the machine learning-based anomaly pattern recognition module includes: Use multi-source monitoring data under normal historical operating conditions to train time series prediction models or unsupervised learning models to learn the normal change patterns of each monitoring parameter and their interrelationships. The real-time collected data is input into the trained model to calculate the reconstruction error or prediction deviation. When the error or deviation exceeds the preset abnormal threshold, the corresponding data is marked as abnormal. When an anomaly is detected, contribution analysis is used to identify one or more variables that contribute the most to the anomaly, providing clues for tracing the source of the fault.
[0076] In the above embodiments, the fusion decision-making mechanism is used to determine whether the current data anomaly is caused by a real contamination event or by an instrument malfunction, including: When the multi-dimensional data shows consistency, conforms to the laws of atmospheric physicochemical processes, and the anomalies in the machine learning model are widespread and corroborated by meteorological conditions and data from surrounding stations, it is diagnosed as a real pollution event. When the total mass of particulate matter is normal but chemical components are missing, the comparison values of multiple volatile organic compounds show a systematic shift, or the concentration of reactive gases is consistently lower than the model prediction value and its correlation with other parameters is less than at least one of the set values, the corresponding instrument malfunction type is diagnosed.
[0077] In the above embodiments, after the alarm output, there is also a knowledge accumulation and model self-learning process: the diagnosis results confirmed by the operation and maintenance personnel are used as labeled data and fed back to the machine learning model and rule base for model retraining and dynamic adjustment of rule thresholds.
[0078] In the above embodiments, after real-time collection and fusion of multi-source monitoring data from environmental monitoring superstations and before data health assessment, a data preprocessing step is also included: time alignment of multi-source monitoring data from different sampling periods with a unified time benchmark, and interpolation or aggregation filling of missing values in the data.
[0079] In the above embodiments, the auxiliary parameter data includes: five meteorological parameters of the station, environmental parameters of the station building, and internal status parameters of each instrument; and data from different sources and formats are converted into a standard format through a unified data interface.
[0080] The system provided in this embodiment is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.
[0081] In one embodiment of the present invention, a computing device is provided. This computing device can be a terminal and may include a processor, a communication interface, memory, a display screen, and an input device. The processor, communication interface, and memory communicate with each other via a communication bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. When the computer programs are executed by the processor, they implement the methods described in the above embodiments. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, a management network, NFC (Near Field Communication), or other technologies. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad mounted on the casing of the computing device, or an external keyboard, touchpad, or mouse. The processor can call logical instructions stored in the memory.
[0082] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0083] In one embodiment of the present invention, a computer program product is provided, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to perform the methods provided in the above-described method embodiments.
[0084] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided, which stores server instructions that cause a computer to perform the methods provided in the above embodiments.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for cross-source data quality control and intelligent instrument fault diagnosis, characterized in that, include: Real-time collection and fusion of multi-source monitoring data from environmental monitoring superstations. Multi-source monitoring data includes conventional pollutant concentration data, particulate matter physicochemical properties data, volatile organic compound data, and auxiliary parameter data; The fused multi-source monitoring data is input in parallel into the intelligent data health assessment model, which includes a rule base verification module based on physical and chemical correlations and an anomaly pattern recognition module based on machine learning, to perform multi-dimensional cross-validation and anomaly detection on the multi-source monitoring data respectively, and output the rule verification results and anomaly recognition results. Based on the combined rule verification results and anomaly identification results, a fusion decision-making mechanism is used to determine whether the current data anomaly is caused by a real contamination event or by an instrument malfunction, and to trace and locate the type of instrument malfunction to generate a preliminary diagnostic conclusion. The preliminary diagnostic conclusions and related data information will be used to generate a visual diagnostic report and generate alarms.
2. The method for multi-source data cross-quality control and intelligent instrument fault diagnosis as described in claim 1, characterized in that, The rule base validation module based on physicochemical correlation includes the following multiple cross-validation sub-modules that work in parallel: The particulate matter chemical quality closure verification submodule is used to calculate in real time the ratio of the sum of the mass concentrations of the chemical components of particulate matter to the monitored value of the total mass concentration of particulate matter, and to determine whether the ratio is within a preset reasonable range, so as to verify the integrity of the particulate matter quality data. The characteristic element and volatile organic compound species comparison value verification submodule is used to monitor the normal deviation of the ratio of specific inorganic elements and the stability of the ratio of homologous volatile organic compound characteristic species pairs, so as to diagnose the anomalies of the sampling or detection system. The chemical and physical property correlation verification submodule is used to verify the stability of mass scattering efficiency between particulate matter scattering coefficient and particulate matter mass concentration and particle size distribution, as well as to verify the correlation between particulate matter concentration and relative humidity and visibility. The photochemical coupling verification submodule is used to incorporate free radical precursors into the photochemical system, verify the correlation between their diurnal concentration variation and solar radiation intensity, and analyze the spatiotemporal coupling relationship between the ozone generation potential calculated from volatile organic compounds and the measured ozone concentration.
3. The method for multi-source data cross-quality control and intelligent instrument fault diagnosis as described in claim 1, characterized in that, The specific processing steps of the machine learning-based anomaly pattern recognition module include: Use multi-source monitoring data under historical normal operating conditions to train time series prediction models or unsupervised learning models to learn the normal change patterns of each monitoring parameter and their interrelationships. The real-time collected data is input into the trained model to calculate the reconstruction error or prediction deviation. When the error or deviation exceeds the preset abnormal threshold, the corresponding data is marked as abnormal. When an anomaly is detected, contribution analysis is used to identify one or more variables that contribute the most to the anomaly, providing clues for tracing the source of the fault.
4. The method for multi-source data cross-quality control and intelligent instrument fault diagnosis as described in claim 1, characterized in that, A fusion decision-making mechanism is used to determine whether the current data anomaly is caused by a real contamination event or an instrument malfunction, including: When the multi-dimensional data shows consistency, conforms to the laws of atmospheric physicochemical processes, and the anomalies in the machine learning model are widespread and corroborated by meteorological conditions and data from surrounding stations, it is diagnosed as a real pollution event. When the total mass of particulate matter is normal but chemical components are missing, the comparison values of multiple volatile organic compounds show a systematic shift, or the concentration of reactive gases is consistently lower than the model prediction value and its correlation with other parameters is less than at least one of the set values, the corresponding instrument malfunction type is diagnosed.
5. The method for multi-source data cross-quality control and intelligent instrument fault diagnosis as described in claim 1, characterized in that, After the alarm is output, there is also a knowledge accumulation and model self-learning process: the diagnosis results confirmed by the operation and maintenance personnel are used as labeled data and fed back to the machine learning model and rule base for model retraining and dynamic adjustment of rule thresholds.
6. The method for multi-source data cross-quality control and intelligent instrument fault diagnosis as described in claim 1, characterized in that, After real-time collection and fusion of multi-source monitoring data from environmental monitoring superstations, and before data health assessment, the process also includes data preprocessing steps: time alignment of multi-source monitoring data from different sampling periods with a unified time benchmark, and interpolation or aggregation to fill in missing values in the data.
7. The method for multi-source data cross-quality control and intelligent instrument fault diagnosis as described in claim 1, characterized in that, The auxiliary parameter data includes: five meteorological parameters of the station, environmental parameters of the station building, and internal status parameters of each instrument; and data from different sources and formats are converted into standard formats through a unified data interface.
8. A multi-source data cross-quality control and instrument fault intelligent diagnosis system, characterized in that, include: The data acquisition and fusion module collects and fuses multi-source monitoring data from the environmental monitoring superstation in real time. The multi-source monitoring data includes conventional pollutant concentration data, particulate matter physicochemical properties data, volatile organic compound data, and auxiliary parameter data. The data health assessment module inputs the fused multi-source monitoring data into the intelligent data health assessment model in parallel. This model includes a rule base verification module based on physical and chemical correlations and an anomaly pattern recognition module based on machine learning, so as to perform multi-dimensional cross-validation and anomaly detection on the multi-source monitoring data respectively, and output the rule verification results and anomaly recognition results. The intelligent diagnosis and output module integrates rule verification results and anomaly identification results, and uses a fusion decision-making mechanism to determine whether the current data anomaly is caused by a real contamination event or an instrument malfunction, and traces and locates the type of instrument malfunction to generate a preliminary diagnostic conclusion. The preliminary diagnostic conclusions and related data information will be used to generate a visual diagnostic report and generate alarms.
9. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described in claims 1 to 7.
10. A computing device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described in claims 1 to 7.