Efficient data statistical processing method

By constructing a multi-dimensional quantitative indicator system and an improved hierarchical analysis method, combined with soil type correction coefficients, we have achieved accurate classification and processing of soil pollution monitoring data. This solves the problems of single-dimensional outlier judgment and poor adaptability in existing technologies, and improves the refinement and operability of data processing.

CN121958745APending Publication Date: 2026-05-01BEIJING MUNICIPAL RES INST OF ENVIRONMENT PROTECTION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING MUNICIPAL RES INST OF ENVIRONMENT PROTECTION
Filing Date
2026-01-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for processing soil pollution monitoring data are insufficient in terms of accuracy, efficiency, and applicability. They cannot effectively distinguish the intensity and scope of impact of outliers and lack quantitative indicators and standardized traceability records.

Method used

A multi-dimensional quantitative indicator system is constructed, and the weights are determined by an improved analytic hierarchy process. Combined with soil type correction coefficients, a comprehensive scoring formula is used to achieve quantitative classification of outliers, and a dedicated processing strategy is implemented to output a traceable report.

Benefits of technology

It improves the accuracy of outlier identification and the sophistication of data processing, adapts to different soil types and pollutant types, and meets the requirements of verifiability and traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958745A_ABST
    Figure CN121958745A_ABST
Patent Text Reader

Abstract

The invention provides an efficient data statistical processing method, and belongs to the technical field of soil environment monitoring data processing. The method comprises the following steps: cleaning detected soil pollutant concentration original data, and calculating a multi-dimensional quantitative index of an abnormal value; an improved analytic hierarchy process (AHP) is adopted to determine the weight of the multi-dimensional quantitative index; based on the multi-dimensional quantitative indexes and the corresponding weights, a comprehensive scoring formula is constructed, an abnormal value comprehensive score is calculated through the comprehensive scoring formula, and the comprehensive scoring formula comprises a concentration non-negativity punishment mechanism; classifying abnormal value grades according to the comprehensive score; executing corresponding processing strategies for different abnormal value levels; and outputting a traceable report containing abnormal value basic information, an index calculation process, an abnormal level judgment result and a processing record. According to the method, the problems that an existing method is single in abnormal value processing mode, poor in weight adaptability and free of a quantitative grading system are solved, and the accuracy of abnormal value grading and the refinement degree of data processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of soil environmental monitoring data processing technology, and in particular to an efficient data statistical processing method. Background Technology

[0002] Statistical processing of soil pollution concentration data is a core technical aspect of soil environmental monitoring and pollution control. Its accuracy and efficiency directly affect the scientific validity and effectiveness of control measures. Currently, with the continuous deepening of soil pollution prevention and control efforts, the amount of monitoring data is increasing dramatically, and the types of pollutants are becoming increasingly complex, covering multiple land use types such as agricultural land, residential land, public management and public service land, commercial service land, and industrial and mining land. Existing statistical methods are insufficient in terms of accuracy, efficiency, and applicability.

[0003] Existing soil monitoring data outlier handling techniques mostly rely on single statistical methods such as the Grubbs test and Dixon test, failing to consider the physical characteristics of soil data, such as the non-negativity of concentration and the continuity of regional gradients. This easily leads to misjudging valid high values ​​or missing systematic outliers. Simply classifying data as normal or abnormal fails to differentiate the intensity, scope of impact, and difficulty of correction of outliers, resulting in a "one-size-fits-all" approach, such as using the same correction method for minor and extreme outliers. The outlier determination process lacks quantitative indicators and cannot reproduce the grading logic. It fails to consider the differences between different soil types and pollutant species, resulting in poor universality of grading standards. Furthermore, the processing lacks standardized traceability documentation, making it difficult to effectively retain the identification basis, correction parameters, and calculation processes for outliers, thus failing to meet the requirements for verifiability and traceability of environmental monitoring data. Summary of the Invention

[0004] The purpose of this invention is to provide an efficient data statistical processing method. By constructing a multi-dimensional quantitative indicator system, an improved weight allocation algorithm, a comprehensive scoring hierarchical logic, and a dedicated processing strategy, it achieves a technological breakthrough in the qualitative judgment and quantitative classification of outliers, improves the accuracy of outlier classification and the refinement of data processing, and is adaptable to different soil types, pollutant types, and monitoring scenarios.

[0005] To achieve the above objectives, this invention proposes an efficient data statistical processing method, comprising the following steps: Step S1: Clean the raw data of detected soil pollutant concentrations and calculate multi-dimensional quantitative indicators for outliers. These multi-dimensional quantitative indicators include: statistical significance indicators, physical compliance indicators, and data impact indicators. Step S2: Use the improved analytic hierarchy process (AHP) to determine the weights of the multi-dimensional quantitative indicators. Specifically, the improved AHP involves introducing a soil type correction coefficient to dynamically adjust the importance of the multi-dimensional quantitative indicators. Step S3: Based on multi-dimensional quantitative indicators and corresponding weights, construct a comprehensive scoring formula, calculate the comprehensive score of outliers using the comprehensive scoring formula, and include a concentration non-negativity penalty mechanism. Step S4: Classify outlier levels based on the comprehensive score; Step S5: Implement corresponding handling strategies for different outlier levels; Step S6: Output a traceable report containing basic information about outliers, the calculation process of indicators, the results of the anomaly level determination, and the processing records.

[0006] Preferably, in step S1, the statistical significance index is the standardized Grubbs deviation value, calculated using the following formula: ; in, G std To standardize the Grubbs deviation value, G The Grubbs test statistic for outliers. G cr This represents the Grubbs critical value for the corresponding sample size. The physical compliance indicator is the regional concentration gradient violation rate, calculated using the following formula: ; in, S vio For the regional concentration gradient violation rate, S th For the concentration gradient threshold preset based on soil type, C i The pollutant concentration at the anomaly point. C j The pollutant concentrations at monitoring points adjacent to the anomaly point. D ij The distance between the anomaly point and its adjacent monitoring points; The data impact index is the neighborhood data disturbance rate, and the calculation formula is: ; in, P dis The neighborhood data perturbation rate, The value represents the average concentration of pollutants at five neighboring monitoring points around the anomaly point.

[0007] Preferably, step S2 includes the following steps: Step S21: Construct an indicator importance judgment matrix based on pollutant type, using the following formula: ; in, A This is a matrix for judging the importance of indicators. For the first m Class 1 indicators relative to the first n The importance of these indicators K The number of indicator categories; Step S22: Calculate the largest eigenvalue of the judgment matrix, and perform consistency checks using consistency index, random consistency index, and consistency ratio to determine the validity of the matrix; Step S23: Introduce a soil type correction coefficient to adjust the basic weights that have passed the consistency test, and calculate the final weights using the following formula: ; in, W k The final weights for multi-dimensional quantitative indicators, The basic weights for multi-dimensional quantitative indicators to pass the consistency test. k Number the indicator category. δ This is the soil type correction factor.

[0008] Preferably, when the monitoring object is heavy metal pollutants, the corresponding basic weights are statistical significance index 0.3, physical compliance index 0.4, and data impact index 0.3.

[0009] Preferably, in step S3, the comprehensive scoring formula is: ; Where Y is the comprehensive score of outliers. , X k For the first k The original calculated value of the category indicator, I This is the highest score in the overall outlier assessment.

[0010] Preferably, if the outlier concentration is less than zero, the outlier comprehensive score is increased. v The specific division is as follows: .

[0011] Preferably, in step S4, the outlier levels include: minor outlier, moderate outlier, severe outlier, and extreme outlier; the classification criteria for outlier levels are: Level I Minor Outlier: 0 <Y≤ Grade II moderate abnormality: <Y≤ Grade III severe abnormality: <Y≤ Level IV extreme anomaly: <Y≤ .

[0012] Preferably, in step S5, corresponding processing strategies are implemented for different outlier levels, specifically: Level I minor outliers are marked with only the outlier attributes and the original data is retained; Level II moderate outliers are corrected using the LOWESS local weighted regression formula; Level III severe outliers are corrected using the regional mean interpolation formula; and Level IV extreme outliers are removed and the data is reconstructed based on the spatial weighted mean of the surrounding effective monitoring points.

[0013] Preferably, the LOWESS locally weighted regression formula is used to correct for level II moderate outliers, and the Gaussian kernel function is used to calculate the weights. The calculation formula is as follows: ; ; in, These are pollutant concentration values ​​corrected for moderate outliers. Q This refers to the number of monitoring points within a local window, i.e., the width of the local window. q The numbers of the nearby monitoring points, For the first q The weight of each neighboring monitoring point C q For the first q Pollutant concentration values ​​at neighboring monitoring points d q For moderate outliers and the first q The spatial straight-line distance between neighboring monitoring points h For kernel function bandwidth, max (·) represents the maximum value; The regional mean interpolation formula was used to correct the Level III severe outlier values. The formula is as follows: ; in, These are the pollutant concentration values ​​corrected for severe anomalies. This represents the average pollutant concentration at all valid monitoring points within the sub-region where the anomaly point is located. This is the system outlier correction factor. Determined based on soil type; After removing Level IV extreme outliers, the data is reconstructed based on the spatially weighted mean of surrounding valid monitoring points, using the following formula: ; ; in, These are the pollutant concentration values ​​reconstructed after removing extreme outliers. u Number the valid monitoring points around the extreme outliers. For the first u Spatial weights of each surrounding effective monitoring point Cu For the first u Pollutant concentration values ​​at several effective monitoring points in the surrounding area. U This refers to the number of non-abnormal monitoring points selected that are closest to the abnormal point. d u For extreme outliers and the first u The spatial straight-line distance between the surrounding effective monitoring points.

[0014] Preferably, in step S6, the traceability report includes the following basic information about the anomaly: monitoring area, coordinates, pollutant type and original concentration of the anomaly; the indicator calculation process includes: specific values ​​and calculation basis of statistical significance indicators, specific values ​​and calculation basis of physical compliance indicators and data impact indicators; the grade determination result includes comprehensive score, final grade and judgment basis; and the processing record includes correction algorithm, key parameters and concentration comparison before and after processing.

[0015] Therefore, this invention proposes an efficient data statistical processing method, the beneficial effects of which are as follows: (1) This invention integrates three types of indicators: statistical significance, physical compliance, and data impact, and constructs a multi-dimensional indicator system, which solves the problem of misjudgment of effective data or omission of abnormal data and improves the accuracy of outlier identification.

[0016] (2) This invention introduces a soil type correction coefficient to optimize AHP weight allocation, adapting to different soil types and pollutant types, and meeting the needs of multiple scenarios, thus avoiding the problem of poor adaptability caused by fixed weights.

[0017] (3) This invention achieves gradient classification of outlier intensity through a comprehensive scoring formula, combined with a dedicated processing strategy, to avoid data distortion caused by correction and achieve refined data processing.

[0018] (4) This invention clarifies the values ​​and calculation logic of various parameters, requiring no professional programming skills and can be completed using only basic statistical tools. It is highly operable and improves the operational efficiency of grassroots monitoring agencies.

[0019] (5) This invention fully records the entire process of indicator calculation, weight determination, and hierarchical processing, and has traceability, meeting industry needs. Attached Figure Description

[0020] Figure 1 This is a flowchart of an efficient data statistical processing method. Detailed Implementation

[0021] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0023] Example 1 The application scenario is based on the processing of soil pollution monitoring data from urban industrial heritage construction sites. This site was originally used for chemical production and is now planned for Class II land use.

[0024] The total monitoring area is 4.8 km². 2 The monitoring units were divided into 100m × 100m grids, with a total of 480 monitoring points. Soil sample testing revealed the soil type to be brown soil, with a porosity of 43% and a permeability coefficient of 3.5 × 10⁻⁶. -6 The concentration of cadmium (Cd) was measured at m / s, which is consistent with the basic properties of brown soil. The pollutant detected was cadmium (Cd), a heavy metal pollutant. The screening value for Class II land use was 65 mg / kg, and the control value was 140 mg / kg. The original data of cadmium concentration from 480 monitoring points were collected, totaling 480 data points, with a data range of 0.2-89.6 mg / kg. After preliminary cleaning and removal of 3 duplicate data points, 477 valid data points were obtained.

[0025] like Figure 1 As shown, this invention provides an efficient data statistical processing method, specifically including the following steps: Step S1: Clean the raw data of detected soil pollutant concentrations, screen out 12 suspected outliers, number them Y1-Y12, and calculate multi-dimensional quantitative indicators for the 12 suspected outliers. These multi-dimensional quantitative indicators include: statistical significance indicators, physical compliance indicators, and data impact indicators, specifically: The statistical significance index is the standardized Grubbs bias value, calculated using the following formula: ; in, G std To standardize the Grubbs deviation value, G The Grubbs test statistic for outliers. G cr The Grubbs critical value is for a sample size of 477. The Grubbs critical value is 3.28, as determined from the Grubbs test critical value table. The physical compliance indicator is the regional concentration gradient violation rate, calculated using the following formula: ; in, S vio For the regional concentration gradient violation rate, S thThe concentration gradient threshold, preset based on soil type, is set to 0.10 mg / (kg·m). C i The pollutant concentration at the anomaly point. C j The pollutant concentrations at monitoring points adjacent to the anomaly point. D ij The distance between the anomaly point and its adjacent monitoring points; The data impact index is the neighborhood data disturbance rate, and the calculation formula is: ; in, P dis The neighborhood data perturbation rate, The values ​​represent the average concentrations of pollutants at five neighboring monitoring points around the anomaly point; the calculation results of multi-dimensional indicators for suspected anomalies are shown in Table 1. Table 1. Calculation results of multi-dimensional indicators for suspected anomalies

[0026] Step S2: Determine the weights of the multi-dimensional quantitative indicators using the improved Analytic Hierarchy Process (AHP). Specifically, the improved AHP involves introducing a soil type correction coefficient to dynamically adjust the importance of the multi-dimensional quantitative indicators, including the following steps: Step S21: Construct an indicator importance judgment matrix based on pollutant type, using the following formula: ; in, A This is a matrix for judging the importance of indicators. For the first m Class 1 indicators relative to the first n The importance of these indicators K The number of indicator categories; Step S22: The maximum eigenvalue of the judgment matrix is ​​3.009. The consistency index is 0.0045, the random consistency index is 0.58, and the consistency ratio is 0.0078, which is less than 0.1. The matrix is ​​then judged to be valid. Step S23: Introduce a soil type correction coefficient to adjust the basic weights that passed the consistency test. Normalize the judgment matrix and calculate the eigenvectors to obtain the basic weights: statistical significance index 0.3, physical compliance index 0.4, and data impact index 0.3. Calculate the final weights using the following formula: ; in, W k The final weights for multi-dimensional quantitative indicators, The basic weights for multi-dimensional quantitative indicators to pass the consistency test.k is the index category number, δ is the soil type correction coefficient. The soil type in this area is cinnamon soil, and the correction coefficient is 1.0. The final weights are 0.3 for the statistical significance index, 0.4 for the physical compliance index, and 0.3 for the data impact index.

[0027] Step S3: Based on the multi-dimensional quantification indicators and their corresponding weights, construct a comprehensive scoring formula, and calculate the comprehensive score of the outlier through the comprehensive scoring formula. The comprehensive scoring formula includes a concentration non-negativity penalty mechanism, and the comprehensive scoring formula is: ; where Y is the comprehensive score result of the outlier, , V take 100, X k is the k original calculated value of the I th type of indicator, v is the highest comprehensive score of the outlier; if the concentration of the outlier point is less than zero, the comprehensive score result of the outlier is increased by points, specifically: <oo00247>Table 2 Comprehensive scores of outlier points

[0028] Step S4: Divide the outlier levels according to the comprehensive score. The outlier levels include: slight anomaly, moderate anomaly, severe anomaly, and extreme anomaly; the classification criteria for the anomaly levels are: Level I slight anomaly: 0 < Y ≤ 20; Level II moderate anomaly: 20 < Y ≤ 50; Level III severe anomaly: 50 < Y ≤ 80; Level IV extreme anomaly: 80 < Y ≤ 100; the classification results are shown in Table 3: Table 3 Judgment results of outlier levels

[0029] Step S5: Execute corresponding processing strategies for different outlier levels; specifically: Level I slight outlier values only mark the outlier attributes and retain the original data, Level II moderate outlier values are corrected using the LOWESS locally weighted regression formula, Level III severe outlier values are corrected using the regional mean interpolation formula, and after the Level IV extreme outlier values are removed, the data is reconstructed based on the spatial weighted mean of the surrounding valid monitoring points; Use the LOWESS locally weighted regression formula to correct the Level II moderate outlier values, and calculate the weights using the Gaussian kernel function. The calculation formula is: ; [[ID=**47]] ; where, is the pollutant concentration value after correcting the moderate outlier point,Q The number of monitoring points within the local window, i.e., the width of the local window, is taken as... Q =5, q The numbers of the nearby monitoring points, For the first q The weight of each neighboring monitoring point C q For the first q Pollutant concentration values ​​at neighboring monitoring points d q For moderate outliers and the first q The spatial straight-line distance between neighboring monitoring points h For kernel function bandwidth, max (·) represents the maximum value; The regional mean interpolation formula was used to correct the Level III severe outlier values. The formula is as follows: ; in, These are the pollutant concentration values ​​corrected for severe anomalies. This represents the average pollutant concentration at all valid monitoring points within the sub-region where the anomaly point is located. This is the system outlier correction factor. Determined based on soil type; After removing Level IV extreme outliers, the data is reconstructed based on the spatially weighted mean of surrounding valid monitoring points, using the following formula: ; ; in, These are the pollutant concentration values ​​reconstructed after removing extreme outliers. u Number the valid monitoring points around the extreme outliers. For the first u Spatial weights of each surrounding effective monitoring point C u For the first u Pollutant concentration values ​​at several effective monitoring points in the surrounding area. U This refers to the number of non-abnormal monitoring points selected that are closest to the abnormal point. d u For extreme outliers and the first u The spatial straight-line distance between the surrounding effective monitoring points.

[0030] The corrected concentration results are shown in Table 4: Table 4. Corrected concentration results

[0031] Step S6: Output a traceable report containing basic information about outliers, the calculation process of the indicators, the results of the anomaly level determination, and the processing records, specifically: Basic information about anomalies includes: monitoring area, coordinates, pollutant type, and original concentration of the anomaly. The indicator calculation process includes: the specific values ​​and calculation basis of the statistical significance indicator, the specific values ​​and calculation basis of the physical compliance indicator, and the specific values ​​and calculation basis of the data impact indicator; The grading results include a comprehensive score, a final grade, and the basis for the judgment. The processing records include the correction algorithm, key parameters, and a comparison of concentrations before and after processing.

[0032] It is worth noting that all contents not described in detail in this invention are existing technologies and are well known to those skilled in the art.

[0033] Therefore, this invention provides an efficient data statistical processing method. It calculates three quantitative indicators—statistical significance, physical compliance, and data impact—and uses an improved analytic hierarchy process (AHP) incorporating a soil type correction coefficient to determine indicator weights. A comprehensive scoring formula with a concentration-based non-negative penalty mechanism is then used to obtain a comprehensive outlier score. This results in four outlier levels, the implementation of specific processing strategies, and the output of a traceable report. This invention solves the problems of existing technologies, such as limited outlier judgment dimensions, poor weight adaptability, and the lack of a quantitative grading system. It improves the accuracy of outlier grading and the refinement of data processing, is adaptable to various scenarios including construction land and industrial sites, and is suitable for monitoring heavy metals and organic pollutants. It possesses strong practicality and operability.

[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A highly efficient data statistical processing method, characterized in that, Includes the following steps: Step S1: Clean the raw data of detected soil pollutant concentrations and calculate multi-dimensional quantitative indicators for outliers. These multi-dimensional quantitative indicators include: statistical significance indicators, physical compliance indicators, and data impact indicators. Step S2: Use the improved analytic hierarchy process (AHP) to determine the weights of the multi-dimensional quantitative indicators. Specifically, the improved AHP involves introducing a soil type correction coefficient to dynamically adjust the importance of the multi-dimensional quantitative indicators. Step S3: Based on multi-dimensional quantitative indicators and corresponding weights, construct a comprehensive scoring formula, calculate the comprehensive score of outliers using the comprehensive scoring formula, and include a concentration non-negativity penalty mechanism. Step S4: Classify outlier levels based on the comprehensive score; Step S5: Implement corresponding handling strategies for different outlier levels; Step S6: Output a traceable report containing basic information about outliers, the calculation process of indicators, the results of the anomaly level determination, and the processing records.

2. The efficient data statistical processing method according to claim 1, characterized in that: In step S1, the statistical significance index is the standardized Grubbs bias value, calculated using the following formula: ; in, G std To standardize the Grubbs deviation value, G This is the Grubbs test statistic for outliers. G cr This represents the Grubbs critical value for the corresponding sample size. The physical compliance indicator is the regional concentration gradient violation rate, calculated using the following formula: ; in, S vio For the regional concentration gradient violation rate, S th For the concentration gradient threshold preset based on soil type, C i The pollutant concentration at the anomaly point. C j The pollutant concentrations at monitoring points adjacent to the anomaly point. D ij The distance between the anomaly point and its adjacent monitoring points; The data impact index is the neighborhood data disturbance rate, and the calculation formula is: ; in, P dis The neighborhood data perturbation rate, The value represents the average concentration of pollutants at five neighboring monitoring points around the anomaly point.

3. The efficient data statistical processing method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: Construct an indicator importance judgment matrix based on pollutant type, using the following formula: ; in, A This is a matrix for judging the importance of indicators. For the first m Class 1 indicators relative to the first n The importance of these indicators K The number of indicator categories; Step S22: Calculate the largest eigenvalue of the judgment matrix, and perform consistency checks using consistency index, random consistency index, and consistency ratio to determine the validity of the matrix; Step S23: Introduce a soil type correction coefficient to adjust the basic weights that have passed the consistency test, and calculate the final weights using the following formula: ; in, W k The final weights for multi-dimensional quantitative indicators, The basic weights for multi-dimensional quantitative indicators to pass the consistency test. k Number the indicator category. δ This is the soil type correction factor.

4. The efficient data statistical processing method according to claim 3, characterized in that: When the monitoring target is heavy metal pollutants, the corresponding basic weights are statistical significance index 0.3, physical compliance index 0.4, and data impact index 0.

3.

5. The efficient data statistical processing method according to claim 1, characterized in that: In step S3, the comprehensive scoring formula is: ; Where Y is the comprehensive score of outliers. , X k For the first k The original calculated value of the category indicator, I This is the highest score in the overall outlier assessment.

6. The efficient data statistical processing method according to claim 5, characterized in that: If the outlier concentration is less than zero, the outlier comprehensive score will be added. v The specific division is as follows: .

7. The efficient data statistical processing method according to claim 1, characterized in that: In step S4, the outlier levels include: minor outlier, moderate outlier, severe outlier, and extreme outlier; the classification criteria for outlier levels are: Level I Minor Outlier: 0 <Y≤ Grade II moderate abnormality: <Y≤ Grade III severe abnormality: <Y≤ Level IV extreme anomaly: <Y≤ .

8. The efficient data statistical processing method according to claim 1, characterized in that: In step S5, corresponding processing strategies are implemented for different outlier levels. Specifically, for Level I minor outliers, only the outlier attributes are marked and the original data is retained; for Level II moderate outliers, the LOWESS local weighted regression formula is used for correction; for Level III severe outliers, the regional mean interpolation formula is used for correction; and for Level IV extreme outliers, after removal, the data is reconstructed based on the spatial weighted mean of the surrounding valid monitoring points.

9. The efficient data statistical processing method according to claim 8, characterized in that: The Lowess locally weighted regression formula was used to correct for level II moderate outliers, and the Gaussian kernel function was used to calculate the weights. The calculation formula is as follows: ; ; in, These are pollutant concentration values ​​corrected for moderate outliers. Q This refers to the number of monitoring points within a local window, i.e., the width of the local window. q The numbers of the nearby monitoring points, For the first q The weight of each neighboring monitoring point C q For the first q Pollutant concentration values ​​at neighboring monitoring points d q For moderate outliers and the first q The spatial straight-line distance between neighboring monitoring points h For kernel function bandwidth, max (·) represents the maximum value; The regional mean interpolation formula was used to correct the Level III severe outlier values. The formula is as follows: ; in, These are the pollutant concentration values ​​corrected for severe anomalies. This represents the average pollutant concentration at all valid monitoring points within the sub-region where the anomaly point is located. This is the system outlier correction factor. Determined based on soil type; After removing Level IV extreme outliers, the data is reconstructed based on the spatially weighted mean of surrounding valid monitoring points, using the following formula: ; ; in, These are the pollutant concentration values ​​reconstructed after removing extreme outliers. u Number the valid monitoring points around the extreme outliers. For the first u Spatial weights of each surrounding effective monitoring point C u For the first u Pollutant concentration values ​​at several effective monitoring points in the surrounding area. U This refers to the number of non-abnormal monitoring points selected that are closest to the abnormal point. d u For extreme outliers and the first u The spatial straight-line distance between the surrounding effective monitoring points.

10. The efficient data statistical processing method according to claim 1, characterized in that: In step S6, the traceability report includes the following basic information about anomalies: monitoring area, coordinates, pollutant type, and original concentration of the anomaly. The indicator calculation process includes the specific values ​​and calculation basis of statistical significance indicators, physical compliance indicators, and data impact indicators. The grade determination results include the comprehensive score, final grade, and judgment basis. The processing record includes the correction algorithm, key parameters, and concentration comparison before and after processing.