A grid air pollution concentration data space heterogeneity evaluation method

By constructing site pairs and distance segments, and combining logarithmic distance with equal weights, the multi-scale problem of assessing the spatial heterogeneity of grid air pollution concentration data was solved, enabling accurate diagnosis and evaluation of data at different spatial scales.

CN122414883APending Publication Date: 2026-07-17NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies cannot effectively assess the spatial heterogeneity of grid-based air pollution concentration data. They can only reflect the overall accuracy but cannot assess and integrate the results by distance segment to measure the spatial heterogeneity of the data.

Method used

By acquiring gridded air pollution concentration data and ground station air pollution concentration data within the same geographical area, station pairs are constructed, concentration differences and geographical distances are calculated, and distance segments are divided using specific rules. Correlation and error indices are calculated, and a comprehensive evaluation is conducted by combining logarithmic distance with equal weights.

Benefits of technology

It enables accurate multi-scale diagnosis of gridded air pollution concentration data, fairly and balancedly evaluates the performance of data at different spatial scales, and provides a more accurate and reliable method for assessing spatial heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122414883A_ABST
    Figure CN122414883A_ABST
Patent Text Reader

Abstract

This invention discloses a method for evaluating the spatial heterogeneity of gridded air pollution concentration data, belonging to the field of environmental monitoring. The method includes: acquiring gridded data and station data for the same area; constructing station pairs and calculating concentration differences and distances; dividing the station pairs into multiple consecutive distance segments by equal number of bins based on distance; within each distance segment, calculating correlation indicators and normalized error indicators reflecting concentration gradient differences based on valid station pairs; finally, calculating weights based on the logarithmic span of each distance segment, and weighting and integrating the segmented indicators to obtain a comprehensive error indicator and a comprehensive correlation indicator for evaluating the spatial heterogeneity of the gridded data. This invention achieves refined evaluation through "distance binning" and ensures the balance and scientific validity of the evaluation conclusions through a "logarithmic distance equal weighting" strategy, providing a precise and reliable evaluation benchmark for the spatial heterogeneity quality of gridded data products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of environmental monitoring and environmental management, and specifically relates to a method for evaluating the spatial heterogeneity of grid-based air pollution concentration data. Background Technology

[0002] Air pollution concentration data is an important support for environmental management. Air pollution concentration data can be divided into two categories: (1) point-based air pollution concentration data, such as air pollution concentration data based on observation stations; (2) grid-based air pollution concentration data. The basic unit of this type of data is a grid, and a single grid covers a certain spatial range (such as 1 km × 1 km). The grid air pollution concentration value represents the average pollution concentration of that spatial range. The above two types of air pollution concentration data each have their own advantages and disadvantages. For air pollution concentration data based on observation stations, the data is usually more accurate and the observation frequency is higher (usually hourly observations are possible). However, the station air pollution concentration data usually only represents a limited spatial range near the station. For grid-based air pollution concentration data, it is usually a fusion of multi-source data such as satellite observations and meteorological fields. The inversion algorithm is used to generate air pollution concentration values ​​covering a large area. The spatial resolution of the air pollution concentration value corresponds to the grid spatial coverage range (such as 1 km × 1 km). Although grid air pollution concentration data can cover a large area, the observation frequency is low (usually only daily averages are provided). Nevertheless, the grid air pollution concentration data has advantages and disadvantages. Data remains widely used in environmental management, health assessment, and other fields. Before using gridded air pollution concentration data, it is necessary to evaluate its accuracy. Existing methods typically rely on station observation data combined with quantitative indicators such as correlation coefficients and root mean square error to evaluate the accuracy of gridded air pollution concentration data. However, these quantitative indicators only reflect the overall accuracy of the data and do not provide an assessment of its spatial heterogeneity. Spatial heterogeneity refers to the uneven and significant differences in air pollutant concentrations across different locations due to variations in emission source distribution, topography and meteorological conditions, land use, and human activities. In summary, existing technologies have problems including relying on station-level air pollution concentrations to evaluate gridded air pollution concentration data, and their evaluation indicators only reflect overall accuracy, failing to assess and integrate results for different distance segments to measure the spatial heterogeneity of the data. Summary of the Invention

[0003] This invention addresses the problems existing in the prior art by providing a method for evaluating the spatial heterogeneity of gridded air pollution concentration data. This method can assess the spatial heterogeneity of gridded air pollution concentration data for different distance segments and can ultimately integrate the evaluation results of different distance segments into a single evaluation result.

[0004] To address the above technical problems, this invention provides the following technical solution: a method for evaluating the spatial heterogeneity of gridded air pollution concentration data, comprising the following steps:

[0005] Step 1: Basic data preparation and site pair construction: Obtain gridded air pollution concentration data and ground station air pollution concentration data within the same geographical area, construct all site pairs, and calculate the concentration difference and geographical distance for each site pair.

[0006] Preferably, step 1 includes the following sub-steps:

[0007] Step 1.1: Data Acquisition and Alignment

[0008] Two sets of data are obtained within the same geographical area: (1) gridded air pollution concentration data; (2) ground station air pollution concentration data. According to the time scale type (annual scale / monthly scale / daily scale), the datasets within the corresponding time period are extracted to ensure that the grid data is accurately aligned with the geographical location of each station, forming a sequence of air pollution concentration data pairs at the station locations.

[0009] Step 1.2: Screening the validity of air pollution concentration data at the monitoring station

[0010] Perform time-scale quality control on the air pollution concentration data of the stations.

[0011] For annual scale: For each station, count the number of valid days of air pollution within one year and set a validity threshold (e.g., valid days percentage ≥ 75%). For monthly scale: For each station, count the number of valid days of air pollution within a single month and set dual validity thresholds for each month: For months with 30 / 31 days: valid days ≥ 20 days and valid days percentage ≥ 75% (approximately 0.75 × number of days in the month); For months with 28 / 29 days: valid days ≥ 18 days and valid days percentage ≥ 65% (approximately 0.65 × number of days in the month). For daily scale: For each station, analyze the data for each day.

[0012] Set a validity threshold (e.g., valid days percentage ≥ 75%) to mark a site as a "valid site" (otherwise it will be an "invalid site"), and retain complete spatial distance structure information.

[0013] Step 1.3: Calculation of mean values ​​and generation of site pairs at different time scales

[0014] For the selected timescale type, calculate the mean and station pairs within the corresponding period:

[0015] For an annual scale: For each station, calculate the average of its gridded air pollution concentration and the average of the air pollution concentration at ground stations over one year to obtain the annual average air pollution concentration at ground stations. and gridded annual average air pollution concentration For monthly scales: For each month and each station, calculate the average of its gridded air pollution concentration and the average of the air pollution concentration at ground stations within that month to obtain the monthly average air pollution concentration at ground stations. and gridded monthly average air pollution concentration (m represents the month, such as January to December); If it is a daily scale: for each single day and each station, calculate the average of the ground gridded air pollution concentration and the average of the ground station air pollution concentration within that day to obtain the daily average air pollution concentration of the ground station. and gridded daily average air pollution concentration (d represents the date).

[0016] All within the region Each site is paired up (excluding itself) to generate... A unique site pair .

[0017] Step 1.4: Difference Sequence and Distance Calculation

[0018] For each site pair ,calculate:

[0019] (a) Air pollution concentration difference at ground stations ;

[0020] (b) Gridded air pollution concentration difference ;

[0021] (c) Geographical distance between sites High-precision great circle distance is adopted Formula calculation.

[0022] Simultaneously, record the validity status of the two stations in the station pair. Step 2: Multi-scale distance segmentation based on specific rules: Divide all station pairs into K consecutive distance segments by equal number of bins according to their geographical distance; execute the steps of this module independently for each time scale (processing year / month / day scales separately).

[0023] Preferably, step 2 includes the following sub-steps:

[0024] Step 2.1: Set the lower limit for distance analysis

[0025] Define minimum effective analysis distance km, all Site pairs marked as "near-distance removal pairs" will not participate in subsequent binning analysis.

[0026] Step 2.2: Divide into equal-quantity boxes and determine boundaries

[0027] To satisfy There are station pairs, according to their distance Sort in ascending order. Divide it into groups using the "specific rule equal quantity binning method". A series of consecutive distance segments ( (Can be set to 10). The binning rule is: find the quantile points of distance such that the total number of station pairs contained in each bin is approximately equal. Specifically, the distance values ​​corresponding to the quantiles of [0, 10%, 20%, ..., 100%] are used as the initial bin boundaries. Subsequently, a critical correction is performed: the lower boundary of the first bin is forcibly set to... This ensures that the analysis strictly begins from the set lower limit. The upper boundary of each bin is the maximum distance plus a small increment. Step 2.3: Filtering valid data within the distance segment.

[0028] For the distance segments ( =1, 2, ..., ), filter out site pairs that meet both of the following conditions:

[0029] (a) its distance It falls within the boundary of that distance segment.

[0030] (b) Site and sites All of these were marked as "valid stations" in step 1.2. These station pairs constitute the "set of valid analysis station pairs" for this distance segment, and their corresponding... and The sequence will be used for index calculations within this distance segment. Let the number of station pairs in this set be denoted as . (monthly scale) The daily scale is ) .

[0031] Step 3: Calculation of core heterogeneity index for distance segment: Within each distance segment, select two station pairs where both stations are valid stations as valid analysis station pairs. Based on the concentration difference data of the valid analysis station pairs, calculate the heterogeneity index for that distance segment. The heterogeneity index includes a correlation index reflecting the correlation of concentration gradients and a normalized error index reflecting the error of concentration gradients.

[0032] Preferably, step 3 specifically includes the following sub-steps:

[0033] Step 3.1: Calculation of correlation indicators:

[0034] For the A distance segment, based on its right Data, calculations: (a) Rank correlation coefficient The assessment evaluates the consistency of the gridded air pollution concentration gradient with the station-level air pollution concentration gradient in terms of direction and order, and is insensitive to outliers. (b) Correlation coefficient To assess the strength of the linear correlation between the gridded air pollution concentration gradient and the site-specific air pollution concentration gradient. Step 3.2: Error index calculation:

[0035] For the For each distance segment, calculate: (a) Root mean square error : (a) The absolute error between the gridded air pollution concentration gradient and the station air pollution concentration gradient. : This represents the variation in air pollution concentration at stations within this distance range. (c) Normalized root mean square error : This indicator is dimensionless and standardizes the error relative to the variation of air pollution concentration at the station, making the errors between different distance segments comparable. It can be directly interpreted as "the proportion of the gridded air pollution concentration gradient error to the station's air pollution concentration gradient". Step 4: Comprehensive evaluation based on logarithmic distance equal weight and statistical transformation: Based on the distance span of each distance segment on the logarithmic scale, calculate the integration weight of each distance segment, and use the weight to perform weighted integration of the correlation index and normalized error index of all distance segments to obtain the comprehensive correlation index and comprehensive error index used to evaluate the spatial heterogeneity of grid air pollution concentration data.

[0036] Preferably, step 4 specifically includes the following sub-steps:

[0037] Step 4.1: Calculation of logarithmic distance equal weights

[0038] For each distance segment Calculate a combined weight This weight does not depend on the sample size. It is based not on the distance segment's span on a logarithmic scale, but on the distance segment's length.

[0039]

[0040]

[0041] Logarithmic distance equal weighting assigns equal weights to equal-length intervals on the logarithmic distance axis, ensuring that each order of magnitude scale from 1 km to 1000 km (e.g., 1-10 km, 10-100 km, 100-1000 km) has similar weight in the final composite index. This allows for a fair and balanced evaluation of the data's performance across all spatial scales, overcoming the problem of long-distance segment dominance that may result from simple sample size weighting.

[0042] Step 4.2: Comprehensive Error Index calculate

[0043] It reflects the overall gradient error level of grid data across multiple scales.

[0044]

[0045] It is a dimensionless comprehensive index. The smaller its value, the more accurately the gridded air pollution describes the concentration gradient across all spatial scales, and the stronger its ability to reproduce spatial heterogeneity.

[0046] Step 4.3: Comprehensive correlation indicators calculate

[0047] This reflects the level of consistency between gridded air pollution concentrations and station-level air pollution concentration gradient changes across multiple scales. It employs a method based on... Robust methods for transformation:

[0048] (a) For each distance segment Correlation coefficient (or Correlation coefficient )conduct Transformation:

[0049] (b) Using the log-distance equal weights calculated in step 4_1 ,calculate Weighted average: .

[0050] (c) Perform the reverse Transformation yields the comprehensive correlation coefficient: .

[0051] Step 4.4: Obtain core results for final evaluation

[0052] (a) Scale-based diagnostic report:

[0053] Includes the distance range for each distance segment, , , , It is used to identify the advantages and disadvantages of gridded air pollution concentration at a specific scale.

[0054] (b) Two core comprehensive indicators:

[0055] (Comprehensive gradient error): The theoretical range is [0, +∞), and in practice, the smaller the better. Different datasets or models can be directly compared. (Comprehensive gradient correlation): The value range is [-1, 1], and the closer to 1, the better. It indicates the overall consistency between the gridded spatial variation pattern of air pollution and the changing trend of air pollution at the monitoring stations.

[0056] Compared with the prior art, the beneficial technical effects of the present invention using the above technical solution are as follows:

[0057] 1. This invention achieves multi-scale accurate diagnosis of spatial heterogeneity. Existing technologies only provide an overall accuracy score and cannot distinguish the differences in data performance at different spatial scales. This invention decomposes the evaluation into multiple distinct spatial scales through "site-to-site difference analysis" and "distance binning".

[0058] 2. This invention employs a "logarithmic distance equal weighting" integration strategy, solving the challenge of scientifically integrating multi-scale evaluation results. Traditional methods, when attempting multi-scale evaluation, typically use arithmetic mean or sample size weighting when integrating results. This can lead to biased conclusions due to the large number of long-distance site pairs, masking shortcomings at shorter distances. The core innovation of this invention is the calculation of weights for each distance segment using "logarithmic distance equal weighting," assigning equal weights to intervals of equal length on the logarithmic distance coordinate system. Experimental comparisons show that for the same data, the traditional sample size weighting gives a lower overall correlation coefficient (…). The correlation coefficient is 0.375, while the comprehensive correlation coefficient given by the method of this invention is... The value was 0.267, which, by balancing the discourse power of each scale, more truthfully and fairly reflects the overall situation of the data's poor performance at short-range scales.

[0059] In summary, this invention achieves refined evaluation through "distance binning" and ensures the balance and scientific nature of the evaluation conclusions through "logarithmic distance equal weighting," providing a more accurate and reliable "benchmark" for the spatial heterogeneity quality of grid air pollution data products, and effectively guiding the improvement and applicability judgment of data products. Attached Figure Description

[0060] Figure 1 This is a flowchart of the present invention.

[0061] Figure 2This is a schematic diagram illustrating how the correlation changes with distance in the embodiment.

[0062] Figure 3 This is a schematic diagram illustrating how the error index changes with distance in the embodiment. Detailed Implementation

[0063] To better understand the technical content of the present invention, specific embodiments are described below in conjunction with the accompanying drawings.

[0064] In this invention, various aspects of the invention are described with reference to the accompanying drawings, in which numerous illustrative embodiments are shown. Embodiments of the invention are not limited to those depicted in the drawings. It should be understood that the invention is implemented through any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed herein are not limited to any particular implementation. Furthermore, some aspects of the invention disclosed may be used alone or in any suitable combination with other aspects of the invention disclosed.

[0065] like Figure 1 As shown, taking the PM2.5 pollution in a certain city in 2020 as an example, daily PM2.5 concentration data from 171 stations in the city were collected. The gridded PM2.5 pollution concentration data came from the CHAP public dataset (https: / / zenodo.org / records / 6398971), and the spatial resolution of the gridded data was 1 km × 1 km. Using the technical framework of this invention, the spatial heterogeneity of the gridded PM2.5 pollution concentration data was evaluated using the station PM2.5 concentration data. A specific example is as follows:

[0066] (I) Basic Data Preparation and Site Pair Construction

[0067] Obtain gridded air pollution concentration data and ground-based station air pollution concentration data within the same geographic area, construct all station pairs, and calculate the concentration difference and geographic distance for each station pair. Specifically, this includes the following sub-steps:

[0068] Step 1.1: Data Acquisition and Alignment

[0069] Ground-based air pollution concentration data, gridded air pollution concentration data, and station information were read and verified. Verification showed that all three datasets contained 171 stations, spanned 366 days, and were spatially and temporally aligned, forming a sequence of data pairs at each station location.

[0070] Step 1.2: Screening the validity of air pollution concentration data at the monitoring station

[0071] Replace the invalid value -999 in the data with NaN. Count the number of valid days for each site, and set the validity threshold to more than 275 valid days (approximately 75% of the total days). As a result, 169 sites met the criteria for both ground and gridded data and were marked as "valid sites"; the remaining 2 sites were marked as "invalid sites" due to significant data loss.

[0072] This embodiment uses an annual timescale as an example. For each "effective station," the average value of its ground-based air pollution concentration and the gridded air pollution concentration over 366 days is calculated to obtain the annual average value of the ground-based station. and gridded annual average All 171 sites were paired up to generate... A unique site pair .

[0073] Step 1.3: Difference Sequence and Distance Calculation

[0074] For each site pair ,calculate:

[0075] a) Air pollution concentration difference at ground stations ;

[0076] b) Gridded air pollution concentration difference ;

[0077] c) Geographical distance between sites ,use The formula is based on the latitude and longitude of the station.

[0078] At the same time, the validity status of the two sites in the middle is recorded.

[0079] (II) Multi-scale distance segmentation based on specific rules

[0080] All site pairs are binned equally according to their geographical distance, forming K consecutive distance segments; the steps of this module are executed independently for each time scale (year / month / day scales are processed separately), specifically including the following sub-steps:

[0081] Step 2.1: Set the lower limit for distance analysis

[0082] Define minimum effective analysis distance The minimum distance between all station pairs was calculated to be 0.742 kilometers. One station pair with a distance of less than 1 kilometer was marked as a "near-distance exclusion pair" and will not be included in the subsequent binning analysis.

[0083] Step 2.2: Divide into equal-quantity boxes and determine boundaries

[0084] For the remainder One satisfies The stations are arranged according to their distance. Sort in ascending order. Divide the data into 10 segments using an equal-number binning method (each segment containing approximately 10% of the site pairs). Then, forcibly adjust the lower boundary of the first bin to... The upper boundary of the 10th compartment is the maximum distance. This determined the distance boundaries of the 10 sub-boxes. Step 2.3: Filtering valid data within the distance segment

[0085] In each distance segment Within this section, site pairs where both sites are considered "valid sites" are selected, forming the valid analysis set for that segment. Ultimately, the total number of valid site pairs participating in the core indicator calculation is [number missing]. Yes, the overall effectiveness rate is approximately The distance range of each sub-container, the number of station pairs included, and the calculated center distance of each sub-container are shown in Table 1 below:

[0086] Table 1

[0087]

[0088] (III) Calculation of Core Heterogeneity Indicators for Distance Segments

[0089] Within each distance segment, station pairs where both stations are valid are selected as valid analysis station pairs. Based on the concentration difference data of these valid analysis station pairs, a heterogeneity index for that distance segment is calculated. The heterogeneity index includes a correlation index reflecting the correlation of concentration gradients and a normalized error index reflecting the error of concentration gradients. Specifically, this includes the following sub-steps:

[0090] Step 3.1: Calculation of correlation indicators

[0091] For the A distance segment, based on its right Data, calculations: (a) Rank correlation coefficient The assessment evaluates the consistency of the gridded air pollution concentration gradient with the station-level air pollution concentration gradient in terms of direction and order, and is insensitive to outliers. (b) Correlation coefficient To assess the strength of the linear correlation between the gridded air pollution concentration gradient and the site-specific air pollution concentration gradient. Step 3.2: Error Index Calculation

[0092] For the For each distance segment, calculate: (a) Root mean square error : (a) The absolute error between the gridded air pollution concentration gradient and the station air pollution concentration gradient. : This represents the variation in air pollution concentration at stations within this distance range. (c) Normalized root mean square error : This indicator is dimensionless and standardizes the error relative to the variation of air pollution concentration at the station, making the errors between different distance segments comparable. It can be directly interpreted as "the proportion of the gridded air pollution concentration gradient error to the station's air pollution concentration gradient".

[0093] (iv) Comprehensive evaluation based on logarithmic distance equal weighting and statistical transformation

[0094] Based on the distance span of each distance segment on a logarithmic scale, an integration weight is calculated for each distance segment. This weight is then used to weight and integrate the correlation and normalized error indices of all distance segments, resulting in comprehensive correlation and error indices for evaluating the spatial heterogeneity of gridded air pollution concentration data. The core innovation of this stage lies in how to scientifically integrate the data. The indicators for each distance segment are used to form a global evaluation conclusion. Specifically, this includes the following sub-steps:

[0095] Step 4.1: Calculation of logarithmic distance equal weights

[0096] For each distance segment Calculate a combined weight This weight does not depend on the sample size. It is based not on the distance segment's span on a logarithmic scale, but on the distance segment's length.

[0097]

[0098]

[0099] Logarithmic distance equal weighting assigns equal weights to equal-length intervals on the logarithmic distance axis, ensuring that each order of magnitude scale from 1 km to 1000 km (e.g., 1-10 km, 10-100 km, 100-1000 km) has similar weight in the final composite index. This allows for a fair and balanced evaluation of the data's performance across all spatial scales, overcoming the problem of long-distance segment dominance that may result from simple sample size weighting.

[0100] Step 4.2: Comprehensive Error Index calculate

[0101] It reflects the overall gradient error level of grid data across multiple scales.

[0102]

[0103] It is a dimensionless comprehensive index. The smaller its value, the more accurately the gridded air pollution describes the concentration gradient across all spatial scales, and the stronger its ability to reproduce spatial heterogeneity.

[0104] Step 4.3: Comprehensive correlation indicators calculate

[0105] This reflects the level of consistency between gridded air pollution concentrations and station-level air pollution concentration gradient changes across multiple scales. It employs a method based on... Robust methods for transformation:

[0106] (a) For each distance segment Correlation coefficient (or Correlation coefficient )conduct Transformation: .

[0107] (b) Using the log-distance equal weights calculated in step 4.1 ,calculate Weighted average: .

[0108] (c) Perform the reverse Transformation yields the comprehensive correlation coefficient: =0.267.

[0109] Step 4.4: Obtain core results for final evaluation

[0110] (a) Scale-based diagnostic report:

[0111] Includes the distance range for each distance segment, , , , ,like Figure 2 and Figure 3 As shown, it has advantages and disadvantages in identifying gridded air pollution concentrations at specific scales.

[0112] (b) Two core comprehensive indicators:

[0113] (Comprehensive gradient error): The theoretical range is [0, +∞), and in practice, the smaller the better. Different datasets or models can be directly compared. (Comprehensive Gradient Correlation): The value range is [-1, 1], the closer to 1 the better. Global Comprehensive Evaluation Index: Comprehensive Gradient Error Comprehensive gradient correlation This indicates the overall consistency between the gridded spatial variation pattern of air pollution and the changing trend of air pollution at the monitoring stations.

[0114] As can be seen from the above embodiments, this invention employs "distance binning," transforming the evaluation from "a single overall score" into "multiple distance segment reports," which accurately reveals the performance of data across different distances, such as short, medium, and long distances. Simultaneously, the proposed "logarithmic equal weighting" algorithm creatively uses "logarithmic distance span" to calculate weights when synthesizing results from each distance segment, ensuring that each distance scale (e.g., 1-10 km, 10-100 km) has equal weight in the final evaluation, resulting in a fairer outcome.

[0115] While the present invention has been described above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.

Claims

1. A method for evaluating the spatial heterogeneity of gridded air pollution concentration data, characterized in that, Includes the following steps: Step 1: Obtain gridded air pollution concentration data and ground station air pollution concentration data within the same geographical area, construct all station pairs, and calculate the concentration difference and geographical distance for each station pair; Step 2: Divide all station pairs into K consecutive distance segments by equal number of bins according to their geographical distance; Step 3: Within each distance segment, select station pairs where both stations are valid as valid analysis station pairs. Based on the concentration difference data of the valid analysis station pairs, calculate the heterogeneity index for that distance segment; the heterogeneity index includes a correlation index reflecting the correlation of concentration gradients and a normalized error index reflecting the error of concentration gradients; Step 4: Based on the distance span of each distance segment on a logarithmic scale, calculate the integration weight for each distance segment, and use the weight to perform weighted integration of the correlation index and normalized error index for all distance segments to obtain a comprehensive correlation index and comprehensive error index for evaluating the spatial heterogeneity of gridded air pollution concentration data.

2. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 1, characterized in that, Step 1 includes the following sub-steps: Step 1.1: Obtain gridded air pollution concentration data and ground station air pollution concentration data within the same geographical area and time period, and align the grid data with the geographical locations of each station to form a data pair sequence at the station location; Step 1.2: Perform validity screening on the ground station air pollution concentration data by time scale, and mark the stations that meet the preset validity threshold as valid stations; Step 1.3: For the selected time scale, calculate the average ground station concentration and the average gridded concentration for each station within that time period; Step 1.4: Combine all stations in pairs to construct unique station pairs. And calculate the ground station concentration difference for each station pair. Gridded concentration difference and the geographical distance between stations .

3. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 2, characterized in that, In the validity screening in step 1.2, the time scale includes an annual scale, a monthly scale, or a daily scale; the preset validity threshold is set according to the time scale, which is ≥75% for an annual scale, and ≥18 days with a valid percentage of valid days or ≥65% for a monthly scale, or ≥20 days with a valid percentage of valid days or ≥75%.

4. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 1, characterized in that, Step 2 includes the following sub-steps: Step 2.1: Set the minimum effective analysis distance km, geographical distance Step 2.2: Sort the remaining site pairs in ascending order of their geographical distance, and divide them into K distance segments using the equal number binning method, so that each distance segment contains an equal number of site pairs; Step 2.3: Within each distance segment k, select site pairs where both sites are valid sites to form the set of valid analysis site pairs for that distance segment.

5. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 4, characterized in that, In step 2.2, the lower boundary of the first bin is forcibly set to the minimum effective analysis distance. .

6. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 1, characterized in that, In step 3, the correlation indicators include Rank correlation coefficient and / or Correlation coefficient The normalized error index is the normalized root mean square error. It can be calculated using the following formula: , in, The root mean square error of the difference between the effective analysis station pairs within the k-th distance segment is given. denoted as the standard deviation of the ground station concentration difference for the effective analysis station pair within the k-th distance segment.

7. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 6, characterized in that, In step 4, the integration weight is calculated based on the distance span of each distance segment on a logarithmic scale. Specifically: , in, , and These are the upper and lower limits of the distance for the k-th distance segment, respectively.

8. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 7, characterized in that, In step 4, the comprehensive error index is calculated as follows: 。 9. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to claim 7, characterized in that, In step 4, comprehensive correlation indicators are used. By based on The weighted average method for the transformation is calculated as follows: The correlation index for each distance segment is... Transformation: , calculate Weighted average: , right Performing an inverse transformation yields the comprehensive correlation index: .

10. The method for evaluating the spatial heterogeneity of gridded air pollution concentration data according to any one of claims 1-9, characterized in that, Step 4 is followed by: outputting the distance range, number of valid site pairs, correlation index, and normalization error index for each distance segment, in order to identify the performance of gridded air pollution concentration data at different spatial scales.