Air quality monitoring station site selection optimization method, medium and equipment

Through the random forest model correction of remote sensing data and a multi-dimensional evaluation system, the location selection of air quality monitoring sites is optimized, which solves the problem of insufficient multi-dimensional comprehensive evaluation of air quality monitoring networks in the existing technology, and realizes the scientific nature of the layout of air quality monitoring sites and the overall consideration of health protection, providing accurate data support.

CN120494188APending Publication Date: 2025-08-15CENT SOUTH UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510624139.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing air quality monitoring network is difficult to fully characterize the spatial heterogeneity of regional pollution, lacks multi-dimensional comprehensive evaluation, and fails to effectively consider future urban development and changes and people's health protection needs, resulting in insufficient scientificity and reliability of the optimization results of monitoring site.

Method used

The random forest model is used to correct remote sensing monitoring data, combine consistency measurement indicators and multi-dimensional spatial statistical indicator system to optimize the site selection of monitoring sites, consider representativeness, redundancy and rationality, and combine crowd exposure risks and future development trends to optimize site layout.

Benefits of technology

It has achieved scientific and forward-looking improvement in the layout of air quality monitoring sites, improved the coverage capacity and health protection effect of the monitoring network, provided accurate data support, and provided comprehensive support for environmental management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494188A_ABST
    Figure CN120494188A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of air quality monitoring, in particular to an air quality monitoring station site selection optimization method, medium and equipment, and the method comprises the steps: S1, obtaining remote sensing monitoring data and ground station monitoring data of pollutant concentration of a certain research area; s2, correcting the remote sensing monitoring data based on a random forest model to obtain corrected monitoring data; s3, evaluating the credibility of the corrected monitoring data based on the consistency measurement index, and if the consistency measurement index of the corrected monitoring data is higher than the consistency measurement index of the remote sensing monitoring data, entering S4; otherwise, returning to S2; and S4, according to the corrected monitoring data, optimizing the monitoring station at the pollution view angle, and obtaining a monitoring station site selection result after the pollution view angle optimization. According to the invention, overall consideration of monitoring efficiency and health protection is realized, future development change is considered, and more accurate and comprehensive data support can be provided for air quality monitoring and environment management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air quality monitoring, and in particular to an air quality monitoring site selection optimization method, medium and equipment. Background Art

[0002] Air pollution has become a major environmental issue of global concern. To accurately monitor air quality and assess the spatiotemporal distribution of pollutants, existing air quality monitoring networks of varying sizes have been established. However, these existing networks suffer from the following shortcomings:

[0003] ① Existing monitoring site evaluations often rely on discrete ground-based monitoring data or pollution source distribution data, making it difficult to fully characterize the spatial heterogeneity of regional pollution. Although remote sensing data offers the advantages of macroscopic and continuous monitoring, the local credibility assessment of remote sensing data needs to be improved, resulting in limited application in monitoring site evaluation and optimization. ② Traditional site evaluation methods often overly focus on single indicators (such as coverage and site density), lacking comprehensive consideration of multiple dimensions such as monitoring capacity and pollution distribution characteristics. Existing research lacks monitoring network optimization schemes based on multidimensional collaborative evaluation, limiting the scientific nature and reliability of the optimization results. ③ Existing site optimization models primarily optimize based on historical monitoring data, with limited consideration of future changes in pollution patterns, population distribution, and building density brought about by urban development. Furthermore, insufficient attention is paid to the health protection needs of the population, making it difficult to achieve a comprehensive balance between environmental monitoring and health protection. Summary of the Invention

[0004] The present invention aims to provide a method for optimizing the site selection of air quality monitoring stations that can take into account both environmental monitoring and health protection. The specific technical solution is as follows:

[0005] The present invention provides a method for optimizing the site selection of air quality monitoring stations, comprising the following steps:

[0006] S1: Obtain remote sensing monitoring data and ground station monitoring data of pollutant concentrations in a certain study area;

[0007] S2: Correct the remote sensing monitoring data based on the random forest model to obtain the corrected monitoring data;

[0008] S3: Evaluate the credibility of the corrected monitoring data based on the consistency measurement index. If the consistency measurement index of the corrected monitoring data is higher than the consistency measurement index of the remote sensing monitoring data, proceed to S4; otherwise, return to S2.

[0009] S4: Optimize the monitoring sites from a pollution perspective based on the corrected monitoring data to obtain the site selection results of the monitoring sites after the pollution perspective optimization.

[0010] Optionally, the S2 includes:

[0011] S2.1. Use natural and human environmental data as modeling variables and perform preliminary screening of these variables using significance analysis and collinearity tests to obtain a selected subset of modeling features. Preliminary screening involves excluding modeling variables with a P-value > 0.05 and a VIF > 10. Natural environmental data include DEM data, temperature, wind direction, humidity, and wind speed; and human environmental data include land use, population density, point of interest (POI) data, road traffic data, and emission inventories.

[0012] S2.2. Use all features in the current modeling feature subset to train the pre-trained model;

[0013] S2.3. Use recursive feature elimination to eliminate unimportant features in the modeling feature subset;

[0014] S2.4. Use cross-validation to calculate the model error of the current modeling feature subset;

[0015] S2.5. Repeat S2.2 to S2.4 until the model error reaches the minimum value, and a random forest model is obtained;

[0016] S2.6. Input the remote sensing monitoring data into the random forest model, and output the corrected monitoring data;

[0017] In S3, the consistency measurement indicators include goodness of fit, root mean square error, high deviation rate, uncertainty and spatial differentiation characteristics.

[0018] Optionally, the S4 includes:

[0019] S4.1. Eliminate sites with insufficient representation and sites with high redundancy and low coverage. Retain sites with a representation radius greater than 4000m, even if the site has a high redundancy count. A site with insufficient representation is defined as one with a spatial representation range less than or equal to the first radius threshold. A site with high redundancy and low coverage is defined as one with a redundancy count greater than the median and a spatial representation range radius less than the second radius threshold. The redundancy count refers to the number of times the comprehensive redundancy evaluation index of the site that overlaps with other monitoring ranges is greater than the median.

[0020] S4.2. Divide the study area into multiple grids and calculate the rationality value of each grid based on a multi-dimensional spatial statistical indicator system;

[0021] S4.3. Sort the grids from high to low rationality and prioritize the site placement in the study area based on the rationality values. Specifically, the top 30% of the rationality values are assigned high priority; the middle 40% of the rationality values are assigned medium priority; and the bottom 30% of the rationality values are assigned low priority.

[0022] S4.4. A greedy algorithm based on multi-criteria constraints is used to select new sites according to the site layout priority, and the site selection results of the monitoring sites are obtained after optimization from the pollution perspective.

[0023] Optionally, the spatial representative range in S4.1 is obtained by using buffer zone analysis and goodness of fit R 2 Calculate the spatial representativeness of each site in the study area, including:

[0024] With each monitoring site i as the center, buffer zones with radii of 100m, 200m, 300m, 400m, 500m, 600m, 700m, 800m, 900m, 1000m, 1500m, 2000m, 2500m, 3000m, 3500m, and 4000m are generated, denoted as B. s (r), where r is the buffer radius;

[0025] For each buffer B i (r), extract the remote sensing grid data within its range and calculate the mean of the remote sensing grid data

[0026] Calculate the pollutant concentration monitoring value S at the site i and buffer mean Goodness of fit between the mean

[0027] Goodness of fit based on the mean Determine the spatial representativeness of the monitoring sites, specifically:

[0028] For each monitoring site i, traverse all buffer zones with radius r and select The minimum radius r min As the spatial representative range of the site, the specific formula is as follows:

[0029] r min =min{r|R 2 (r)<0.9};

[0030] If there is no radius r that satisfies R 2 (r) ≥ 0.9, then the site space representative range is 0. If all radii r satisfy R 2 If (r) ≥ 0.9, the spatial representative range of the site is marked as a maximum value of 4000 m;

[0031] In S4.1, the method for obtaining the comprehensive evaluation index of site redundancy is to calculate the comprehensive evaluation index of site redundancy of sites within the overlapping area based on the similarity of geographical environment and the similarity of ground site monitoring data, specifically:

[0032] ① Calculate geographical environment similarity, including:

[0033] Select geographical elements and divide them into five levels using the natural breakpoint method;

[0034] The explanatory power of discretized geographical elements on geographical phenomena is quantified by calculating the q value. The specific formula is as follows:

[0035]

[0036] Among them, N h and are the number of samples and variance of the h-th layer, N and σ respectively 2 are the total number of samples and the overall variance, respectively, and L is the number of strata;

[0037] Select geographic features with q values greater than the set threshold or ranked in the top X as input variables for similarity measurement;

[0038] Calculate the multidimensional Euclidean distance D between sites based on the input variables of the similarity metric ik , the specific formula is as follows:

[0039]

[0040] Where: i and k are site numbers, i≠k; X ij is the jth geographic feature value of the i-th site, X kj is the jth geographic element value of the kth site; Z(.) is the normalization process; m is the number of geographic elements;

[0041] Normalize the multidimensional Euclidean distance to the range [0,1] to obtain the geographical similarity

[0042]

[0043] Among them, max(D) is the maximum value of the Euclidean distance of all site pairs;

[0044] ②Calculate the similarity of monitoring data, including:

[0045] Calculate the Pearson correlation coefficient r between site i and site k ik :

[0046]

[0047] in, and are the mean pollutant concentrations at site i and site k, S i (t) and S k(t) are the mean pollutant concentrations at site i and site k on day t respectively; T is the number of days with statistically valid values of ground site monitoring data of pollutant concentration;

[0048] Normalize the correlation coefficient to the range of [0,1] to obtain the similarity of monitoring data The specific formula is as follows:

[0049]

[0050] ③Calculate the overlapping area of the monitoring site pairs with overlapping monitoring ranges, specifically:

[0051] A double circle geometric model is established for site i and site k, with radii r i and r k ;

[0052] Calculate the distance d between site i and site k ik , when d ik ≥r i +r k When d ik ≤|r i -r k |, it is determined that site i and site k completely overlap; in other cases, site i and site k partially overlap;

[0053] Calculate the overlapping area of site i and site k;

[0054] ④ Based on geographical similarity Similarity with monitoring data Calculate the comprehensive evaluation index of the redundancy degree of sites within the overlapping area. The specific formula is as follows:

[0055]

[0056] Among them, ω is the geographical environment similarity weight.

[0057] Optionally, the S4.2 includes:

[0058] S4.2.1. Divide the study area into multiple grids;

[0059] S4.2.2. Calculate the information entropy of each indicator in the multidimensional spatial statistical indicator system within each grid, where the multidimensional spatial statistical indicator system includes the annual average concentration of pollutants Temporal variability Spatial variability Annual average concentration exceeding the standard E annual (g) and the annual daily average exceeding frequency F daily (g), the specific formula is as follows:

[0060]

[0061] Where: g is the grid number, f is the index number, and N is the number of grids;

[0062] S4.2.3. Calculate the weight of each indicator ω based on information entropy f , the specific formula is as follows:

[0063]

[0064] Where: F is the number of indicators;

[0065] S4.2.4. Calculate the rationality value Q of each grid based on the weight of the indicator g , the specific calculation formula is as follows:

[0066]

[0067] Among them: F daily (g) is the annual daily average exceeding the standard frequency, ω1 is The weight of ω2 is The weight of ω3 is The weight of E annual The weight of (g), ω5 is F daily The weight of (g).

[0068] Optionally, the S4.4 includes:

[0069] S4.4.1. Add the number of new sites based on the grid's site layout priorities, specifically:

[0070] N high =0.5N new ;

[0071] N mid =0.3N new ;

[0072] N low =0.2N new ;

[0073] Where: N high is the number of new sites in the high-priority grid, N mid N is the number of new sites in the medium priority grid. low N is the number of new sites in the low-priority grid; new Set the number of sites to a value between 90% and 110% of the original number of sites;

[0074] S4.4.2. Use a greedy algorithm to find the optimal address, specifically:

[0075] Arrange the grids in descending order of rationality value, put the grids with covered points into the existing site set, and put the grids with uncovered points into the candidate set;

[0076] Iteratively select the grid with the largest rationality value that meets the distance constraint in the candidate set as the new site to be added to the existing site set; and dynamically update the existing site set and candidate set in the study area, where: the distance constraint is that the Haversine distance between any sites is greater than or equal to 1000m.

[0077] Optionally, the method further includes optimizing the monitoring sites from a risk perspective based on the site selection results of the monitoring sites optimized from a pollution perspective, specifically:

[0078] S5.1. Calculate the comprehensive exposure risk index (ERI) for each grid, specifically:

[0079] S5.1.1. Construct four exposure risk indicators based on the distribution of patients with cardiovascular and cerebrovascular diseases, allergic diseases, and respiratory diseases, as well as population density and pollutant distribution data;

[0080] S5.1.2. Calculate the exposure risk index corresponding to the exposure risk indicator. The specific calculation formula is as follows:

[0081]

[0082] Where: b is the number of the exposure risk indicator, T gb is the exposure risk index of the b-th indicator in the g-th grid in the study area, P gb is the population of the b-th category indicator in the g-th grid, C g is the pollution concentration value of the g-th grid;

[0083] S5.1.3. Standardize the exposure risk index and perform equal weight fusion on the standardized exposure risk to obtain the comprehensive exposure risk index ERI. g , the specific calculation formula is as follows:

[0084]

[0085] Where: T gb * is the standardized exposure risk index;

[0086] S5.2. Constructing a comprehensive site optimization index Z based on the risk perspective based on the comprehensive exposure risk index g , the specific formula is as follows:

[0087] Z g =0.4ERI g +0.2R p +0.2R b+0.2R pop ;

[0088] Where: R p is the pollutant concentration prediction, R b For building density prediction, R pop For population density prediction;

[0089] S5.3. Taking district or county as the basic spatial unit, sort the monitoring stations in each unit by comprehensive index, and remove the Z g <Z threshold 's site;

[0090] S5.4. Based on the greedy algorithm, the monitoring site selection results after pollution perspective optimization are further optimized to obtain the comprehensive point optimization results.

[0091] Optionally, the 5.4 includes:

[0092] S5.4.1. Select all the items that satisfy Z g ≥Z threshold The grid is used as the candidate site selection point;

[0093] S5.4.2. Count the number of sites in each unit. If the number of sites in each unit does not exceed Q, the candidate points in each unit are sorted according to Z. g Arrange in descending order and iteratively select the site optimization comprehensive index Z that meets the spatial constraints in each unit g The largest candidate point is selected as the new site until the number of sites in each cell exceeds Q, where Q is a set value and the spatial constraint is that the distance between the new site and the existing site is no less than 1.5 km;

[0094] S5.4.3. Select sites based on the entire study area and iteratively select sites within the study area that meet the spatial constraints and optimize the comprehensive index Z. g The largest candidate point is selected as the new site; the termination condition is that the total number of sites in the study area reaches the upper limit of the total number of sites or the total number of sites reaches the lower limit of the total number of sites, and the grids with the top 50% of the site optimization comprehensive index in the study area have sites deployed; finally, the comprehensive point optimization result taking into account pollution and risk is obtained; where: the total number of sites N final ∈[150,183].

[0095] The present invention also provides a readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implements the above-mentioned method for optimizing the site selection of air quality monitoring stations.

[0096] The present invention also provides an electronic device, characterized in that it includes: at least one processor, at least one memory and computer program instructions stored in the memory, when the computer program instructions are executed by the processor, the air quality monitoring station site selection optimization method as described above is performed.

[0097] The technical solution of the present invention analyzes the consistency between the remote sensing data of pollutant concentration and the monitoring data of the ground station, deeply explores the spatial differentiation characteristics of the credibility of the remote sensing data and its changing laws, and corrects the remote sensing data, providing reliable data support for the optimization of the air quality monitoring station points; and establishes a multi-dimensional evaluation system including representativeness, redundancy and rationality based on the pollution perspective, and optimizes the site layout, breaking through the limitations of traditional single-dimensional optimization, and systematically evaluating the spatial coverage capacity, redundancy characteristics and layout rationality of the monitoring station; on this basis, the population exposure risk and future development trends are taken into consideration, while ensuring the pollution monitoring capability, further strengthening the coverage of high-risk areas for population exposure, achieving a comprehensive balance between monitoring efficiency and health protection, and taking into account future development changes, which can provide more accurate and comprehensive data support for air quality monitoring and environmental management. The present invention provides a feasible technical path for improving the scientificity, foresight and service efficiency of the site selection of air quality monitoring stations.

[0098] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0100] Figure 1 It is a technical roadmap of the air quality monitoring station site selection optimization method in an embodiment of the present invention. DETAILED DESCRIPTION

[0101] To make the objectives, features, and advantages of the present invention more readily apparent, the following detailed description of the present invention is provided with reference to the accompanying drawings. The accompanying drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein.

[0102] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways as defined and covered by the claims.

[0103] In one embodiment, a method for optimizing the site selection of air quality monitoring stations, focusing on the impact of pollution, includes the following steps:

[0104] S1: Obtain remote sensing monitoring data and ground station monitoring data of pollutant concentrations in a certain study area;

[0105] The pollutants in this example are PM 2.5 The remote sensing monitoring data selected are the CHAP dataset with relatively high precision and spatiotemporal resolution. The ground monitoring data mainly include 10 national monitoring stations from the China National Environmental Monitoring Center and 157 encrypted stations from the Changsha Ecological Environment Monitoring Center in Hunan Province from January 1, 2019 to December 31, 2022. 2.5 Concentration data.

[0106] S2: Correct the remote sensing monitoring data based on the random forest model to obtain the corrected monitoring data;

[0107] The S2 includes:

[0108] S2.1. Use natural and human environmental data as modeling variables and perform preliminary screening using significance analysis and collinearity tests to obtain a selected modeling feature subset. Preliminary screening involves removing modeling variables with a significance P-value > 0.05 and a variance inflation factor (VIF) > 10. Natural environmental data include DEM data, temperature, wind direction, humidity, and wind speed; human environmental data include land use, population density, point of interest (POI) data, road traffic data, and emission inventories. Preliminary screening reduces unnecessary noise, avoids interference between variables, and improves model accuracy.

[0109] Random forest is an ensemble learning method based on decision trees. It improves the generalization ability and accuracy of the model by constructing multiple decision trees and combining the prediction results of each tree.

[0110] In random forests, feature importance reflects the contribution of each feature in reducing model error. This is achieved by evaluating the impact of each feature on the decision tree node split. For each feature, feature importance is the sum of the mean squared error (MSE) reduced by that feature in the split process of all trees.

[0111] To further improve model accuracy, recursive feature elimination with cross-validation (RFECV) is used for feature selection. Recursive Feature Elimination (RFE) repeatedly builds a model to calculate the importance of each feature and gradually eliminates the least important features until the optimal feature subset is selected. After each feature is eliminated, cross-validation (CV) is used to calculate the model error of the current feature subset. By calculating the cross-validation error of each feature subset, the feature combination with the smallest error is selected. The specific modeling process is shown in S2.2 to S2.6:

[0112] S2.2. Use all features in the current modeling feature subset to train the pre-trained model;

[0113] S2.3. Use recursive feature elimination to eliminate unimportant features in the modeling feature subset;

[0114] S2.4. Use cross-validation to calculate the model error of the current modeling feature subset;

[0115] S2.5. Repeat S2.2 to S2.4 until the model error reaches the minimum value, and a random forest model is obtained;

[0116] S2.6. Input the remote sensing monitoring data into the random forest model, and output the corrected monitoring data;

[0117] S3: Evaluate the credibility of the corrected monitoring data based on the consistency measurement index. If the consistency measurement index of the corrected monitoring data is higher than the consistency measurement index of the remote sensing monitoring data, proceed to S4; otherwise, return to S2.

[0118] In S3, the consistency measurement indicators include goodness of fit, root mean square error, high deviation rate, uncertainty and spatial differentiation characteristics; specifically:

[0119] The consistency between remote sensing monitoring data and site monitoring data is analyzed from the two aspects of fitting effect and deviation degree, and the credibility of remote sensing monitoring data is evaluated, thereby quantifying the suitability of remote sensing methods for local monitoring. The fitting effect is evaluated by goodness of fit and root mean square error, and the deviation degree is characterized by high deviation rate and uncertainty. Among them:

[0120] (1) Goodness of fit

[0121] The goodness of fit R, which characterizes the fitting effect between remote sensing monitoring data and site monitoring data 2 The calculation formula is as follows:

[0122]

[0123] Where S i For site monitoring data, is the mean of the site monitoring data, R i Remote sensing monitoring data.

[0124] (2) Root mean square error

[0125] RMSE is used to measure the fitting effect of remote sensing monitoring data and site monitoring data. The calculation formula is as follows:

[0126]

[0127] Where n is the number of samples, S i is the site monitoring data, R i Remote sensing monitoring data.

[0128] (3) High deviation rate

[0129] Extract the remote sensing data values at the monitoring station, calculate the deviation between the two, sort all deviation values in descending order, and take the value in the top 2% (27.3) as the threshold. Deviations greater than this threshold are considered high deviations. The high deviation rate refers to the ratio of days within the statistical period D where the deviation between the remote sensing monitoring data and the station monitoring data is greater than 27.3. The calculation formula is as follows:

[0130]

[0131] Where, HD i is the high deviation rate at site i during the statistical period, The number of days when the deviation between remote sensing monitoring data and site monitoring data is greater than 27.3.

[0132] (4) Uncertainty

[0133] The uncertainty calculation formula that characterizes the overall deviation level between remote sensing monitoring data and site monitoring data is as follows:

[0134]

[0135] Where U i,d is the uncertainty at site i on day d, R i,d is the daily average value of remote sensing monitoring data at site i on day d, S i,d is the daily average of the monitoring data at site i on day d.

[0136] Uncertainty U at site i i The total number of days D in the statistical period for all U i,d The average value is calculated as follows:

[0137]

[0138] (5) Credibility spatial differentiation feature identification

[0139] To reveal PM 2.5 The spatial differentiation characteristics of the credibility of concentration remote sensing data were analyzed. According to the direction angle and wind direction of the national monitoring station and the infilled station, the data of the infilled station when it was in the downwind state relative to the national monitoring station were selected. According to the distance between the infilled station and the nearest national monitoring station, the data were divided into 6 groups: less than or equal to 1km, 1-5km, 5-10km, 10-15km, 15-20km, and greater than 20km. 2.5 The consistency analysis of concentration remote sensing data and ground monitoring data of each group of encrypted stations was carried out.

[0140] ① Closest distance matching

[0141] By calculating the Haversine distance between each encrypted site and the national monitoring site, the nearest national monitoring site to each encrypted site is determined. The distance calculation formula is as follows:

[0142]

[0143] Where R is the radius of the Earth, lat1 and lat2 are the latitudes of the two points, and Δlat and Δlon are the latitude and longitude differences between the two points, respectively.

[0144] ② Direction angle calculation

[0145] Use the following formula to calculate the direction angle between the encrypted site and the nearest national control site:

[0146] θ=arctan2(Δx,Δy);

[0147] in:

[0148] Δx=sin(Δlon)·cos(lat2);

[0149] Δy=cos(lat1)·sin(lat2)-sin(lat1)·cos(lat2)·cos(Δlon);

[0150] ③Determination of tailwind and headwind

[0151] Based on the azimuth angle and wind direction between the national monitoring station and the intensified station, determine whether the intensified station is in a tailwind or headwind relative to the national monitoring station. The relationship between wind direction and azimuth angle is determined using the following rule: a tailwind is determined when the difference between wind direction and azimuth angle is less than 90° or greater than 270°; a headwind is determined when the difference between 90° and 270°. Wind direction angles are calculated using the following rule: 0° is for east wind, and other wind directions increase in a clockwise direction.

[0152] S4: Optimize the monitoring sites from a pollution perspective based on the corrected monitoring data to obtain the site selection results of the monitoring sites after the pollution perspective optimization;

[0153] The S4 includes:

[0154] S4.1. Eliminate sites with insufficient representative range and high-redundancy low-coverage sites, and retain sites with a representative range radius greater than 4000m, even if the site has a high redundancy number. Among them: sites with insufficient representative range are sites with a spatial representative range less than or equal to the first radius threshold; sites with high redundancy low coverage are sites with a redundancy number greater than the median and a spatial representative range radius less than the second radius threshold. The redundancy number refers to the number of times the value of the comprehensive evaluation index of the redundancy degree of sites that overlap with other monitoring ranges is greater than the median; in this embodiment, the first radius threshold is 500m and the second radius threshold is 2000m.

[0155] In S4.1, the spatial representative range is obtained by using buffer analysis and goodness of fit R 2 Calculate the spatial representativeness of each site in the study area, including:

[0156] With each monitoring site i as the center, buffer zones with radii of 100m, 200m, 300m, 400m, 500m, 600m, 700m, 800m, 900m, 1000m, 1500m, 2000m, 2500m, 3000m, 3500m, and 4000m are generated, denoted as B. i (r), where r is the buffer radius;

[0157] For each buffer B i (r), extract the remote sensing grid data within its range and calculate the mean of the remote sensing grid data The specific formula is as follows:

[0158]

[0159] Among them, R g Buffer B i PM of the g-th grid in (r) 2.5 Concentration value, N r Buffer B i (r) The number of grid cells within;

[0160] Calculate the pollutant concentration monitoring value S at the site i and buffer mean Goodness of fit between the mean The specific formula is as follows:

[0161]

[0162] Where: S i PM for each site 2.5 Monitoring value, PM for all sites 2.5 Mean of monitored values; is the mean value of remote sensing data within the buffer zone radius r corresponding to site i, and I is the total number of sites.

[0163] Goodness of fit based on the mean Determine the spatial representativeness of the monitoring sites, specifically:

[0164] For each monitoring site s, traverse all buffer zones with radius r and select The minimum radius r min As the spatial representative range of the site, the specific formula is as follows:

[0165] r min =min{r|R 2 (r)<0.9};

[0166] If there is no radius r that satisfies R 2 (r) ≥ 0.9, then the site space representative range is 0. If all radii r satisfy R 2 If (r) ≥ 0.9, the spatial representative range of the site is marked as a maximum value of 4000 m;

[0167] In S4.1, the method for obtaining the comprehensive evaluation index of site redundancy is to calculate the comprehensive evaluation index of site redundancy of sites within the overlapping area based on the similarity of geographical environment and the similarity of ground site monitoring data, specifically:

[0168] ① Calculate geographical environment similarity, including:

[0169] Select geographical elements and divide them into 5 levels using the natural breakpoint method. In this embodiment, 87 geographical elements such as annual average temperature, annual average humidity, and the number of POI types in different distance buffer zones are selected for screening. See Table 1 for details.

[0170] Table 1 Redundancy evaluation index table

[0171]

[0172] The explanatory power of discretized geographical elements on geographical phenomena is quantified by calculating the q value. The specific formula is as follows:

[0173]

[0174] Among them, N h and are the number of samples and variance of the h-th layer, N and σ respectively 2 are the total number of samples and the overall variance, respectively, and L is the number of strata;

[0175] Select geographic features with q values greater than the set threshold or ranked in the top X as input variables for similarity measurement;

[0176] Calculate the multidimensional Euclidean distance D between sites based on the input variables of the similarity metric ik , the specific formula is as follows:

[0177]

[0178] Where: i and k are site numbers, i≠k; X ij is the jth geographic feature value of the i-th site, X kj is the jth geographic element value of the kth site; Z(.) is the normalization process; m is the number of geographic elements;

[0179] Normalize the multidimensional Euclidean distance to the range [0,1] to obtain the geographical similarity

[0180]

[0181] Among them, max(D) is the maximum value of the Euclidean distance of all site pairs;

[0182] In this embodiment, the geographical detector is used to detect the influence of PM 2.5 The geographical environmental factors of pollution distribution are screened, and the geographical factors with significant impact are selected as input variables for subsequent similarity measurement.

[0183] ②Calculate the similarity of monitoring data, including:

[0184] Calculate the Pearson correlation coefficient r between site i and site k ik :

[0185]

[0186] in, and are the mean pollutant concentrations at site i and site k, S i (t) and S k (t) are the mean pollutant concentrations at site i and site k on day t respectively; T is the number of days with statistically valid values of ground site monitoring data of pollutant concentration;

[0187] Normalize the correlation coefficient to the range of [0,1] to obtain the similarity of monitoring data The specific formula is as follows:

[0188]

[0189] ③Calculate the overlapping area of the monitoring site pairs with overlapping monitoring ranges, specifically:

[0190] A double circle geometric model is established for site i and site k, with radii r i and r k ;

[0191] Calculate the distance d between site i and site k ik , when d ik ≥r i +r k When d ik ≤|r i -r k |, it is determined that site i and site k completely overlap; in other cases, site i and site k partially overlap;

[0192] Calculate the overlapping area A between site i and site k overlap :

[0193] A overlap =A1+A2-A3;

[0194] in:

[0195]

[0196] ④ Based on geographical similarity Similarity with monitoring data Calculate the comprehensive evaluation index of the redundancy degree of sites within the overlapping area. The specific formula is as follows:

[0197]

[0198] Wherein, ω is the geographical environment similarity weight. In this embodiment, ω=0.5.

[0199] S4.2. Divide the study area into multiple grids and calculate the rationality value of each grid based on a multi-dimensional spatial statistical indicator system;

[0200] The S4.2 includes:

[0201] S4.2.1. Divide the study area into multiple grids;

[0202] S4.2.2. Calculate the information entropy of each indicator in the multidimensional spatial statistical indicator system within each grid, where the multidimensional spatial statistical indicator system includes the annual average concentration of pollutants Temporal variability Spatial variability Annual average concentration exceeding the standard Eannual (g) and the annual daily average exceeding frequency F daily (g), the specific formula is as follows:

[0203] (1)PM 2.5 Average annual concentration

[0204] The scientific and reasonable layout of air quality monitoring stations should ensure that there are sufficient monitoring points covering each concentration gradient area, and PM 2.5 The annual average concentration is an indicator of the long-term distribution of pollution and can smooth out the impact of factors such as short-term weather changes and instantaneous pollution emissions. Compared with single-day or short-term concentration values, PM 2.5 The annual average concentration can better represent the overall pollution status of a region, help identify pollution hotspots, background concentration areas and transition areas, intuitively display the spatial gradient distribution of pollution levels, and reduce monitoring blind spots. The calculation formula is as follows:

[0205]

[0206] Where D g is the effective observation days of grid g, is the validity indicator function (valid value is 1, otherwise it is 0).

[0207] (2) Temporal variability

[0208] Temporal variability measures the PM 2.5 The degree of fluctuation of concentration in the time dimension, compared with the annual average concentration, can identify the frequent occurrence areas of short-term pollution events through time variability, reflecting the PM 2.5 The dynamic characteristics of pollution can be monitored to ensure that monitoring stations can effectively capture areas with large pollution fluctuations and improve the monitoring network's ability to respond to pollution events. In addition, high-variability areas are usually affected by complex meteorological conditions or unstable pollution sources, such as industrial emissions, traffic emissions, or seasonal biomass burning. Station layout needs to give priority to covering such areas to improve pollution source tracing capabilities and provide real-time data support for emergency management. In this study, the annual variance at each grid location is used to measure the PM 2.5 The concentration time variability is as follows:

[0209]

[0210] (3) Spatial variability

[0211] To further analyze PM 2.5Spatial variability of concentration: calculate the annual mean value of the variance of pollution concentration between each grid point and its adjacent grids. Select the 8-neighborhood grid of grid g (upper, lower, left, right and four diagonal directions). For the grid at the boundary, if the number of neighbors is less than 8, all available adjacent grids are used for calculation. Spatial variability reflects the PM 2.5 The uneven distribution of pollution in geographical space, that is, the variation of pollution levels between adjacent areas. Temporal variability focuses on the dynamic changes of pollution concentration over time and is suitable for evaluating the long-term stability of a site, while spatial variability focuses on the trend of pollution changes at different spatial locations. It can reveal the spatial heterogeneity of pollution levels in a region, and thus effectively guide the optimal layout of air quality monitoring sites and improve the representativeness and monitoring efficiency of the entire monitoring network. In areas with large spatial variability, PM 2.5 Local variations in pollution concentrations are more dramatic, requiring a higher density of monitoring stations to capture the details of pollution changes.

[0212] For each grid, the daily average PM 2.5 The local spatial variance of concentration is calculated using the following formula:

[0213]

[0214] in is the valid neighborhood set (Moore neighborhood, containing 3-8 cells) of grid g at time t, and v is the grid number in the valid neighborhood set.

[0215] The formula for calculating the annual average is:

[0216]

[0217] (4) Annual average concentration exceeding the standard E annual

[0218] According to relevant regulations, PM 2.5 The annual average concentration limit is 25 μg / m 3 Using this as the limit, the annual average concentration of each grid is judged to determine whether it exceeds the standard, and the areas where the annual average concentration exceeds the standard are identified.

[0219]

[0220] Where: E annual (g,f) is the annual average concentration exceeding the standard for index f in grid g;

[0221] (5) Annual daily average exceeding standard frequency F daily

[0222] According to the Ambient Air Quality Standard (GB3095-2012), PM 2.5 The daily average concentration limit is 35 μg / m3 (Level 1 standard) and 75 μg / m 3 (Secondary standard), calculate the number of days in which the standard is exceeded in each grid in 2022. Compared with short-term pollution fluctuations, the annual daily average frequency of exceeding the standard can more stably reflect the long-term exposure level of regional pollution. From the perspective of spatial heterogeneity, the spatial distribution of the frequency of exceeding the standard can reveal the aggregation characteristics of pollution sources and the law of pollution diffusion. Areas with high frequency of exceeding the standard usually correspond to areas with long-term accumulation and high emissions of pollutants, and require a higher density of stations for monitoring to accurately capture the pollution level and its changing trend, which will help to trace the source of pollution and control it. In areas with low frequency of exceeding the standard, the number of monitoring points can be appropriately reduced to optimize resource allocation. In addition, the frequency of exceeding the standard under different concentration thresholds can reflect the spatial distribution characteristics of different pollution levels, which helps to distinguish the pollution monitoring priorities of the region and more comprehensively evaluate the rationality of site coverage.

[0223]

[0224] Among them: F daily (g,f) The frequency of exceeding the standard for the annual daily mean value of grid g index f;

[0225] (6) Construction of comprehensive indicators based on entropy method

[0226] To scientifically and rationally quantify PM 2.5 To investigate the spatial variability of pollution and assess the rationality of monitoring site layout, this study used the entropy method to construct a comprehensive indicator system. The entropy method is an objective weighting method based on information entropy theory. It automatically determines the weight of each indicator based on its degree of dispersion, avoiding bias that can be introduced by subjective weighting. It is highly scientific and reliable.

[0227] The entropy method originates from information theory. Its core idea is to evaluate the information content of an indicator by calculating the degree of dispersion of the indicator data (i.e., information entropy). The greater the information entropy, the smaller the dispersion of the indicator, the less information it provides, and its weight should be reduced accordingly. Conversely, the smaller the information entropy, the greater the dispersion of the indicator, the more information it provides, and its weight should be increased accordingly. The mathematical expression of the entropy method is as follows:

[0228] ① Standardization processing:

[0229] Each indicator is normalized to the minimum and maximum values to eliminate the dimension difference:

[0230]

[0231] Among them, X gf is the fth index value of the gth grid, X fmin and X fmax are the minimum and maximum values of the f-th indicator respectively.

[0232] ②Calculate the proportion of indicators:

[0233] Calculate the proportion p of index f of each grid g gf :

[0234]

[0235] ③Calculate information entropy:

[0236]

[0237] Where: g is the grid number, f is the index number, and N is the number of grids;

[0238] S4.2.3. Calculate the weight of each indicator ω based on information entropy f , the specific formula is as follows:

[0239]

[0240] Where: F is the number of indicators, F = 5;

[0241] S4.2.4. Calculate the rationality value Q of each grid based on the weight of the indicator i , the specific calculation formula is as follows:

[0242]

[0243] Among them: F daily (g) is the annual daily average exceeding the standard frequency, ω1 is The weight of ω2 is The weight of ω3 is The weight of E annual The weight of (g), ω5 is F daily The weight of (g).

[0244] S4.3. Sort the grids from high to low rationality and prioritize the site placement in the study area based on the rationality values. Specifically, the top 30% of the rationality values are assigned high priority; the middle 40% of the rationality values are assigned medium priority; and the bottom 30% of the rationality values are assigned low priority.

[0245] S4.4. A greedy algorithm based on multi-criteria constraints is used to select new sites according to the site layout priority, and the site selection results of the monitoring sites are obtained after optimization from the pollution perspective.

[0246] S4.4 includes:

[0247] S4.4.1. Add the number of new sites based on the grid's site layout priorities, specifically:

[0248] Nhigh =0.5N new ;

[0249] N mid =0.3N new ;

[0250] N low =0.2N new ;

[0251] Where: N high is the number of new sites in the high-priority grid, N mid N is the number of new sites in the medium priority grid. low N is the number of new sites in the low-priority grid; new Set the number of sites to a value between 90% and 110% of the original number of sites;

[0252] S4.4.2. Use a greedy algorithm to find the optimal address, specifically:

[0253] Arrange the grids in descending order of rationality value, put the grids with covered points into the existing site set, and put the grids with uncovered points into the candidate set;

[0254] Iteratively select the grid with the largest rationality value that meets the distance constraint in the candidate set as the new site to be added to the existing site set; and dynamically update the existing site set and candidate set in the study area, where: the distance constraint is that the Haversine distance between any sites is greater than or equal to 1000m.

[0255] In this embodiment, the high-value area coverage and weighted Gini coefficient are used to evaluate the location of the optimized monitoring stations. The specific formula is as follows:

[0256] ① Coverage of high-value areas

[0257]

[0258] Among them, Cov is the coverage of high-value areas of rationality or site optimization comprehensive index, N is N is the number of grids that rank in the top 30% of the rationality or comprehensive index i of the site optimization within the monitoring range of the site. ih The number of grids that rank in the top 30% of the rationality or site optimization comprehensive index of the study area.

[0259] ② Weighted Gini coefficient

[0260] The Gini Coefficient (GC) is a classic indicator for measuring spatial distribution balance. It was originally used in economics to analyze income distribution and was later introduced into geography to evaluate the fairness of spatial distribution. This study used the Weighted Gini Coefficient (WGC) to assess the degree of balance in the spatial layout of monitoring stations. By coupling the spatial location of the stations with the calculation results of the rationality of the station layout in the study area, the balance of the optimized station spatial distribution and its match with environmental requirements were quantified, achieving a quantitative evaluation of the optimized station layout.

[0261]

[0262] in

[0263]

[0264] Where Q g(1)g(2) is the rationality mean of grid g(1) and grid g(2), is the rationality mean in the study area, Q g(1) is the rationality of the grid g(1), Q g(2) is the rationality of the grid g(2); x g(1) is the number of monitoring stations in grid g(1), x g(2) is the number of monitoring stations in grid g(2), is the average number of monitoring stations in each grid.

[0265] In addition, the impact of risks is also considered, including S5, as follows:

[0266] S5: Based on the site selection results of the monitoring stations optimized from the pollution perspective, the monitoring stations are optimized from the risk perspective to obtain comprehensive site optimization results that take into account "pollution-risk" and "present-future".

[0267] The S5 includes:

[0268] S5.1. Calculate the comprehensive exposure risk index (ERI) for each grid, specifically:

[0269] S5.1.1. Construct four exposure risk indicators based on the distribution of patients with cardiovascular and cerebrovascular diseases, allergic diseases, and respiratory diseases, as well as population density and pollutant distribution data;

[0270] S5.1.2. Calculate the exposure risk index corresponding to the exposure risk indicator. The specific calculation formula is as follows:

[0271]

[0272] Where: b is the number of the exposure risk indicator, T gbis the exposure risk index of the b-th indicator in the g-th grid in the study area, P gb is the population of the b-th category indicator in the g-th grid, C g is the pollution concentration value of the g-th grid;

[0273] S5.1.3. Standardize the exposure risk index and perform equal weight fusion on the standardized exposure risk to obtain the comprehensive exposure risk index ERI. g , the specific calculation formula is as follows:

[0274]

[0275] Where: T gb * is the standardized exposure risk index;

[0276] S5.2. Constructing a comprehensive site optimization index Z based on the risk perspective based on the comprehensive exposure risk index g , the specific formula is as follows:

[0277] Z g =0.4ERI g +0.2R p +0.2R b +0.2R pop ;

[0278] Where: R p is the pollutant concentration prediction, R b For building density prediction, R pop For population density prediction;

[0279] S5.3. Taking district or county as the basic spatial unit, sort the monitoring stations in each unit by comprehensive index, and remove the Z g <Z threshold In this embodiment, Z threshold Take the 50% quantile of the composite index;

[0280] S5.4. Based on the greedy algorithm, the site selection results of the monitoring stations after the pollution perspective optimization are further optimized to obtain the comprehensive site optimization results that take into account "pollution-risk" and "present-future".

[0281] The 5.4 includes:

[0282] S5.4.1. Select all the items that satisfy Z g ≥Z threshold The grid is used as the candidate site selection point;

[0283] S5.4.2. Count the number of sites in each unit. If the number of sites in each unit does not exceed Q, the candidate points in each unit are sorted according to Z.g Arrange in descending order and iteratively select the site optimization comprehensive index Z that meets the spatial constraints in each unit g The largest candidate point is selected as a new site until the number of sites in each cell exceeds Q, where Q is a set value and the spatial constraint is that the distance between the new site and the existing site is no less than 1.5 km; in this embodiment, Q is 8.

[0284] S5.4.3. Select sites based on the entire study area and iteratively select sites within the study area that meet the spatial constraints and optimize the comprehensive index Z. g The largest candidate point is selected as the new site; the termination condition is that the total number of sites in the study area reaches the upper limit of the total number of sites or the total number of sites reaches the lower limit of the total number of sites, and the grids with the top 50% of the site optimization comprehensive index in the study area have sites deployed; finally, the comprehensive point optimization result taking into account "pollution-risk" and "present-future" is obtained; where: the total number of sites N final ∈[150,183].

[0285] In this embodiment, after S5.4.3, the high-value area coverage and weighted Gini coefficient can be used to evaluate the optimization results of the monitoring site locations.

[0286] The monitoring site selection optimization method in this embodiment is to analyze the PM 2.5 The consistency between the concentration remote sensing data and the ground monitoring data of national monitoring and encrypted stations reveals that PM 2.5 The spatial differentiation characteristics of the credibility of concentration remote sensing data are analyzed, and the feasibility of using quantitative remote sensing data for optimizing monitoring site locations is analyzed. On this basis, the impact of PM 2.5 The key factor in the accuracy differentiation of concentration remote sensing data is to introduce a random forest model with recursive feature elimination and cross-validation to improve its local fitting accuracy. By comparing the consistency measurement indicators of remote sensing data before and after correction, it is found that the credibility of remote sensing data at national control sites is higher, while the credibility at encrypted sites has significant spatial heterogeneity. After the random forest model correction of this embodiment, the credibility of remote sensing data is significantly improved, R 2 The RMSE increased from 0.87 to 0.98, and the RMSE increased from 8.59 μg / m 3 Reduced to 3.08 μg / m 3 The high deviation rate was reduced from 2.01% to 0.04%, and the uncertainty was reduced from 18.93% to 8.27%, providing a reliable pollution distribution data basis for subsequent site optimization.

[0287] This example also establishes a three-dimensional evaluation system based on a pollution perspective: "Representativeness (spatial representation range) - Redundancy (comprehensive evaluation index of site redundancy) - Rationality." This system systematically assesses the spatial coverage, redundancy characteristics, and layout rationality of monitoring sites, and optimizes site locations based on the evaluation results. In this example, the original 167 monitoring sites exhibited significant spatial imbalance: 46.7% of the sites had a representative range of less than 100 meters, indicating poor representativeness. Furthermore, 45 pairs of redundant sites with overlapping monitoring ranges and highly similar geographic environments and monitoring data were identified. Using a greedy algorithm, 86 inefficient sites were eliminated and 102 new sites were added, bringing the total number of optimized sites to 183. The average representative range radius of the sites increased from 1.89 km to 3.93 km, representing a 37.71% increase in representative range area. Coverage of heavily polluted areas increased by 33.34%, and the Gini coefficient, weighted by rationality, decreased by 68.57%. This represents a shift from a centralized distribution of monitoring sites to a more balanced, pollution-oriented distribution, providing reliable data support for refined urban pollution management.

[0288] This embodiment also integrates disease medical treatment data with PM 2.5 Pollution distribution and the establishment of an exposure risk assessment system for sensitive populations. Based on the optimization of monitoring site locations from a pollution perspective, coupled with comprehensive exposure risk and future development forecasts, a comprehensive "pollution-risk" and "present-future" optimization model was constructed. After optimization, the representative range of the sites expanded to 2419.81 km², a 33.21% increase compared to before optimization. The average representative range radius of the sites increased from 1.89 km to 3.96 km, increasing coverage of severely polluted areas by 23.14% and coverage of high-risk areas by 24.95%. Furthermore, the Gini coefficient, weighted by the comprehensive risk index, decreased from 0.12 to 0.07, significantly improving the spatial equity of monitoring resource allocation. The adjusted site layout, while maintaining monitoring capacity, increases coverage of high-exposure risk areas and can adapt to future changes in urban development patterns, providing more forward-looking monitoring data support for environmental management.

[0289] This embodiment also includes a readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the above-mentioned method for optimizing the site selection of air quality monitoring stations is implemented.

[0290] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0291] This embodiment also includes an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory. When the computer program instructions are executed by the processor, the air quality monitoring station site selection optimization method as described above is performed.

[0292] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0293] The electronic device may be a computing device such as a mobile phone, desktop computer, laptop, PDA, or cloud server. The electronic device may include, but is not limited to, a processor and memory. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0294] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire electronic device using various interfaces and lines.

[0295] The memory can be used to store the computer program and / or module, and the processor implements the computer program by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0296] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device that can carry the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium, etc.

[0297] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for optimizing the site selection of air quality monitoring stations, characterized in that: The following steps are involved: S1: Obtain remote sensing monitoring data and ground station monitoring data of pollutant concentrations in a certain study area; S2: Correct the remote sensing monitoring data based on the random forest model to obtain the corrected monitoring data; S3: Evaluate the credibility of the corrected monitoring data based on the consistency measurement index. If the consistency measurement index of the corrected monitoring data is higher than the consistency measurement index of the remote sensing monitoring data, proceed to S4; otherwise, return to S2. S4: Optimize the monitoring sites from a pollution perspective based on the corrected monitoring data to obtain the site selection results of the monitoring sites after the pollution perspective optimization.

2. The method for optimizing the site selection of air quality monitoring stations according to claim 1, characterized in that: The S2 includes: S2.

1. Use natural and human environmental data as modeling variables and perform preliminary screening of these variables using significance analysis and collinearity tests to obtain a selected subset of modeling features. Preliminary screening involves excluding modeling variables with a P-value > 0.05 and a VIF > 10. Natural environmental data include DEM data, temperature, wind direction, humidity, and wind speed; and human environmental data include land use, population density, point of interest (POI) data, road traffic data, and emission inventories. S2.

2. Use all features in the current modeling feature subset to train the pre-trained model; S2.

3. Use recursive feature elimination to eliminate unimportant features in the modeling feature subset; S2.

4. Use cross-validation to calculate the model error of the current modeling feature subset; S2.

5. Repeat S2.2 to S2.4 until the model error reaches the minimum value, and a random forest model is obtained; S2.

6. Input the remote sensing monitoring data into the random forest model, and output the corrected monitoring data; In S3, the consistency measurement indicators include goodness of fit, root mean square error, high deviation rate, uncertainty and spatial differentiation characteristics.

3. The method for optimizing the site selection of air quality monitoring stations according to claim 2, characterized in that: The S4 includes: S4.

1. Eliminate sites with insufficient representation and sites with high redundancy and low coverage. Retain sites with a representation radius greater than 4000m, even if the site has a high redundancy count. A site with insufficient representation is defined as one with a spatial representation range less than or equal to the first radius threshold. A site with high redundancy and low coverage is defined as one with a redundancy count greater than the median and a spatial representation range radius less than the second radius threshold. The redundancy count refers to the number of times the comprehensive redundancy evaluation index of the site that overlaps with other monitoring ranges is greater than the median. S4.

2. Divide the study area into multiple grids and calculate the rationality value of each grid based on a multi-dimensional spatial statistical indicator system; S4.

3. Sort the grids from high to low rationality and prioritize the site placement in the study area based on the rationality values. Specifically, the top 30% of the rationality values are assigned high priority; the middle 40% of the rationality values are assigned medium priority; and the bottom 30% of the rationality values are assigned low priority. S4.

4. A greedy algorithm based on multi-criteria constraints is used to select new sites according to the site layout priority, and the site selection results of the monitoring sites are obtained after optimization from the pollution perspective.

4. The method for optimizing the site selection of air quality monitoring stations according to claim 3, characterized in that: The spatial representative range in S4.1 is obtained by using buffer zone analysis and goodness of fit R 2 Calculate the spatial representativeness of each site in the study area, including: With each monitoring site i as the center, buffer zones with radii of 100m, 200m, 300m, 400m, 500m, 600m, 700m, 800m, 900m, 1000m, 1500m, 2000m, 2500m, 3000m, 3500m, and 4000m are generated, denoted as B. s (r), where r is the buffer radius; For each buffer B i (r), extract the remote sensing grid data within its range and calculate the mean of the remote sensing grid data Calculate the pollutant concentration monitoring value S at the site i and buffer mean Goodness of fit between the mean Goodness of fit based on the mean Determine the spatial representativeness of the monitoring sites, specifically: For each monitoring site i, traverse all buffer zones with radius r and select The minimum radius r min As the spatial representative range of the site, the specific formula is as follows: r min =min{r|R 2 (r)<0.9}; If there is no radius r that satisfies R 2 (r) ≥ 0.9, then the site space representative range is 0. If all radii r satisfy R 2 If (r) ≥ 0.9, the spatial representative range of the site is marked as a maximum value of 4000 m; In S4.1, the method for obtaining the comprehensive evaluation index of site redundancy is to calculate the comprehensive evaluation index of site redundancy of sites within the overlapping area based on the similarity of geographical environment and the similarity of ground site monitoring data, specifically: ① Calculate geographical environment similarity, including: Select geographical elements and divide them into five levels using the natural breakpoint method; The explanatory power of discretized geographical elements on geographical phenomena is quantified by calculating the q value. The specific formula is as follows: Among them, N h and are the number of samples and variance of the h-th layer, N and σ respectively 2 are the total number of samples and the overall variance, respectively, and L is the number of strata; Select geographic features with q values greater than the set threshold or ranked in the top X as input variables for similarity measurement; Calculate the multidimensional Euclidean distance D between sites based on the input variables of the similarity metric ik , the specific formula is as follows: Where: i and k are site numbers, i≠k; X ij is the jth geographic feature value of the i-th site, X kj is the jth geographic element value of the kth site; Z(.) is the normalization process; m is the number of geographic elements; Normalize the multidimensional Euclidean distance to the range [0,1] to obtain the geographical similarity Among them, max(D) is the maximum value of the Euclidean distance of all site pairs; ②Calculate the similarity of monitoring data, including: Calculate the Pearson correlation coefficient r between site i and site k ik : in, and are the mean pollutant concentrations at site i and site k, S i (t) and S k (t) are the mean pollutant concentrations at site i and site k on day t respectively; T is the number of days with statistically valid values of ground site monitoring data of pollutant concentration; Normalize the correlation coefficient to the range of [0,1] to obtain the similarity of monitoring data The specific formula is as follows: ③Calculate the overlapping area of the monitoring site pairs with overlapping monitoring ranges, specifically: A double circle geometric model is established for site i and site k, with radii r i and r k ; Calculate the distance d between site i and site k ik , when d ik ≥r i +r k When d ik ≤|r i -r k |, it is determined that site i and site k completely overlap; in other cases, site i and site k partially overlap; Calculate the overlapping area of site i and site k; ④ Based on geographical similarity Similarity with monitoring data Calculate the comprehensive evaluation index of the redundancy degree of sites within the overlapping area. The specific formula is as follows: Among them, ω is the geographical environment similarity weight.

5. The method for optimizing the site selection of air quality monitoring stations according to claim 4, characterized in that: The S4.2 includes: S4.2.

1. Divide the study area into multiple grids; S4.2.

2. Calculate the information entropy of each indicator in the multidimensional spatial statistical indicator system within each grid, where the multidimensional spatial statistical indicator system includes the annual average concentration of pollutants Temporal variability Spatial variability Annual average concentration exceeding the standard E annual (g) and the annual daily average exceeding frequency F daily (g), the specific formula is as follows: Where: g is the grid number, f is the index number, and N is the number of grids; S4.2.

3. Calculate the weight of each indicator ω based on information entropy f , the specific formula is as follows: Where: F is the number of indicators; S4.2.

4. Calculate the rationality value Q of each grid based on the weight of the indicator g , the specific calculation formula is as follows: Among them: F daily (g) is the annual daily average exceeding the standard frequency, ω1 is The weight of ω2 is The weight of ω3 is The weight of E annual The weight of (g), ω5 is F daily The weight of (g).

6. The method for optimizing the site selection of air quality monitoring stations according to claim 5, characterized in that: The S4.4 includes: S4.4.

1. Add the number of new sites based on the grid's site layout priorities, specifically: N high =0.5N new ; <h2 style=";text-align:left;direction:ltr">N<h2 style=";text-align:left;direction:ltr"> mid <h2 style=";text-align:left;direction:ltr"> <0.3N<h2 style=";text-align:left;direction:ltr"> new <h2 style=";text-align:left;direction:ltr"> ; <h2 style=";text-align:left;direction:ltr">N<h2 style=";text-align:left;direction:ltr"> low <h2 style=";text-align:left;direction:ltr"> <0.2N<h2 style=";text-align:left;direction:ltr"> new <h2 style=";text-align:left;direction:ltr"> ; Where: N high is the number of new sites in the high-priority grid, N mid N is the number of new sites in the medium priority grid. low N is the number of new sites in the low-priority grid; new Set the number of sites to a value between 90% and 110% of the original number of sites; S4.4.

2. Use a greedy algorithm to find the optimal address, specifically: Arrange the grids in descending order of rationality value, put the grids with covered points into the existing site set, and put the grids with uncovered points into the candidate set; Iteratively select the grid with the largest rationality value that meets the distance constraint in the candidate set as the new site to be added to the existing site set; and dynamically update the existing site set and candidate set in the study area, where: the distance constraint is that the Haversine distance between any sites is greater than or equal to 1000m.

7. The method for optimizing the site selection of air quality monitoring stations according to claim 6, characterized in that: It also includes the steps for optimizing monitoring sites from a risk perspective based on the site selection results of the monitoring sites optimized from a pollution perspective, specifically: S5.

1. Calculate the comprehensive exposure risk index (ERI) for each grid, specifically: S5.1.

1. Construct four exposure risk indicators based on the distribution of patients with cardiovascular and cerebrovascular diseases, allergic diseases, and respiratory diseases, as well as population density and pollutant distribution data; S5.1.

2. Calculate the exposure risk index corresponding to the exposure risk indicator. The specific calculation formula is as follows: Where: b is the number of the exposure risk indicator, T bg is the exposure risk index of the b-th indicator in the g-th grid in the study area, P gb is the population of the b-th category indicator in the g-th grid, C g is the pollution concentration value of the g-th grid; S5.1.

3. Standardize the exposure risk index and perform equal weight fusion on the standardized exposure risk to obtain the comprehensive exposure risk index ERI. g , the specific calculation formula is as follows: Where: T gb * is the standardized exposure risk index; S5.

2. Constructing a comprehensive site optimization index Z based on the risk perspective based on the comprehensive exposure risk index g , the specific formula is as follows: Z g =0.4ERI g +0.2R p +0.2R b +0.2R pop ; Where: R p is the pollutant concentration prediction, R b For building density prediction, R pop For population density prediction; S5.

3. Taking district or county as the basic spatial unit, sort the monitoring stations in each unit by comprehensive index, and remove the Z g <Z threshold 's site; S5.

4. Based on the greedy algorithm, the monitoring site selection results after pollution perspective optimization are further optimized to obtain the comprehensive point optimization results.

8. The method for optimizing the site selection of air quality monitoring stations according to claim 7, characterized in that: The 5.4 includes: S5.4.

1. Select all the items that satisfy Z g ≥Z threshold The grid is used as the candidate site selection point; S5.4.

2. Count the number of sites in each unit. If the number of sites in each unit does not exceed Q, the candidate points in each unit are sorted according to Z. g Arrange in descending order and iteratively select the site optimization comprehensive index Z that meets the spatial constraints in each unit g The largest candidate point is selected as the new site until the number of sites in each cell exceeds Q, where Q is a set value and the spatial constraint is that the distance between the new site and the existing site is no less than 1.5 km; S5.4.

3. Select sites based on the entire study area and iteratively select sites within the study area that meet the spatial constraints and optimize the comprehensive index Z. g The largest candidate point is selected as the new site; the termination condition is that the total number of sites in the study area reaches the upper limit of the total number of sites or the total number of sites reaches the lower limit of the total number of sites, and the grids with the top 50% of the site optimization comprehensive index in the study area have sites deployed; finally, the comprehensive point optimization result taking into account pollution and risk is obtained; where: the total number of sites N final ∈[150,183].

9. A readable storage medium, characterized in that: Computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, the method for optimizing the site selection of an air quality monitoring station according to any one of claims 1 to 8 is implemented.

10. An electronic device, characterized in that: include: At least one processor, at least one memory, and computer program instructions stored in the memory, when the computer program instructions are executed by the processor, the air quality monitoring site selection optimization method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Atmospheric monitoring stationing optimization identification method, device and equipment and storage medium

    CN121031344A

  • Building project block chain sandbox evaluation system and method

    CN121304096A

  • A blockchain sandbox assessment system and method for construction projects

    CN121304096B

  • Water body monitoring station layout optimization method, equipment and computer readable medium

    CN121480889A