Atmospheric pollution analysis method and device based on big data portrait, equipment and medium

By constructing a pollutant concentration analysis model based on big data and using cluster analysis, and utilizing SHAP values ​​to reflect the degree of contribution of pollution characteristics, the problem of one-sidedness in pollution analysis of streets and towns has been solved, enabling precise source tracing and pollution category classification, and improving the accuracy of environmental monitoring and collaborative management capabilities.

CN120832539BActive Publication Date: 2026-01-27BEIJING MUNICIPAL ENVIRONMENTAL MONITORING CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511316967.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-01-27
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

In existing technologies, the analysis of air pollution in streets and towns lacks verification of actual environmental data, resulting in one-sided research. Traditional methods rely on subjective experience and cannot identify specific pollution characteristics.

Method used

By acquiring multi-source data, a pollutant concentration analysis model is constructed. The SHAP value is used to reflect the contribution of pollution characteristics, a feature vector is constructed, and a pollution profile is generated through cluster analysis, thereby achieving accurate source tracing and pollution category classification.

Benefits of technology

It enhances the interpretability of the model, improves the accuracy of environmental monitoring, and enables the identification of pollution transmission channels and promotes collaborative pollution control across streets and townships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832539B_ABST
    Figure CN120832539B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of based on big data portrait atmospheric pollution analysis method, device, equipment and medium, it is related to atmospheric environment monitoring technical field, the method includes obtaining the multi-source data of street and township, constructs pollutant concentration analytical model, its target variable is pollutant concentration, and its pollution feature is meteorological data and pollution source emission related data;Based on pollutant concentration analytical model, obtain each pollution feature SHAP value;Based on SHAP value, construct the feature vector of each street, township;Based on feature vector, obtain the pollution type of each street, township by cluster analysis, obtain feature matrix by matrix construction;Based on cluster analysis result and feature matrix, generate the portrait reflecting the pollution situation of each street, township, based on the portrait carries out atmospheric pollution analysis.The scheme solves the problem of difficult quantification and evaluation limitation of atmospheric pollution related parameter evaluation weight assignment, improves the accuracy of atmospheric pollution analysis, and can accurately trace the source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring technology, and in particular to an air pollution analysis method, device, equipment, and medium based on big data profiling. Background Technology

[0002] As air pollution control efforts continue and pollutant concentrations decline, the space for sustained improvement in ambient air quality is shrinking. The scope for localized emission reduction needs to be clearly defined. Under this new demand for spatial management, streets and towns, as important management units within cities, have become key targets for implementing refined pollution reduction management measures. Simultaneously, this demand also sets higher standards and requirements for the analysis of pollution characteristics in streets and towns.

[0003] With the acceleration of urbanization, traditional monitoring and evaluation methods for air pollution control often suffer from limitations due to their reliance on single analytical indicators, failing to identify specific pollution characteristics. Furthermore, traditional evaluation studies frequently rely on subjective experience, lacking objective evaluation methods. Currently, street and township profiles lack verification of actual pollution data, and street- and township-level analyses of air pollution exhibit a one-sidedness. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an air pollution analysis method based on big data profiling to address the technical problems in existing environmental monitoring technologies, such as the lack of verification of actual environmental data and the one-sided nature of research in street and town profiling. The method includes:

[0005] Acquire multi-source data from streets and townships, including basic information data of streets and townships, pollutant concentration data, meteorological data, and data related to pollution source emissions;

[0006] A pollutant concentration analysis model is constructed based on the multi-source data. The target variable of the pollutant concentration analysis model is the pollutant concentration, and the pollution characteristics of the pollutant concentration analysis model are meteorological data and pollution source emission-related data.

[0007] Based on the pollutant concentration analysis model, the SHAP value of each pollution feature is obtained, and the SHAP value is used to reflect the degree of contribution of each pollution feature to the pollutant concentration.

[0008] Based on the SHAP value of each pollution feature, construct feature vectors for each street and township;

[0009] Based on the feature vectors of each street and township, the pollution type of each street and township is obtained through cluster analysis, and the feature matrix is ​​obtained through matrix construction.

[0010] Based on the cluster analysis results and feature matrix, a profile reflecting the pollution situation of each street and town is generated, and air pollution analysis is performed based on the profile.

[0011] This invention also provides an air pollution analysis device based on big data profiling to address the technical problems in existing environmental monitoring technologies, such as the lack of verification of actual environmental data for street and town profiles and the one-sided nature of research. The device includes:

[0012] The data acquisition module is used to acquire multi-source data from streets and townships, including basic information data of streets and townships, pollutant concentration data, meteorological data, and data related to pollution source emissions.

[0013] The model building module is used to build a pollutant concentration analysis model based on the multi-source data. The target variable of the pollutant concentration analysis model is the pollutant concentration, and the pollution characteristics of the pollutant concentration analysis model are meteorological data and pollution source emission-related data.

[0014] The SHAP value calculation module is used to obtain the SHAP value of each pollution feature based on the pollutant concentration analysis model. The SHAP value is used to reflect the degree of contribution of each pollution feature to the pollutant concentration.

[0015] The feature vector construction module is used to construct feature vectors for each street and township based on the SHAP value of each pollution feature;

[0016] The clustering module is used to obtain the pollution type of each street and town through cluster analysis based on the feature vectors of each street and town, and to obtain the feature matrix through matrix construction.

[0017] The analysis module is used to generate a profile reflecting the pollution situation of each street and town based on the cluster analysis results and feature matrix, and to perform air pollution analysis based on the profile.

[0018] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned atmospheric pollution analysis methods based on big data profiling, in order to solve the technical problems in the prior art of environmental monitoring, such as the lack of verification of actual environmental data for street and town profiles and the one-sidedness of research.

[0019] This invention also provides a computer-readable storage medium storing a computer program that executes any of the above-described atmospheric pollution analysis methods based on big data profiling, in order to solve the technical problems in existing environmental monitoring, such as the lack of verification of actual environmental data for street and township profiles and the one-sided nature of research.

[0020] Compared with the prior art, the beneficial effects that the above-mentioned at least one technical solution adopted in the embodiments of this specification can achieve include at least the following: The present invention uses SHAP (Shapley Additive exPlanations) values ​​as an extraction method for constructing pollution feature vectors, which solves the problem of difficulty in quantifying the evaluation weight assignment of air pollution-related parameters in streets and towns, and enhances the interpretability of the model; by comprehensively utilizing environmental monitoring data, meteorological data, and pollution emission data, it solves the limitation of traditional methods that mainly rely on human experience for assessment, and improves the accuracy of environmental monitoring; based on the atmospheric pollution feature profile, accurate source tracing can be carried out, and the causes of pollution can be analyzed in combination with meteorological data; in addition, through clustering, it is possible to classify pollution categories in streets and towns, as well as identify pollution transmission channels, and promote cross-street and town collaborative pollution control. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of the air pollution analysis method based on big data profiling provided in an embodiment of the present invention;

[0023] Figure 2 This is another flowchart of the air pollution analysis method based on big data profiling provided in the embodiments of the present invention;

[0024] Figure 3 This is a flowchart of the centroid update process provided in an embodiment of the present invention;

[0025] Figure 4 This is a structural block diagram of a computer device provided in an embodiment of the present invention;

[0026] Figure 5 This is a structural block diagram of an air pollution analysis device based on big data profiling provided in an embodiment of the present invention. Detailed Implementation

[0027] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0028] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] Currently, pollution assessment in streets and townships primarily relies on monitoring results of six air quality parameters. With continuous technological advancements, on the one hand, the availability of various environmental sensing data has been greatly improved, providing a data foundation for big data analysis. On the other hand, artificial intelligence algorithms, represented by machine learning, are constantly being updated and iterated, providing an algorithmic foundation for big data analysis. The method and system proposed in this application, based on big data methods and combined with self-developed algorithms, constructs a detailed profile of air pollution characteristics in streets and townships, providing further technical support for air pollution prevention and control work in streets and townships.

[0030] In this embodiment of the invention, an air pollution analysis method based on big data profiling is provided, such as... Figure 1 As shown, the method includes:

[0031] Step S101: Obtain multi-source data of streets and towns, including basic information data of streets and towns, pollutant concentration data, meteorological data and pollution source emission related data;

[0032] Step S102: Construct a pollutant concentration analysis model based on the multi-source data. The target variable of the pollutant concentration analysis model is the pollutant concentration, and the pollution characteristics of the pollutant concentration analysis model are meteorological data and pollution source emission-related data.

[0033] Step S103: Based on the pollutant concentration analysis model, obtain the SHAP value of each pollution feature, wherein the SHAP value is used to reflect the degree of contribution of each pollution feature to the pollutant concentration;

[0034] Step S104: Based on the SHAP value of each pollution feature, construct the feature vector of each street and township;

[0035] Step S105: Based on the feature vectors of each street and township, obtain the pollution type of each street and township through cluster analysis, and obtain the feature matrix through matrix construction;

[0036] Step S106: Based on the cluster analysis results and feature matrix, generate a profile reflecting the pollution situation of each street and town, and conduct air pollution analysis based on the profile.

[0037] In this embodiment, SHAP (Shapley Additive exPlanations) values ​​are used as the extraction method for constructing pollution feature vectors, which solves the problem of difficulty in quantifying the evaluation weights of air pollution-related parameters in streets and towns, and enhances the interpretability of the model. By comprehensively utilizing environmental monitoring data, meteorological data, and pollution emission data, the limitations of traditional methods that mainly rely on human experience for assessment are overcome, and the accuracy of environmental monitoring is improved. Based on the atmospheric pollution feature profile, the dominant pollution sources are identified by SHAP value ranking and multi-source data fusion, which can achieve accurate source tracing and analyze the causes of pollution by combining meteorological data. In addition, through clustering, the pollution categories of streets and towns can be classified, and pollution transmission channels can be identified, promoting collaborative pollution control across streets and towns.

[0038] In practice, the collection of multi-source data mainly includes four categories: basic information on streets and townships, pollutant concentrations, meteorological data, and data related to pollution source emissions.

[0039] Basic information on streets and townships: Using big data collection methods, basic information such as the names, locations, and administrative divisions of streets and townships are obtained;

[0040] Pollutant concentration: High-density atmospheric monitoring nodes are set up at the street and township levels to obtain PM2.5 and TSP monitoring data;

[0041] Meteorological data: Meteorological data is obtained from meteorological monitoring stations, mainly including wind speed, wind direction, temperature, humidity, etc.

[0042] Pollution Source Emissions: Data related to pollution emissions is obtained through monitoring, surveys, and other means, mainly including five categories: industrial sources, dust sources, mobile sources, residential sources, and natural sources. Industrial sources are obtained by combining key pollution source monitoring data, electricity consumption monitoring, and discharge permit data to determine activity levels or emission intensity. Dust sources include road dust data such as road network data and road dust load data; construction site dust data is obtained through satellite remote sensing data to determine construction site area and covering status; mobile sources mainly collect traffic flow monitoring data; residential sources are mainly catering sources, obtaining catering source data and locations through map POIs, combined with oil fume monitoring data to determine activity levels or emission intensity; natural sources are mainly wildfires and burning, relying on satellite remote sensing technology to obtain fire point data.

[0043] In one embodiment, refer to Figure 2The method further includes: preprocessing the acquired multi-source data, mainly including missing value processing, outlier processing and spatiotemporal registration, to ensure data accuracy and availability.

[0044] Handling missing data: Missing data is filled using methods such as mean imputation and interpolation;

[0045] Data outlier handling: Identify and correct outliers through statistical analysis to ensure the rationality of the data;

[0046] Spatiotemporal registration: By using high-frequency data statistics and low-frequency data amplification methods, different parameters are aligned in the time and space dimensions and unified to a single scale, so that the data has consistency in time and space.

[0047] In one embodiment, a pollutant concentration analytical model is constructed, and machine learning algorithms such as XGBoost are used to establish a functional relationship with pollutant concentration as the dependent variable and meteorological conditions and pollution emissions as independent variables.

[0048] Dependent variable (target variable): PM2.5 and TSP pollutant concentrations in each street and town.

[0049] Independent variables (pollution characteristics): meteorological data (temperature, humidity, wind speed, wind direction, etc.) and pollution source emission data (industrial sources, dust sources, etc.).

[0050] Through model training, a model of the relationship between pollutant concentration and meteorological conditions and pollution source emissions is obtained.

[0051] In one embodiment, obtaining the SHAP value for each contamination feature includes:

[0052] The initial SHAP value for each contamination feature is calculated based on the spatiotemporal weighting coefficient and the data credibility coefficient.

[0053] The initial SHAP value is corrected based on the pollutant concentration threshold to obtain the corrected SHAP value.

[0054] Based on pollutant concentration thresholds, pollutant concentration gradient weighting coefficients are set;

[0055] The corrected SAP values ​​and pollutant concentration gradient weighting coefficients are integrated to obtain the final SAP value for each pollution feature, which serves as the SAP value for each pollution feature.

[0056] In practice, based on the trained pollutant concentration analysis model, the influence and contribution of different parameters (pollution characteristics) on pollutant concentration are calculated. The specific quantitative assessment indicator uses the SHAP value to measure the contribution of each characteristic (such as meteorological conditions, pollution source emissions, etc.) to the model output (pollutant concentration). By calculating the SHAP value of each characteristic, it is possible to determine which factors have the greatest impact on pollutant concentration, as well as the direction and extent of that impact.

[0057] This embodiment introduces a triple correction mechanism: ① Eliminating spatiotemporal bias in data through spatiotemporal weighting coefficients (streets and towns with closer time and distance have higher weights); ② Correcting the impact of missing and outlier values ​​using data reliability coefficients; ③ Dynamically amplifying the contribution of high-concentration scenarios based on the concentration gradient weighting coefficients of PM2.5 and TSP. Finally, the SHAP value is iteratively optimized using dual-pollutant exceedance correction factors to form a final value reflecting the true pollution contribution, and a tree model is used for the optimization framework.

[0058] Furthermore, considering the spatiotemporal characteristics of contamination data and the need for multi-source fusion, a spatiotemporal weighting coefficient and a data reliability coefficient are introduced. The initial SHAP value is calculated using the following formula:

[0059] ,

[0060] in, The initial SHAP value is the i-th pollution characteristic (such as the emission intensity of a factory) of the j-th street or township, representing its contribution to the real-time pollutant concentration of that street or township; The spatiotemporal weight coefficients of the j-th street / township are determined by the time decay factor. Spatial distance factor calculate; The data reliability coefficient for the j-th street / township is dynamically adjusted based on the preprocessing results (missing value imputation, outlier correction). The multi-source fusion weight for the j-th street / township; Real-time concentration change when the feature subset S of the j-th street / township is added to the i-th pollution feature; S is any subset of features that does not contain the i-th pollution feature. It is used to measure the impact of adding pollution feature i to subset S on the model's predicted value when calculating the SHAP value of pollution feature i. Traditional subset weights ensure a fair allocation of feature contributions;

[0061] The contribution of pollutants from both streets and townships is corrected. When the pollutant concentration exceeds a threshold, the SHAP value is corrected according to the following rules. The expression for the corrected SHAP value is as follows:

[0062] ,

[0063] in, C is the corrected SHAP value of the i-th pollution characteristic of the j-th street / township. PM2.5 To measure the concentration of pollutant PM2.5, C t-PM2.5 C represents the threshold concentration of pollutant PM2.5. TSP To measure the concentration of pollutant TSP, C t-TSP The threshold for TSP concentration;

[0064] The pollutant concentration gradient weighting coefficients include the concentration gradient weighting coefficients for PM2.5 and TSP, with the following expressions:

[0065] ,

[0066] ,

[0067] Among them, f g-PM2.5,j f represents the concentration gradient weighting coefficient of PM2.5 in the j-th street / township. g-TSP,j Let C be the concentration gradient weighting coefficient of TSP in the j-th street / township. PM2.5,j C represents the measured PM2.5 concentration in the j-th street / township. TSP,j Let C be the measured TSP concentration of the j-th street / township. t-PM2.5,j Let C be the PM2.5 concentration threshold for the j-th street / township. t-TSP,j Let be the TSP concentration threshold for the j-th street / township.

[0068] In one embodiment, for the constructed XGBoost model, a hierarchical feature grouping algorithm is used to group and calculate the pollution features of streets and towns according to categories (industrial sources, dust sources, etc.). The SHAP value formula, spatiotemporal registration weights, and threshold correction factors are then integrated to obtain a complete formula for calculating the SHAP value contribution of streets and towns. Specifically, the integration of the corrected SHAP value and pollutant concentration gradient weight coefficients to obtain the final SHAP value for each pollution feature includes:

[0069] Each pollution feature of a street or township is grouped according to its pollution source. The adjusted SHAP values ​​after grouping are then integrated with the pollutant concentration gradient weighting coefficients to obtain the final SHAP value for each pollution feature. The expression for the final SHAP value is as follows:

[0070] ,

[0071] in, Let be the final SHAP value of the i-th pollution characteristic of the j-th street / township. G represents the final SHAP value of the g-th feature group of the j-th street / township. j Let G be the set of feature groups for the j-th street / township, and g be the set of feature groups. j The g-th feature group in Let g be the spatiotemporal weight coefficient of the g-th feature group in the j-th street or township (for example, if the data within the industrial source group of the 3rd street or township are all real-time data, then...). ), This is the data reliability coefficient of the g-th feature group in the j-th street or township. When new data is added to the j-th street or township, only the tree path of the corresponding group in that street or township is updated (for example, updating the factory emission data of the 3rd street or township does not affect the calculation of the 4th street or township).

[0072] In one embodiment, the pollution sources include industrial sources, dust sources, mobile sources, residential sources, and natural sources, and the construction of feature vectors for each street and town includes:

[0073] Construct the initial feature vector for each street and town. The expression for the initial feature vector is:

[0074] ,

[0075] in, Let be the initial feature vector of the j-th street / township. Let be the final SHAP value of the industrial source of the j-th street / township. Let be the final SHAP value of the dust source in the j-th street / township. Let be the final SHAP value of the mobile source in the j-th street / township. Let be the final SHAP value of the livelihood source of the j-th street / township. Let be the final SHAP value of the meteorological data for the j-th street / township. Let be the final SHAP value of the natural source of the j-th street / township;

[0076] The initial feature vectors are standardized to obtain the feature vectors for each street and township. The expression for the feature vector is as follows:

[0077] ,

[0078] in, Let j be the feature vector of the j-th street / township. This represents the average pollution characteristics of all streets and towns. The standard deviation is denoted as .

[0079] In practice, based on the SHAP values ​​calculated in the above steps, feature vectors are constructed for each street and township. Each indicator in the feature vector corresponds to a SHAP value for a pollution feature, reflecting the contribution of that pollution feature to the pollutant concentration in that street or township. In this way, the complex pollution situation of each street and township is simplified into a single feature vector, facilitating subsequent analysis and processing.

[0080] Based on the degree of impact of pollutants on streets and towns, independent weighting is used as shown in Table 1 to avoid confusion caused by uniform symbols.

[0081] Table 1 Pollution Characteristic Weight Label Table

[0082]

[0083] The initial feature vector for each street and town is represented as follows:

[0084] ,

[0085] The first six terms of this vector are the calculated pollution features for the j-th street and township, and are the final SHAP values ​​(after spatiotemporal registration, confidence coefficient, and threshold correction).

[0086] To eliminate the influence of different characteristic dimensions, standardization is performed, and the formula is as follows:

[0087] ,

[0088] in This represents the average pollution characteristics of all streets and towns. The standard deviation is denoted as .

[0089] In one embodiment, the cluster analysis employs the K-means algorithm, and obtaining the pollution type of each street and town through cluster analysis includes:

[0090] Based on the elbow rule and the pollutant concentration gradient verification preset rules, the WSS value of the feature vectors of all streets and towns is calculated, and the optimal number of clusters K is determined based on the WSS value.

[0091] Calculate the weighted Euclidean distance between all streets and towns;

[0092] Initialize K centroids and assign each street and town to the nearest cluster based on weighted Euclidean distance;

[0093] Recalculate the weighted mean of the SHAP values ​​of streets and towns within each cluster according to the feature group, and use the weighted mean as the new centroid;

[0094] Repeat the centroid update process until the change in centroid is less than a preset threshold or the maximum number of iterations is reached;

[0095] When the inter-cluster differences in pollutant concentration do not meet the preset rules for pollutant concentration gradient verification, adjust the K value and repeat the centroid update process.

[0096] When the inter-cluster differences in pollutant concentrations meet the preset rules for pollutant concentration gradient verification, the groups of streets and townships are obtained;

[0097] The pollution type of each group is determined based on the grouping of streets and townships.

[0098] In one embodiment, the method further includes:

[0099] When the criteria for determining weighted contribution, consistency of pollutant concentration, and spatiotemporal correlation are met simultaneously, the streets and townships are determined to belong to the same pollution type.

[0100] The criteria for determining the weighted contribution are as follows:

[0101] ,and ,

[0102] The criteria for determining the consistency of pollutant concentrations are as follows:

[0103] ,and ,

[0104] The criteria for determining the spatiotemporal correlation are as follows:

[0105] ,and ,

[0106] Among them, W i W represents the weight of pollution feature of type i. p The weights of the pollution characteristics of class p, For the j-th street / township, μ represents the final SHAP value of the p-th pollution characteristic. PM2.5,k Let μ be the average concentration of PM2.5 within the k-th cluster. TSP,k Let W be the mean concentration of TSP within the k-th cluster, r be the correlation coefficient measuring the spatiotemporal weights of the j-th street and the k-th street, and W be the mean concentration of TSP within the k-th cluster. st,k W represents the spatiotemporal weighting coefficient of the k-th street / township. st,j Let j be the spatiotemporal weight coefficient of the j-th street / township. This represents the average spatiotemporal weights of all streets and towns within the urban area.

[0107] In practice, eigenvector clustering analysis includes the following:

[0108] To reveal the differences in pollution characteristics among different streets and towns, clustering algorithms such as K-means or DBSCAN are used to cluster the feature vectors of each street and town. The purpose of clustering is to group streets and towns with similar pollution characteristics into one category and divide them into several types based on the similarity of their feature vectors, thereby identifying the differences in pollution characteristics among streets and towns.

[0109] In this embodiment, streets and towns with similar pollution characteristics are clustered using the K-means algorithm. The main components of feature vector clustering include: first, determining the optimal number of clusters using the elbow rule to verify the inter-cluster differences in pollutant concentration gradients; then, measuring the similarity of pollution characteristics among streets and towns using weighted Euclidean distance; initializing K centroids, iteratively optimizing the centroids to complete the grouping, and finally determining the pollution type. The specific steps are as follows:

[0110] (1) Determination of the optimal number of clusters K (elbow rule + dual pollutant concentration gradient verification)

[0111] A smaller WSS value indicates more similar pollution characteristics among streets and towns within the cluster. By calculating the corresponding intra-cluster sum of squares for different K values, the inflection point (elbow) of the curve is selected as the optimal K value. The differences between clusters are verified using PM2.5 and TSP concentration gradients to ensure the rationality of the clustering results. The specific formula is as follows:

[0112] ,

[0113] Where WSS is the sum of squares within the cluster, W i For example, W represents the feature weights, such as W for industrial sources. ind k is the number of clusters, cl k For the k-th cluster, μ i,k Let be the centroid of the i-th pollution feature in the k-th cluster (i.e., the mean of the i-th feature of all streets and towns in this cluster).

[0114] PM2.5 concentration gradient verification preset rules: ,

[0115] TSP concentration gradient validation preset rules: ;

[0116] (2) Calculation of comprehensive distance

[0117] Considering the weighted assignment of different pollutant concentrations and different pollution sources, a weighted Euclidean distance is used to measure the similarity of pollution characteristics between streets and towns. The smaller the distance, the more similar the pollution characteristics (weighted differences in SHAP values ​​of each pollution source) are between the two streets and towns.

[0118] ,

[0119] Where, dw For weighted Euclidean distance;

[0120] (3) Iterative clustering and centroid update

[0121] Reference Figure 3 Initialize K centroids (pollution feature benchmarks), assign streets and towns to the nearest cluster, recalculate the weighted average of the SHAP values ​​within the cluster as the new centroids, and repeat until the change in centroids is less than a preset threshold or the maximum number of iterations is reached.

[0122] Centroid initialization: Typical streets and towns with the largest SHAP values ​​for industrial sources and dust sources are selected as the initial centroids.

[0123] Allocation phase: ,

[0124] For samples within the k-th cluster, the new centroid calculation needs to be updated. When updating the centroid, the mean is calculated independently for each feature group:

[0125] ,

[0126] in, Let g be the centroid of the pollution characteristics of the k-th cluster and the g-th group (e.g., g is the pollution characteristic group of the dust source). That is, the centroid of the dust source feature group of all samples in cluster k), |S k | represents the number of streets and towns within the k-th cluster, S k Let k be the set of streets and towns contained in the k-th cluster;

[0127] Iteration termination condition: Change in centroid: , For the updated centroid, The centroid before the update; or the maximum number of iterations has been reached.

[0128] (4) Determination of similar pollution characteristics

[0129] The system determines whether streets and towns belong to the same pollution type based on three dimensions: weighted contribution, pollutant concentration, and spatiotemporal correlation. This ensures that the clustering results not only conform to the SHAP value contribution similar pollution characteristic determination model, but also match the actual concentration level.

[0130] ① Weighted contribution determination:

[0131] For pollution type i, the following must be satisfied:

[0132] ,and ;

[0133] ② Verification of pollutant concentration consistency:

[0134] Compare the deviation between the measured concentration and the average concentration within the cluster:

[0135] ,and ;

[0136] ③ Determination of spatiotemporal correlation:

[0137] Calculate the correlation coefficient of the spatiotemporal weight vectors among streets and towns to ensure that the spatiotemporal characteristics of streets and towns in the same cluster are consistent.

[0138] ,

[0139] and .

[0140] In one embodiment, the generation of images of streets and towns includes the following steps:

[0141] Feature matrix construction: Each row represents a feature vector of a street or town, and each column represents a feature dimension.

[0142] Visual mapping: Each street and town is visualized and mapped to its corresponding pollution characteristics in the form of heat maps and radar maps.

[0143] Heatmap: Displays the numerical distribution of each street and town across different feature dimensions, intuitively reflecting the magnitude of the values ​​through color depth.

[0144] Radar chart: Displays the comprehensive performance of each street and town across multiple characteristic dimensions, showcasing the overall situation of its pollution characteristics through the shape and area of ​​the radar chart.

[0145] For example, based on the clustering results, streets and towns can be divided into multiple types. Taking PM2.5 profiling as an example, the results can be as follows:

[0146] Category 1: Highly polluted areas with high PM2.5 concentrations, largely contributed by industrial sources.

[0147] Category 2: Moderately polluted area, with moderate PM2.5 concentration and significant contribution from dust sources.

[0148] Category 3: Low-pollution areas with low PM2.5 concentrations and significant contributions from natural sources.

[0149] Noise points: Data points that do not conform to any cluster and may have special contamination characteristics or data anomalies.

[0150] Through cluster analysis and profiling, the air quality status and pollution source contributions of each street and town can be displayed intuitively.

[0151] In practice, based on the atmospheric pollution characteristic profile, the dominant pollution source is identified by sorting the SHAP value and fusion of multi-source data (e.g., the SHAP value of industrial sources in a certain street or township is 0.45, and the SHAP value of dust sources is 0.28, indicating that the pollution type is industrial-dominated pollution), and the source is accurately traced; and the causes of pollution are analyzed in conjunction with meteorological data.

[0152] Furthermore, by clustering into categories such as "high-pollution industrial type" and "dust-dominated type," the pollution categories of streets and towns can be classified (e.g., the average PM2.5 value of street or town A is 85 μg / m³). 3 Streets and towns with an industrial source SHAP value of 0.6 are classified as high-pollution industrial areas. By utilizing spatiotemporal registration weights and clustering results, pollution transmission channels are identified (e.g., adjacent streets and towns with a spatiotemporal correlation coefficient > 0.7 are identified as the same pollution-affected area), promoting collaborative pollution control across streets and towns.

[0153] In this embodiment, a computer device is provided, such as... Figure 4 As shown, it includes a memory 401, a processor 402, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned atmospheric pollution analysis methods based on big data profiling.

[0154] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.

[0155] In this embodiment, a computer-readable storage medium is provided, which stores a computer program that executes any of the above-described atmospheric pollution analysis methods based on big data profiling.

[0156] Specifically, computer-readable storage media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media does not include transient media, such as modulated data signals and carrier waves.

[0157] Based on the same inventive concept, this invention also provides an air pollution analysis device based on big data profiling, as described in the following embodiments. Since the principle of the air pollution analysis device based on big data profiling is similar to that of the air pollution analysis method based on big data profiling, the implementation of the air pollution analysis device based on big data profiling can refer to the implementation of the air pollution analysis method based on big data profiling, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0158] Figure 5 This is a structural block diagram of an atmospheric pollution analysis device based on big data profiling, according to an embodiment of the present invention. Figure 5 As shown, it includes: a data acquisition module 501, a model building module 502, a SHAP value calculation module 503, a feature vector construction module 504, a clustering module 505, and an analysis module 506. The structure is described below.

[0159] The data acquisition module 501 is used to acquire multi-source data of streets and townships, including basic information data of streets and townships, pollutant concentration data, meteorological data and pollution source emission related data;

[0160] The model building module 502 is used to build a pollutant concentration analysis model based on the multi-source data. The target variable of the pollutant concentration analysis model is the pollutant concentration, and the pollution characteristics of the pollutant concentration analysis model are meteorological data and pollution source emission-related data.

[0161] The SHAP value calculation module 503 is used to obtain the SHAP value of each pollution feature based on the pollutant concentration analysis model. The SHAP value is used to reflect the degree of contribution of each pollution feature to the pollutant concentration.

[0162] The feature vector construction module 504 is used to construct feature vectors for each street and township based on the SHAP value of each pollution feature.

[0163] Clustering module 505 is used to obtain the pollution type of each street and town through cluster analysis based on the feature vectors of each street and town, and to obtain the feature matrix through matrix construction;

[0164] Analysis module 506 is used to generate a profile reflecting the pollution situation of each street and town based on cluster analysis results and feature matrix, and to perform air pollution analysis based on the profile.

[0165] In one embodiment, the SHAP value calculation module 503 is further configured to:

[0166] The initial SHAP value for each contamination feature is calculated based on the spatiotemporal weighting coefficient and the data credibility coefficient.

[0167] The initial SHAP value is corrected based on the pollutant concentration threshold to obtain the corrected SHAP value.

[0168] Based on pollutant concentration thresholds, pollutant concentration gradient weighting coefficients are set;

[0169] The corrected SHAP values ​​and pollutant concentration gradient weighting coefficients are integrated to obtain the final SHAP value for each pollution feature. This final SHAP value serves as the SHAP value for each pollution feature.

[0170] In one embodiment, the SHAP value calculation module 503 is further configured to:

[0171] The formula for calculating the initial SHAP value is as follows:

[0172] ,

[0173] in, Let be the initial SHAP value of the i-th pollution feature of the j-th street / township. Let j be the spatiotemporal weight coefficient of the j-th street / township. Let be the data reliability coefficient for the j-th street / township. For the multi-source fusion weight of the j-th street / township, The real-time concentration change when the feature subset S of the j-th street / township is added to the i-th pollution feature. Let S be any subset of features that does not contain the i-th contamination feature. Traditional subset weights;

[0174] The expression for the corrected SHAP value is:

[0175] ,

[0176] in, C is the corrected SHAP value of the i-th pollution characteristic of the j-th street / township. PM2.5 To measure the concentration of pollutant PM2.5, C t-PM2.5 C represents the threshold concentration of pollutant PM2.5. TSP To measure the concentration of pollutant TSP, C t-TSP The threshold for TSP concentration;

[0177] The pollutant concentration gradient weighting coefficients include the concentration gradient weighting coefficients for PM2.5 and TSP, with the following expressions:

[0178] ,

[0179] ,

[0180] Among them, f g-PM2.5,j f represents the concentration gradient weighting coefficient of PM2.5 in the j-th street / township. g-TSP,j Let C be the concentration gradient weighting coefficient of TSP in the j-th street / township. PM2.5,j C represents the measured PM2.5 concentration in the j-th street / township. TSP,j Let C be the measured TSP concentration of the j-th street / township. t-PM2.5,j Let C be the PM2.5 concentration threshold for the j-th street / township. t-TSP,j Let be the TSP concentration threshold for the j-th street / township.

[0181] In one embodiment, the SHAP value calculation module 503 is further configured to:

[0182] Each pollution feature of a street or township is grouped according to its pollution source. The adjusted SHAP values ​​after grouping are then integrated with the pollutant concentration gradient weighting coefficients to obtain the final SHAP value for each pollution feature. The expression for the final SHAP value is as follows:

[0183] ,

[0184] in, Let be the final SHAP value of the i-th pollution characteristic of the j-th street / township. G represents the final SHAP value of the g-th feature group of the j-th street / township. j Let G be the set of feature groups for the j-th street / township, and g be the set of feature groups. j The g-th feature group in For the spatiotemporal weight coefficient of the g-th feature group in the j-th street / township, Let g be the data reliability coefficient of the g-th feature group in the j-th street or township.

[0185] In one embodiment, the feature vector construction module 504 is further configured to:

[0186] The pollution sources include industrial sources, dust sources, mobile sources, residential sources, and natural sources. The construction of feature vectors for each street and township includes:

[0187] Construct the initial feature vector for each street and town. The expression for the initial feature vector is:

[0188] ,

[0189] in, Let be the initial feature vector of the j-th street / township. Let be the final SHAP value of the industrial source of the j-th street / township. Let be the final SHAP value of the dust source in the j-th street / township. Let be the final SHAP value of the mobile source in the j-th street / township. Let be the final SHAP value of the livelihood source of the j-th street / township. Let be the final SHAP value of the meteorological data for the j-th street / township. Let be the final SHAP value of the natural source of the j-th street / township;

[0190] The initial feature vectors are standardized to obtain the feature vectors for each street and township. The expression for the feature vector is as follows:

[0191] ,

[0192] in, Let j be the feature vector of the j-th street / township. This represents the average pollution characteristics of all streets and towns. The standard deviation is denoted as .

[0193] In one embodiment, the clustering module 505 is further configured to:

[0194] The cluster analysis employs the K-means algorithm, and the pollution types for each street and town are obtained through cluster analysis, including:

[0195] Based on the elbow rule and the pollutant concentration gradient verification preset rules, the intra-cluster sum of squares is calculated for the feature vectors of all streets and towns, and the optimal number of clusters K is determined based on the intra-cluster sum of squares.

[0196] Calculate the weighted Euclidean distance between all streets and towns;

[0197] Initialize K centroids and assign each street and town to the nearest cluster based on weighted Euclidean distance;

[0198] Recalculate the weighted mean of the SHAP values ​​of streets and towns within each cluster according to the feature group, and use the weighted mean as the new centroid;

[0199] Repeat the centroid update process until the change in centroid is less than a preset threshold or the maximum number of iterations is reached;

[0200] When the inter-cluster differences in pollutant concentration do not meet the preset rules for pollutant concentration gradient verification, adjust the K value and repeat the centroid update process.

[0201] When the inter-cluster differences in pollutant concentrations meet the preset rules for pollutant concentration gradient verification, the groups of streets and townships are obtained;

[0202] The pollution type of each group is determined based on the grouping of streets and townships.

[0203] In one embodiment, the apparatus further includes:

[0204] The determination module is used to determine whether streets and towns belong to the same pollution type when the determination criteria of weighted contribution, consistency of pollutant concentration and spatiotemporal correlation are met simultaneously.

[0205] The criteria for determining the weighted contribution are as follows:

[0206] ,and

[0207] The criteria for determining the consistency of pollutant concentrations are as follows:

[0208] ,and

[0209] The criteria for determining the spatiotemporal correlation are as follows:

[0210] ,and ,

[0211] Among them, W i W represents the weight of pollution feature of type i. p The weights of the pollution characteristics of class p, For the j-th street / township, μ represents the final SHAP value of the p-th pollution characteristic. PM2.5,k Let μ be the average concentration of PM2.5 within the k-th cluster. TSP,k Let W be the mean concentration of TSP within the k-th cluster, r be the correlation coefficient measuring the spatiotemporal weights of the j-th street and the k-th street, and W be the mean concentration of TSP within the k-th cluster. st,k W represents the spatiotemporal weighting coefficient of the k-th street / township. st,j Let j be the spatiotemporal weight coefficient of the j-th street / township. Let be the average spatiotemporal weight of all streets and towns within the urban area.

[0212] The embodiments of this invention achieve the following technical effects: This invention utilizes SHAP (Shapley Additive ex Planations) values ​​as an extraction method for constructing pollution feature vectors, solving the problem of difficulty in quantifying the evaluation weights of air pollution-related parameters in streets and towns, thus enhancing model interpretability; it comprehensively utilizes environmental monitoring data, meteorological data, and pollution emission data, overcoming the limitations of traditional methods that primarily rely on human experience for assessment, and improving assessment accuracy; based on air pollution feature profiles, precise source tracing can be performed, and the causes of pollution can be analyzed in conjunction with meteorological data; furthermore, through clustering, it is possible to classify pollution categories in streets and towns, as well as identify pollution transmission channels, promoting collaborative pollution control across streets and towns.

[0213] Obviously, those skilled in the art should understand that the modules or steps of the above-described embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.

[0214] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for analyzing air pollution based on big data profiling, characterized in that, include: Acquire multi-source data from streets and townships, including basic information data of streets and townships, pollutant concentration data, meteorological data, and data related to pollution source emissions; A pollutant concentration analysis model is constructed based on the multi-source data. The target variable of the pollutant concentration analysis model is the pollutant concentration, and the pollution characteristics of the pollutant concentration analysis model are meteorological data and pollution source emission-related data. Based on the pollutant concentration analysis model, the SHAP value of each pollution feature is obtained, and the SHAP value is used to reflect the degree of contribution of each pollution feature to the pollutant concentration. The process of obtaining the SHAP value for each pollution feature includes: calculating an initial SHAP value for each pollution feature based on a spatiotemporal weighting coefficient and a data reliability coefficient; correcting the initial SHAP value based on a pollutant concentration threshold to obtain a corrected SHAP value; setting a pollutant concentration gradient weighting coefficient based on a pollutant concentration threshold; and integrating the corrected SHAP value and the pollutant concentration gradient weighting coefficient to obtain a final SHAP value for each pollution feature, which serves as the SHAP value for each pollution feature. The formula for calculating the initial SHAP value is as follows: , in, Let be the initial SHAP value of the i-th pollution feature of the j-th street / township. Let j be the spatiotemporal weight coefficient of the j-th street / township. Let be the data reliability coefficient for the j-th street / township. For the multi-source fusion weight of the j-th street and township, The real-time concentration change when the feature subset S of the j-th street / township is added to the i-th pollution feature. Let S be any subset of features that does not contain the i-th contamination feature. Traditional subset weights; The expression for the corrected SHAP value is: , in, C is the corrected SHAP value of the i-th pollution characteristic of the j-th street / township. PM2.5 To measure the concentration of pollutant PM2.5, C t-PM2.5 C represents the threshold concentration of pollutant PM2.

5. TSP To measure the concentration of pollutant TSP, C t-TSP The threshold for TSP concentration; The pollutant concentration gradient weighting coefficients include the concentration gradient weighting coefficients for PM2.5 and TSP, with the following expressions: , , Among them, f g-PM2.5,j f represents the concentration gradient weighting coefficient of PM2.5 in the j-th street / township. g-TSP,j Let C be the concentration gradient weighting coefficient of TSP in the j-th street / township. PM2.5,j C represents the measured PM2.5 concentration in the j-th street / township. TSP,j Let C be the measured TSP concentration of the j-th street / township. t-PM2.5,j Let C be the PM2.5 concentration threshold for the j-th street / township. t-TSP,j Let TSP be the threshold concentration of pollutants in the j-th street or township. Based on the SHAP value of each pollution feature, construct feature vectors for each street and township; Based on the feature vectors of each street and township, the pollution type of each street and township is obtained through cluster analysis, and the feature matrix is ​​obtained through matrix construction. Based on the cluster analysis results and feature matrix, a profile reflecting the pollution situation of each street and town is generated, and air pollution analysis is performed based on the profile.

2. The air pollution analysis method based on big data profiling as described in claim 1, characterized in that, The process of integrating the corrected SHAP value and the pollutant concentration gradient weighting coefficients to obtain the final SHAP value for each pollution feature includes: Each pollution feature of a street or township is grouped according to its pollution source. The adjusted SHAP values ​​after grouping are then integrated with the pollutant concentration gradient weighting coefficients to obtain the final SHAP value for each pollution feature. The expression for the final SHAP value is as follows: , in, Let be the final SHAP value of the i-th pollution characteristic of the j-th street / township. G represents the final SHAP value of the g-th feature group of the j-th street / township. j Let G be the set of feature groups for the j-th street / township, and g be the set of feature groups. j The g-th feature group in For the spatiotemporal weight coefficient of the g-th feature group in the j-th street / township, Let g be the data reliability coefficient of the g-th feature group in the j-th street or township.

3. The air pollution analysis method based on big data profiling as described in claim 2, characterized in that, The pollution sources include industrial sources, dust sources, mobile sources, residential sources, and natural sources. The construction of feature vectors for each street and township includes: Construct the initial feature vector for each street and town. The expression for the initial feature vector is: , in, Let be the initial feature vector of the j-th street / township. Let be the final SHAP value of the industrial source of the j-th street / township. Let be the final SHAP value of the dust source in the j-th street / township. Let be the final SHAP value of the mobile source in the j-th street / township. Let be the final SHAP value of the livelihood source of the j-th street / township. Let be the final SHAP value of the meteorological data for the j-th street / township. Let be the final SHAP value of the natural source of the j-th street / township; The initial feature vectors are standardized to obtain the feature vectors for each street and township. The expression for the feature vector is as follows: , in, Let j be the feature vector of the j-th street / township. This represents the average pollution characteristics of all streets and towns. The standard deviation is denoted as .

4. The air pollution analysis method based on big data profiling as described in claim 3, characterized in that, The cluster analysis employs the K-means algorithm, and the pollution types for each street and town are obtained through cluster analysis, including: Based on the elbow rule and the pollutant concentration gradient verification preset rules, the intra-cluster sum of squares is calculated for the feature vectors of all streets and towns, and the optimal number of clusters K is determined based on the intra-cluster sum of squares. Calculate the weighted Euclidean distance between all streets and towns; Initialize K centroids and assign each street and town to the nearest cluster based on weighted Euclidean distance; Recalculate the weighted mean of the SHAP values ​​of streets and towns within each cluster according to the feature group, and use the weighted mean as the new centroid; Repeat the centroid update process until the change in centroid is less than a preset threshold or the maximum number of iterations is reached; When the inter-cluster differences in pollutant concentration do not meet the preset rules for pollutant concentration gradient verification, adjust the K value and repeat the centroid update process. When the inter-cluster differences in pollutant concentrations meet the preset rules for pollutant concentration gradient verification, the groups of streets and townships are obtained; The pollution type of each group is determined based on the grouping of streets and townships.

5. The air pollution analysis method based on big data profiling as described in claim 4, characterized in that, The method further includes: When the criteria for determining weighted contribution, consistency of pollutant concentration, and spatiotemporal correlation are met simultaneously, the streets and townships are determined to belong to the same pollution type. The criteria for determining the weighted contribution are as follows: ,and The criteria for determining the consistency of pollutant concentrations are as follows: ,and The criteria for determining the spatiotemporal correlation are as follows: ,and , Among them, W i W represents the weight of pollution feature of type i. p The weights of the pollution characteristics of class p, The final SHAP value of the p-th pollution characteristic of the j-th street / township, μ PM2.5,k Let μ be the average concentration of PM2.5 within the k-th cluster. TSP,k Let W be the mean concentration of TSP within the k-th cluster, r be the correlation coefficient measuring the spatiotemporal weights of the j-th street and the k-th street, and W be the mean concentration of TSP within the k-th cluster. st,k W represents the spatiotemporal weighting coefficient of the k-th street / township. st,j Let j be the spatiotemporal weight coefficient of the j-th street / township. This represents the average spatiotemporal weights of all streets and towns within the urban area.

6. An air pollution analysis device based on big data profiling, characterized in that, include: The data acquisition module is used to acquire multi-source data from streets and townships, including basic information data of streets and townships, pollutant concentration data, meteorological data, and data related to pollution source emissions. The model building module is used to build a pollutant concentration analysis model based on the multi-source data. The target variable of the pollutant concentration analysis model is the pollutant concentration, and the pollution characteristics of the pollutant concentration analysis model are meteorological data and pollution source emission-related data. The SHAP value calculation module is used to obtain the SHAP value of each pollution feature based on the pollutant concentration analysis model. The SHAP value is used to reflect the degree of contribution of each pollution feature to the pollutant concentration. The process of obtaining the SHAP value for each pollution feature includes: calculating an initial SHAP value for each pollution feature based on a spatiotemporal weighting coefficient and a data reliability coefficient; correcting the initial SHAP value based on a pollutant concentration threshold to obtain a corrected SHAP value; setting a pollutant concentration gradient weighting coefficient based on a pollutant concentration threshold; and integrating the corrected SHAP value and the pollutant concentration gradient weighting coefficient to obtain a final SHAP value for each pollution feature, which serves as the SHAP value for each pollution feature. The formula for calculating the initial SHAP value is as follows: , in, Let be the initial SHAP value of the i-th pollution feature of the j-th street / township. Let j be the spatiotemporal weight coefficient of the j-th street / township. Let be the data reliability coefficient for the j-th street / township. For the multi-source fusion weight of the j-th street and township, The real-time concentration change when the feature subset S of the j-th street / township is added to the i-th pollution feature. Let S be any subset of features that does not contain the i-th contamination feature. Traditional subset weights; The expression for the corrected SHAP value is: , in, C is the corrected SHAP value of the i-th pollution characteristic of the j-th street / township. PM2.5 To measure the concentration of pollutant PM2.5, C t-PM2.5 C represents the threshold concentration of pollutant PM2.

5. TSP To measure the concentration of pollutant TSP, C t-TSP The threshold for TSP concentration; The pollutant concentration gradient weighting coefficients include the concentration gradient weighting coefficients for PM2.5 and TSP, with the following expressions: , , Among them, f g-PM2.5,j f represents the concentration gradient weighting coefficient of PM2.5 in the j-th street / township. g-TSP,j Let C be the concentration gradient weighting coefficient of TSP in the j-th street / township. PM2.5,j C represents the measured PM2.5 concentration in the j-th street / township. TSP,j Let C be the measured TSP concentration of the j-th street / township. t-PM2.5,j Let C be the PM2.5 concentration threshold for the j-th street / township. t-TSP,j Let TSP be the threshold concentration of pollutants in the j-th street or township. The feature vector construction module is used to construct feature vectors for each street and township based on the SHAP value of each pollution feature; The clustering module is used to obtain the pollution type of each street and town through cluster analysis based on the feature vectors of each street and town, and to obtain the feature matrix through matrix construction. The analysis module is used to generate a profile reflecting the pollution situation of each street and town based on the cluster analysis results and feature matrix, and to perform air pollution analysis based on the profile.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the air pollution analysis method based on big data profiling as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that performs the atmospheric pollution analysis method based on big data profiling as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Urban street scale atmospheric pollution assessment method and device and storage medium

    CN118133640A

  • Surface water multi-point source pollution source tracing method and system

    CN120217089A