Air quality evaluation method and system based on multi-source data
Through multi-source data collection and processing, combined with geographical information and meteorological data, accurate interpolation and prediction of air quality are achieved, contributions to pollution sources are quantitatively analyzed, health risk warning is supported, and problems of insufficient space coverage and inaccurate pollution traceability in traditional methods are solved, and the scientificity and accuracy of air quality management are improved.
Patent Information
- Application Number
- CN202510425693.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing air quality monitoring methods rely on fixed monitoring stations, lack of space coverage, and it is difficult to take into account time and spatial resolution. The pollution traceability method lacks quantitative analysis and targetedness, and the early warning mechanism lacks spatial and temporal precision and population exposure risk considerations.
The air quality evaluation method based on multi-source data is adopted, and data from fixed air monitoring stations, micro-environment stations and portable sensors are collected, outlier value identification and correction processing is performed, and fine grids are divided into weighted calculations to generate air quality interpolation distribution maps, predict future pollutant concentrations, track pollutant transmission paths, determine pollution contribution values, and calculate health risk levels to generate hierarchical warning information.
It realizes accurate interpolation and future trend prediction of pollutant concentrations under uneven distribution of monitoring points, quantitative analysis of the contribution rate of pollution sources, supports accurate health risk warning and control decisions, and improves the scientificity and accuracy of air quality management.
Smart Images

Figure CN120182069A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to an air quality assessment method and system based on multi-source data. Background Art
[0002] Traditional air quality monitoring mainly relies on the construction of fixed monitoring station networks. By deploying a certain number of high-precision monitoring stations, regional air pollutant data can be collected. These monitoring stations are usually equipped with standardized sampling and analysis equipment, which can accurately measure the concentrations of conventional pollutants. Most existing air quality assessment methods use a single data source. For example, only the data of national control stations are used for time series analysis and trend prediction, or only satellite remote sensing data are used for large-scale spatial distribution assessment. In terms of pollution source tracing, traditional methods mainly rely on pollution source inventories and simple meteorological condition analysis, and it is difficult to accurately quantify the contribution rates of different source areas to the target area. The early warning mechanism is mostly based on the concentration thresholds of single pollutants, and adopts unified control measures for the whole city, lacking spatio-temporal fineness and pertinence.
[0003] However, there are obvious deficiencies in the existing technologies. First, the construction and operation and maintenance costs of fixed monitoring stations are high, and the number of stations is limited, resulting in insufficient spatial coverage and unable to fully reflect the spatial differences in air quality within the region. Second, the single data source analysis method is difficult to take into account both time and spatial resolutions at the same time, and the complex relationship between meteorological factors and pollutant diffusion and transmission is not fully considered. Third, traditional pollution source tracing methods mostly rely on empirical judgments, and it is difficult to quantitatively analyze the contribution rates of different pollution sources and accurately trace specific pollution events. Finally, the early warning mechanism lacks consideration of the population exposure risk, and the control measures are extensive, which may lead to insufficient prevention and control in high-risk areas and unnecessary social and economic losses in low-risk areas. These deficiencies seriously limit the scientific nature and accuracy of air quality management. Summary of the Invention
[0004] This application provides an air quality assessment method and system based on multi-source data, which is used to achieve accurate interpolation of pollutant concentrations, future trend prediction, and quantitative analysis of pollution source contribution rates in the case of uneven distribution of monitoring points, and further support accurate health risk early warning and control decisions.
[0005] In a first aspect, the present application provides an air quality assessment method based on multi-source data, and the air quality assessment method based on multi-source data includes: collecting pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors, and performing outlier identification and correction processing on the data to obtain a multi-source air quality original data set; according to the multi-source air quality original data set, the target area is divided into fine grids in urban areas and large-scale grids in suburbs, and terrain, land use, roads, traffic, population and pollution source information are combined to form an environmental feature grid database; the air quality data in the environmental feature grid database is weighted according to data accuracy, environmental similarity and spatial distance to generate an interpolation distribution map of air quality for the entire region; according to the interpolation distribution map of air quality for the entire region, combined with historical data for the same period and meteorological forecast information, the predicted value of pollutant concentration for future time periods is calculated to form a spatiotemporal dynamic prediction result; using the spatiotemporal dynamic prediction result, combined with wind field data and pollution source distribution, the pollutant transmission path is tracked, and the pollution contribution value of each source area to the target area is determined; according to the pollution contribution value, combined with population distribution and activity patterns, the health risk level of different regions is calculated, and graded warning information and control measures are recommended.
[0006] In a second aspect, the present application provides an air quality assessment system based on multi-source data, the air quality assessment system based on multi-source data comprising: The correction module is used to collect pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors, and perform outlier identification and correction processing on the data to obtain a multi-source air quality original data set; A partitioning module is used to divide the target area into fine grids in urban areas and large-scale grids in suburban areas according to the multi-source air quality original data set, and to form an environmental characteristic grid database by combining terrain, land use, roads, transportation, population and pollution source information; A weighting module is used to perform weighted calculation on the air quality data in the environmental feature grid database according to data accuracy, environmental similarity and spatial distance to generate an interpolation distribution map of air quality for the entire region; A prediction module is used to calculate the predicted value of pollutant concentration in the future period according to the interpolation distribution map of air quality in the whole area, combined with historical data of the same period and meteorological forecast information, to form a spatiotemporal dynamic prediction result; A tracking module is used to use the spatiotemporal dynamic prediction results, combined with wind field data and pollution source distribution, to track the pollutant transmission path and determine the pollution contribution value of various source areas to the target area; The control module is used to calculate the health risk level of different areas based on the pollution contribution value, combined with population distribution and activity patterns, and generate graded warning information and control measure recommendations.
[0007] A third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to cause the computer device to execute the above-mentioned air quality assessment method based on multi-source data.
[0008] A fourth aspect of the present invention provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is caused to execute the above-mentioned air quality assessment method based on multi-source data.
[0009] In the technical solution provided by this application, by collecting multi-source pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors, and performing outlier identification and correction processing, the limitations of the traditional single data source have been broken through, the integrity and reliability of the data have been significantly improved, and a solid data foundation has been laid for air quality assessment; the target area is divided into fine grids in urban areas and large-scale grids in suburbs based on terrain, land use, roads, transportation, population and pollution source information to form an environmental feature grid database, which realizes high-resolution spatial expression of environmental elements and fully considers the heterogeneous characteristics of the regional environment; an interpolation distribution map of air quality in the entire region is generated by weighted calculation according to data accuracy, environmental similarity and spatial distance, which solves the problem of uneven spatial distribution of monitoring points and realizes The pollutant concentration field was accurately reconstructed; based on the interpolation distribution map of air quality in the entire region, the predicted pollutant concentration values for future time periods were calculated in combination with historical data for the same period and meteorological forecast information, thus realizing the dynamic prediction of the spatiotemporal distribution of pollutants and providing a time window for early prevention and control; using the spatiotemporal dynamic prediction results, combined with wind field data and pollution source distribution, the pollutant transmission path was tracked, and the pollution contribution value of various source areas to the target area was determined, thus realizing the quantitative and scientific tracing of pollution sources and providing direction for precise pollution control; according to the pollution contribution value, combined with population distribution and activity patterns, the health risk levels of different regions were calculated, and graded warning information and control measures were generated, thus realizing a closed loop from air quality assessment to health risk management, and providing precise support for public health protection and pollution prevention and control decisions. It is particularly worth emphasizing that this solution applies artificial intelligence algorithms in multiple links. For example, the improved geographically weighted regression algorithm used in multi-source data spatial fusion interpolation can adaptively learn the complex nonlinear relationship between environmental characteristics and pollutant concentrations; the random forest algorithm used in the spatiotemporal dynamic prediction model can handle the high-dimensional feature interaction between meteorological factors and pollutant concentrations; the Lagrangian particle model and source contribution matrix algorithm used in pollution source analysis can accurately simulate the complex transmission process of pollutants under different meteorological conditions; the multi-factor risk model used in health risk assessment can comprehensively consider population density, activity patterns and distribution characteristics of sensitive populations. The application of these algorithms enables this solution to process a large amount of heterogeneous data, capture complex spatiotemporal patterns and causal relationships, and realize the intelligent and precise assessment and management of air quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0011] Figure 1A schematic diagram of an embodiment of an air quality assessment method based on multi-source data in an embodiment of the present application; Figure 2 This is a schematic diagram of an embodiment of an air quality assessment system based on multi-source data in an embodiment of the present application; Figure 3 It is a schematic block diagram of the structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The embodiment of the present application provides an air quality assessment method and system based on multi-source data. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0013] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiment of the present application, an embodiment of the air quality assessment method based on multi-source data includes: Step S101, collecting pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors, and performing outlier identification and correction processing on the data to obtain a multi-source air quality original data set; Step S102: Divide the target area into fine grids in urban areas and large-scale grids in suburban areas according to the multi-source air quality original data set, and form an environmental characteristic grid database by combining terrain, land use, roads, transportation, population and pollution source information; Step S103: performing weighted calculation on the air quality data in the environmental feature grid database according to data accuracy, environmental similarity and spatial distance to generate an interpolated distribution map of air quality for the entire region; Step S104: Calculate the predicted pollutant concentration value for the future period based on the interpolated distribution map of air quality in the whole area, combined with historical data for the same period and weather forecast information, to form a spatiotemporal dynamic prediction result; Step S105: Using the spatiotemporal dynamic prediction results, combined with wind field data and pollution source distribution, trace the pollutant transmission path and determine the pollution contribution value of various source areas to the target area; Step S106: Calculate the health risk levels of different regions based on the pollution contribution values, combined with the population distribution and activity patterns, and generate hierarchical warning information and control measure suggestions.
[0014] It can be understood that the execution entity of this application can be an air quality assessment system based on multi-source data, or a terminal or a server. Specifically, it is not limited here. In this embodiment of the application, the server is taken as an example of the execution entity for illustration.
[0015] Specifically, as a high-precision device, the fixed air monitoring station usually collects pollutant indicators once an hour, including data on six major pollutants such as fine particulate matter, inhalable particulate matter, sulfur dioxide, nitrogen dioxide, ozone, and carbon monoxide. The micro environmental station, as a medium-precision device, collects data every thirty minutes, while the portable sensor, as a low-precision device, collects data every ten minutes. For all the collected data, outlier identification processing is carried out. By calculating the deviation of the data from the seasonal historical average, when the data at a certain monitoring point exceeds three times the standard deviation of the seasonal historical average, it is marked as an outlier in the time dimension. At the same time, calculate the percentage deviation of the average value of each monitoring point from the average value of the adjacent five monitoring points. When the deviation exceeds 75%, it is marked as an outlier data point in the space dimension. Remove the data points that are marked as outliers in both the time dimension and the space dimension, and retain the remaining data points. Subsequently, cross-validation is carried out on different precision devices with adjacent distributions to establish the data conversion relationship between different precision devices. For example, the fine particulate matter concentration measured by the fixed monitoring station in the northern region of a certain city is 75 micrograms per cubic meter, while the portable sensor in the same region measures it as 85 micrograms per cubic meter. Through multiple comparisons, a correction coefficient of 0.88 is formed.
[0016] Divide the target area grid according to the multi-source air quality original data set. Due to the dense population and complex pollution sources in the urban area, a fine grid unit of 500 meters by 500 meters is adopted, while in the suburban area, a large-scale grid unit of 2000 meters by 2000 meters is adopted. Extract digital elevation data from the geographic information system, calculate the average altitude and terrain undulation degree of each grid unit to form terrain feature data. At the same time, extract the dominant land type and type composition ratio of each grid from the land use classification data to generate land use feature data. The road feature data is obtained by calculating the total road length and the number of intersections within each grid unit. The traffic feature data records the traffic flow and congestion degree of each grid. The population feature data includes the residential population density and floating population density of each grid. The pollution source feature data marks the number and emission scale of industrial enterprises within the grid. These feature data in different dimensions are associated with the grid through spatial overlay analysis to form an environmental feature grid database.
[0017] Air quality interpolation calculation is carried out based on the environmental feature grid database. First, different weights are assigned to monitoring devices according to precision: the data weight of fixed air monitoring stations is 1.0, that of micro environmental stations is 0.6, and that of portable sensors is 0.3. Then, the environmental similarity between the grid to be interpolated and known monitoring points is calculated. The environmental feature vectors are compared by the cosine similarity method to obtain the environmental similarity weight. The actual accessible distance is calculated based on road network data to form the spatial distance weight. For each grid to be interpolated, the data of monitoring points within a range of 5,000 meters are extracted. If the grid is located downwind of the monitoring point, the weight of this monitoring point is enhanced. The accuracy weight, environmental similarity weight, and spatial distance weight are normalized and then multiplied to obtain the comprehensive weight coefficient, which is used for the weighted average calculation of monitoring point data to generate the air quality interpolation distribution map of the entire region.
[0018] The air quality interpolation distribution map of the entire region is used for prediction. First, historical data for the same period in the past three years are extracted to construct a historical pollutant concentration database for the same period. Meteorological forecast data for the next 72 hours are obtained, including wind direction, wind speed, temperature, humidity, air pressure, and precipitation probability. The existing pollutant concentration data are decomposed by time series to separate the long-term trend and periodic change patterns. The seasonal change law is determined based on historical data for the same period to generate a basic prediction curve. The basic prediction is adjusted using the meteorological forecast data, and the influence of the pollution status in the upwind area is considered to generate a spatio-temporal dynamic prediction result. Based on the spatio-temporal dynamic prediction result, the pollutant transmission path is traced. First, the 72-hour wind field data of the target area are obtained. The wind direction is divided into 16 sectors, and the wind speed is divided into 6 levels to construct a wind field data grid. The location information and emission intensity of the pollution source area are obtained from the environmental protection department, and the Lagrangian particle method is used for backward tracking to generate the pollutant transmission path map. The influence frequency of the source area is statistically analyzed through spatial overlay analysis. Combining the emission intensity and distance attenuation, the pollution contribution value of various source areas is determined.
[0019] The pollution contribution value is used to generate health risk assessment and early warning. First, the population distribution data are obtained to construct a population distribution map. The activity patterns of urban residents are extracted from social survey data to mark the activity characteristics of people at different times. The data of patients with respiratory diseases are obtained from medical institutions to generate a distribution map of sensitive populations. Combining the pollution contribution value, population distribution, activity characteristics, and sensitive population distribution, the population exposure amount of each grid cell at different times is calculated, and the health risk level is set accordingly to generate targeted early warning information and control measure suggestions.
[0020] In the embodiments of the present application, by collecting multi-source pollutant concentration data from fixed air monitoring stations, micro environmental stations and portable sensors, and performing outlier identification and correction processing, the limitations of traditional single data sources are broken through, significantly improving the integrity and reliability of the data, laying a solid data foundation for air quality assessment; combining topographic, land use, road, traffic, population and pollution source information, the target area is divided into urban fine grids and suburban large-scale grids to form an environmental feature grid database, realizing the high-resolution spatial expression of environmental elements and fully considering the heterogeneous characteristics of the regional environment; by performing weighted calculations according to data accuracy, environmental similarity and spatial distance to generate an air quality interpolation distribution map for the entire region, the problem of uneven spatial distribution of monitoring points is solved, and the accurate reconstruction of the pollutant concentration field is realized; based on the air quality interpolation distribution map for the entire region, combined with historical data of the same period and meteorological forecast information, the predicted values of pollutant concentrations for future periods are calculated, realizing the dynamic prediction of the spatio-temporal distribution of pollutants and providing a time window for early prevention and control; using the spatio-temporal dynamic prediction results, combined with wind field data and pollution source distribution to trace the pollutant transmission path, determining the pollution contribution values of various source areas to the target area, realizing the quantification and scientificization of pollution source tracing, and providing a direction for precise pollution control; according to the pollution contribution values, combined with population distribution and activity patterns, the health risk levels of different regions are calculated, generating graded early warning information and control measure suggestions, realizing the closed-loop from air quality assessment to health risk management, and providing precise support for public health protection and pollution prevention and control decision-making. It is particularly worth emphasizing that artificial intelligence algorithms are applied in multiple links of this solution. For example, the improved geographically weighted regression algorithm applied in multi-source data spatial fusion interpolation can adaptively learn the complex non-linear relationship between environmental features and pollutant concentrations; the random forest algorithm applied in the spatio-temporal dynamic prediction model can handle the high-dimensional feature interaction between meteorological factors and pollutant concentrations; the Lagrangian particle model and source contribution matrix algorithm applied in pollution source tracing analysis can accurately simulate the complex transmission process of pollutants under different meteorological conditions; the multi-factor risk model applied in health risk assessment can comprehensively consider population density, activity patterns and the distribution characteristics of sensitive populations. The application of these algorithms enables this solution to process a large amount of heterogeneous data, capture complex spatio-temporal patterns and causal relationships, and realize the intelligence and precision of air quality assessment and management.
[0021] In a specific embodiment, the process of executing step S101 may specifically include the following steps: (1) Obtain fine particulate matter, inhalable particulate matter, sulfur dioxide, nitrogen dioxide, ozone, carbon monoxide atmospheric pollutant concentration data and equipment geographic coordinate information from monitoring devices of different accuracy levels; (2) Establish a monitoring device accuracy level system, classify fixed air monitoring stations as high-precision devices, micro environmental stations as medium-precision devices, and portable sensors as low-precision devices; (3) Calculate the deviation value of the collected pollutant concentration data from the seasonal historical average. When the deviation value exceeds three standard deviations, it is marked as an abnormal data point in the time dimension; (4) Calculate the percentage deviation of each monitoring point from the average of the adjacent five monitoring points. When the percentage exceeds 75%, it is marked as an abnormal data point in the spatial dimension; (5) Remove the data points that are marked as abnormal in both the time dimension and the spatial dimension, and retain the remaining data points; (6) Cross-validate the data of different precision-level devices arranged adjacent to each other, and establish a data conversion equation between devices of different precisions; (7) Fill in the missing data in the time series through linear interpolation, and fill in the missing data in the spatial distribution through inverse distance weighted interpolation to form a multi-source air quality original data set.
[0022] Specifically, obtain the concentration data of various pollutants from monitoring devices of different precision levels. Fixed air monitoring stations are built and maintained by environmental protection departments, equipped with standardized sampling and analysis equipment, and can accurately measure the concentrations of six major air pollutants, namely fine particulate matter, inhalable particulate matter, sulfur dioxide, nitrogen dioxide, ozone, and carbon monoxide, and record the precise geographical coordinates where the devices are located. Miniature environmental stations are usually deployed by district- and county-level environmental protection departments or research institutions. They are small in size and moderate in cost, and can measure major pollutants but with slightly lower precision than fixed stations. Portable sensors are deployed by communities, schools, or citizen science projects. They are low in cost but limited in precision, and mainly measure common pollutants such as fine particulate matter. These three types of devices form a multi-level monitoring network with a coverage range from point to area and a precision level from high to low. Establishing a monitoring device precision level system is to rationally utilize data from different sources. Fixed air monitoring stations, as high-precision devices, collect complete data once an hour, and their measurement results serve as reference standards; miniature environmental stations are classified as medium-precision devices and collect data once every 30 minutes; portable sensors are classified as low-precision devices and collect data once every 10 minutes. The classification of precision levels is based on the calibration frequency, instrument precision, and maintenance level of the devices, which is directly related to the weight allocation in subsequent data processing.
[0023] Identifying outliers in the collected pollutant concentration data is a crucial step in ensuring data quality. First, calculate the deviation from the seasonal historical average, which is the difference between the current measurement and the average of the same month and time period in the past three years, and then compare it with the standard deviation. The standard deviation reflects the fluctuation range of historical data. When the deviation exceeds three times the standard deviation, it indicates that the current value is far from the normal fluctuation range, and at this time, this data point is marked as an outlier in the time dimension. For example, the PM2.5 concentration at a monitoring station at 14:00 on June 15 is 120 micrograms per cubic meter, while the average value of the same station at 14:00 on June 15 in the past three years is 35 micrograms per cubic meter, and the standard deviation is 15 micrograms per cubic meter. Then the deviation value is 85, exceeding three times the standard deviation of 45, so it is marked as an outlier in the time dimension.
[0024] The identification of outliers in the spatial dimension is achieved by calculating the percentage deviation of each monitoring point from the average of the five adjacent monitoring points. Adjacent monitoring points refer to the five monitoring points with the closest spatial distance, which are determined by calculating the Euclidean distance between the monitoring points. When calculating the percentage deviation, first find the average value of the measurement values of the five adjacent monitoring points at the same moment, and then calculate the percentage of the difference between the current monitoring point and this average value to the average value. When the percentage exceeds 75%, it means that the data of the current monitoring point is significantly different from the surrounding environment, and at this time, it is marked as an outlier in the spatial dimension.
[0025] Removing the data points that are simultaneously marked as outliers in the time dimension and the spatial dimension is a double-verification mechanism. Only when it is determined to be abnormal in both the time and spatial dimensions will the data be removed. This can not only eliminate obviously incorrect data but also retain special but real data points such as local pollution events. The retained data points include normal data points, data points that are only abnormal in the time dimension, and data points that are only abnormal in the spatial dimension, which together form the basis for subsequent processing.
[0026] Cross-verifying the data of adjacent devices with different accuracy levels is an important means to solve the systematic errors of different devices. By analyzing the measurement results of different-precision devices deployed adjacent to each other (usually within 100 meters) at the same time period, a linear regression equation is established as the data conversion equation. The form of the conversion equation is Y = aX + b, where Y represents the data of the high-precision device, X represents the data of the low-precision device, and a and b are regression coefficients. It can calibrate the data of the low-precision device to the level equivalent to that of the high-precision device.
[0027] For the missing data in the time series, the linear interpolation method is adopted, that is, assuming a linear change between the two known data points before and after the missing point, and calculating the missing value through the weighted average of the front and back data points. For the missing data in the spatial distribution, the inverse distance weighted interpolation method is adopted, that is, according to the distance from the known points to the point to be interpolated, the weighting coefficient is calculated, and the closer the distance, the greater the weight. Through these two interpolation methods, various types of missing data are filled to form a multi-source air quality original data set.
[0028] In a specific embodiment, the process of performing step S102 may specifically include the following steps: (1) Divide the urban built-up area in the target region into urban fine grid units of 500 meters by 500 meters, and divide the suburban area in the target region into suburban large-scale grid units of 2000 meters by 2000 meters; (2) Extract digital elevation data from the geographic information system, calculate the average altitude and terrain undulation degree of each grid unit, and obtain terrain feature data; (3) Extract the dominant land type and land type composition ratio of each grid unit from the land use classification data to form land use feature data; (4) Obtain road network data from the traffic management department, calculate the road length density and the number of road intersections of each grid unit, and generate road feature data; (5) Obtain traffic data from the traffic flow monitoring system, calculate the average traffic flow and traffic congestion index of each grid unit, and obtain traffic feature data; (6) Extract population distribution information from the census data, calculate the permanent population density and floating population density of each grid unit, and form population feature data; (7) Obtain the list of fixed pollution sources from the environmental protection department, mark the number of industrial enterprises and the emission level in each grid unit, and generate pollution source feature data; (8) Through spatial overlay analysis, associate the terrain feature data, land use feature data, road feature data, traffic feature data, population feature data and pollution source feature data with the grid units to form an environmental feature grid database.
[0029] Specifically, based on the urban planning data, identify the boundary of the urban built-up area, and divide the urban built-up area into urban fine grid units of 500 meters by 500 meters. This scale can not only reflect the spatial variation characteristics of pollutants in the urban area, but also match the spatial resolution of common traffic data and population data. For the suburban area, since the distribution of pollution sources is relatively sparse and changes slowly, large-scale grid units of 2000 meters by 2000 meters are adopted, which reduces the calculation amount and ensures the rationality of spatial expression. The grid division adopts the UTM projection coordinate system to ensure the accuracy of area calculation. Each grid unit is assigned a unique code, which consists of row and column numbers. For example, the urban grid C0512 represents the urban grid unit in the 5th row and the 12th column.
[0030] After the grid division is completed, extract the terrain feature data of each grid. Obtain the Digital Elevation Model (DEM) data from a Geographic Information System, usually with a resolution of 30 meters or 90 meters. For each grid cell, extract the elevation values of all DEM points within its range and calculate the arithmetic mean as the average elevation of the grid. The terrain undulation degree is obtained by calculating the standard deviation of the DEM points within the grid. The larger the standard deviation, the more drastic the terrain change. For example, a grid contains 100 DEM points with elevation values ranging from 45 meters to 75 meters. The calculated average elevation is 60 meters and the standard deviation is 8 meters. These two values together constitute the terrain feature data of the grid.
[0031] The extraction of land use feature data is based on the results of remote sensing image classification or urban planning data. Land use is usually classified into categories such as residential land, commercial land, industrial land, green space, water area, transportation land, etc. For each grid cell, calculate the area proportion of each land type occupied in the grid. The type with the largest proportion is defined as the dominant land type of the grid. At the same time, calculate the composition proportion of each type to form a proportion vector. For example, the land use composition proportion of a grid cell is: residential land 40%, commercial land 30%, green space 15%, transportation land 15%. The dominant land type is residential land, and the land type composition proportion vector is [0.4, 0.3, 0, 0.15, 0, 0.15].
[0032] The road feature data is extracted from the road network data obtained from the traffic management department. The road network data contains the spatial positions and attribute information of roads at all levels, such as road grades, number of lanes, etc. The road length density refers to the total length of roads per square kilometer. The calculation method is to add up the lengths of all roads within the grid and then divide by the grid area. The number of road intersections is obtained by identifying the road intersection points. Intersection points usually mean points of traffic flow change and are closely related to pollutant emissions. For example, the area of a grid cell in an urban area is 0.25 square kilometers, the total length of various roads within it is 3 kilometers, and the number of intersection points is 8. Then its road length density is 12 kilometers per square kilometer, and the road intersection point density is 32 per square kilometer.
[0033] The traffic feature data needs to be obtained from the traffic flow monitoring system. Modern urban traffic monitoring systems collect road traffic flow information through intersection monitoring cameras, electronic police systems, or floating car data. For each grid cell, first identify the main roads inside, and then extract the average traffic flow data of these roads, usually expressed in PCU (Passenger Car Unit). The traffic congestion index is calculated by comparing the ratio of the actual driving speed to the free flow speed of the road. The closer this value is to 1, the smoother the traffic; the closer it is to 0, the more severe the congestion. For roads without direct monitoring data, the average value of similar road types and location conditions is used for estimation.
[0034] Demographic characteristic data is sourced from census or survey data. The permanent population density reflects the nocturnal population distribution and is calculated by dividing the number of permanent residents within a grid by the grid area. The floating population density, on the other hand, reflects the intensity of daytime population activities, and its data sources include indirect indicators such as mobile phone signaling data, bus card swiping data, or commercial POI density. Population density data is closely related to the assessment of the health impacts of air pollution as it reflects the scale of the population exposed to pollutants.
[0035] Pollution source characteristic data needs to obtain the list of fixed pollution sources from the environmental protection department. Fixed pollution sources mainly include point source emission facilities such as industrial enterprises, boiler houses, and waste treatment plants. For each grid cell, count the number of industrial enterprises within it and classify them into three levels: large emission sources, medium emission sources, and small emission sources according to the emissions of the enterprises. The emission level is determined based on the annual total pollutant emissions of the enterprise. For example, an industrial enterprise with an annual emissions of more than 100 tons of a certain type of air pollutant belongs to a large emission source. These data are directly related to the local pollution contribution of the grid.
[0036] By performing spatial overlay analysis, the above-mentioned various types of characteristic data are associated with grid cells to form an environmental characteristic grid database. Spatial overlay analysis is a basic function of geographic information systems, which integrates the attribute information of different layers into a unified spatial unit through spatial position relationships. Specific operations include point-in-polygon judgment, line-and-polygon intersection operations, and polygon-overlap analysis. The data structure of the environmental characteristic grid database includes grid identifiers and various associated characteristic attributes, forming a relational database with a "one grid, one record" relationship, which is convenient for subsequent query and analysis.
[0037] Taking a fine grid cell in the central area of a certain city as an example: The grid is numbered C0305, with an area of 0.25 square kilometers. The average altitude calculated from DEM data is 45 meters, and the terrain undulation degree is 2 meters. Land use analysis shows that the dominant type is commercial land, accounting for 60%, followed by residential land, accounting for 30%, and transportation land, accounting for 10%. Road network analysis shows that there are 2.8 kilometers of main roads and secondary roads in the area, 6 intersections, and the road density is 11.2 kilometers per square kilometer. Traffic monitoring data shows that the average daytime traffic flow on weekdays is 2000 PCU / hour, and the traffic congestion index is 0.6. Population data shows that there are 5000 permanent residents, the permanent population density is 20000 people per square kilometer, the incremental daytime floating population is 8000 people, and the floating population density is 32000 people per square kilometer. The pollution source list shows that there are 2 medium-sized industrial enterprises in the area, with an annual emission of 20 tons of fine particulate matter. These detailed characteristics are associated with grid C0305 through spatial overlay analysis, becoming a complete record in the environmental characteristic grid database, providing multi-dimensional environmental background information for subsequent air quality interpolation calculations.
[0038] In a specific embodiment, the process of executing step S103 may specifically include the following steps: (1) Assign accuracy weight coefficients to the data of each monitoring device, where the weight of the fixed air monitoring station data is 1, the weight of the micro environmental station data is 0.6, and the weight of the portable sensor data is 0.3; (2) Calculate the environmental similarity between the grid cell to be interpolated and the known monitoring points, and obtain the environmental similarity weight by calculating the vector cosine similarity of terrain, land use, roads, traffic, population, and pollution source characteristics; (3) Calculate the actual accessible distance between the grid cell to be interpolated and the monitoring points based on the road network data to form the spatial distance weight; (4) For each grid cell to be interpolated, extract the data of all monitoring points within a radius of five kilometers around it; (5) When the grid cell to be interpolated is located in the downwind position of the monitoring point, enhance the influence weight of the monitoring point according to the wind field data; (6) Multiply the accuracy weight coefficient, environmental similarity weight, and spatial distance weight after normalization to obtain the comprehensive weight coefficient; (7) Use the comprehensive weight coefficient to perform weighted average calculation on the pollutant concentration data of the monitoring points to obtain the pollutant concentration value of the grid cell to be interpolated, and calculate the interpolation uncertainty index; (8) Calculate the pollutant concentration values of each grid cell one by one according to the division of urban fine grids and suburban large-scale grids to form the spatial distribution data of pollutant concentrations; (9) Visualize the spatial distribution data of pollutant concentrations by the contour line method to generate the air quality interpolation distribution map of the whole region.
[0039] Specifically, accuracy weight coefficients are assigned to the data of each monitoring device, and this assignment is based on the measurement accuracy, calibration frequency, and maintenance level of the device. The fixed air monitoring station is maintained by professional technicians, uses instrument equipment that meets national standards, has the highest accuracy, and the weight is set to 1.0; the accuracy of the micro environmental station is the second, and the weight is set to 0.6; due to the low cost and infrequent calibration of the portable sensor, the accuracy is relatively low, and the weight is set to 0.3. This weight assignment ensures that high-precision devices play a greater role in the interpolation calculation.
[0040] For the calculation of the environmental similarity between the grid cell to be interpolated and the known monitoring points, the vector cosine similarity method is adopted. First, the terrain, land use, roads, traffic, population, and pollution source characteristics of the grid cell are used to form a feature vector , and similarly, the characteristics of the grid where the monitoring point is located are used to form a vector , and then the cosine similarity of the two vectors is calculated: ; Among them represents the environmental similarity of the grid and the monitoring points The value range is [0, 1]. The larger the value, the more similar the environment is. represents the total number of feature dimensions, which in this method includes dimensions such as average altitude, terrain undulation degree, proportion of various land uses, road density, traffic flow, population density, pollution source intensity, etc.
[0041] The actual accessible distance calculated based on road network data conforms more to the pollutant propagation law than the straight-line distance. Establish the road network topology structure, regard the road network as a weighted directed graph, where the weight of the edge is the road length. Then use the Dijkstra shortest path algorithm to calculate the shortest path length from the center point of the grid to be interpolated to the monitoring point as the actual accessible distance . Spatial distance weight is calculated through a distance attenuation function: ; where and are attenuation parameters, usually taking = 0.5, = 2, that is, inverse square attenuation is adopted.
[0042] For each grid cell to be interpolated, instead of considering all the monitoring point data, the monitoring point data within a radius of five kilometers around is extracted for calculation. The setting of this range is based on the balance between the diffusion characteristics of pollutants and calculation efficiency. Specifically, when implementing, calculate the straight-line distance between the center point of the grid and all monitoring points, and screen the point set with a distance less than five kilometers : ; where represents the Euclidean distance calculation function.
[0043] The wind direction has a significant impact on pollutant transmission. When the grid cell to be interpolated is downwind of the monitoring point, the pollutant concentration of this monitoring point has a greater impact on the grid. According to the wind field data, calculate the wind direction correction coefficient : ; where is the azimuth angle from the monitoring point to the grid , is the current wind direction angle, is the wind direction enhancement coefficient, with a value of 0.5. When the grid is directly downwind of the monitoring point, , the wind direction correction coefficient is at most 1.5; when the grid is directly upwind of the monitoring point, the correction coefficient is at least 0.5.
[0044] Multiply the precision weight coefficient, environmental similarity weight, spatial distance weight, and wind direction correction coefficient after normalization to obtain the comprehensive weight coefficient : ; Use the comprehensive weight coefficient to perform weighted average calculation on the pollutant concentration data of the monitoring points to obtain the pollutant concentration value of the grid cell to be interpolated : ; where represents the pollutant concentration value measured at the monitoring point . At the same time, calculate the interpolation uncertainty index , which reflects the credibility of the interpolation result: ; where represents the number of monitoring points participating in the calculation The value range is [0, 1], and the larger the value, the higher the uncertainty
[0045] According to the division of urban fine grids and suburban large-scale grids, calculate the pollutant concentration values of each grid cell one by one to form the pollutant concentration spatial distribution data matrix C. Finally, visualize the pollutant concentration spatial distribution data through the contour method to generate the air quality interpolation distribution map of the entire region. The contour method uses a Triangulated Irregular Network (TIN) to construct a continuous surface, and then extracts the contours of specific concentration values to form a visually clear pollutant concentration distribution map
[0046] For example, a certain city has a total of 8 fixed air monitoring stations, 15 micro environmental stations, and 50 portable sensors. The data collection time is 10 am. Interpolation calculation is performed on an urban grid unit located in the commercial area. First, the monitoring devices within a radius of 5 kilometers are extracted, including 2 fixed stations, 4 micro stations, and 12 portable sensors. When calculating the environmental similarity, the feature vector of this grid is [45, 2, 0.6, 0.3, 0.1, 0, 0, 0, 11.2, 6, 2000, 0.6, 20000, 32000, 2, 20], representing features such as an average altitude of 45 meters, a terrain undulation of 2 meters, and a commercial land occupancy ratio of 60%. The calculation result of the environmental similarity with the nearest fixed station is 0.92, indicating that the environmental conditions are close. When calculating the spatial distance weight, the actual road distance from this grid to the nearest fixed station is 2.3 kilometers, and the spatial distance weight is 0.32. The wind direction on that day is northwest, and this grid is located in the southeast direction of one of the fixed stations, in the downwind position, and the wind direction correction coefficient is 1.35. Through comprehensive calculation, the comprehensive weight of this fixed station for the grid is 0.28. Similar calculations are performed on all surrounding monitoring points and normalized. Through weighted average, the PM2.5 concentration of this grid is 85 micrograms per cubic meter, and the interpolation uncertainty index is 0.15, indicating that the result is relatively reliable. Through similar calculations for all grids in the entire region, a PM2.5 concentration distribution map covering the entire city is generated, clearly showing high concentration areas around the commercial center and main traffic arteries, as well as low concentration areas in the suburbs.
[0047] In a specific embodiment, the process of executing step S104 may specifically include the following steps: (1) Extract the interpolation distribution map of the air quality in the entire region for the same period in the past three years in the target area, and construct a historical pollutant concentration database for the same period; (2) Obtain data on wind direction, wind speed, temperature, humidity, air pressure, and precipitation probability for the next seventy-two hours from the meteorological department to form a meteorological forecast data set; (3) Perform time series decomposition on the interpolation distribution map of the air quality in the entire region to separate the long-term trend and periodic change pattern of the pollutant concentration; (4) Determine the seasonal change law of the pollutant concentration according to the historical pollutant concentration database for the same period, and generate a basic prediction curve; (5) Use the meteorological forecast data set to adjust the basic prediction curve to form a meteorological factor correction value; (6) Considering the influence of the pollution status in the upwind area, perform spatial transmission correction on the meteorological factor correction value to generate a spatio-temporal dynamic prediction result.
[0048] Specifically, extract the interpolated distribution maps of the air quality of the entire region for the same period in the past three years to construct a historical concurrent pollutant concentration database. Use the generated interpolated distribution maps of the air quality of the entire region as the basic data to trace back the concurrent data for three years in the time dimension. The interpolated distribution maps of the air quality of the entire region are generated by weighted calculation of the air quality data in the environmental characteristic grid database according to data accuracy, environmental similarity, and spatial distance. For example, when extracting the pollutant data for each July from 2022 to 2024 in a certain city, the concentration values of multiple pollutants in the historical concurrent period are extracted for each grid cell (500 meters × 500 meters in the urban area and 2 kilometers × 2 kilometers in the suburban area), forming a multi-dimensional database containing time stamps, spatial coordinates, and the concentrations of multiple pollutants.
[0049] Obtain the wind direction, wind speed, temperature, humidity, air pressure, and precipitation probability data for the next seventy-two hours from the meteorological department to form a meteorological forecast data set. These meteorological data are stored according to spatial grids, maintaining the same spatial index as the environmental characteristic grids. Each grid cell contains the meteorological forecast data for each hour within the next 72 hours. The wind direction data is represented by angular values (0 - 359 degrees), the wind speed is in meters per second, the temperature is in degrees Celsius, the humidity is in percentage, the air pressure is in hectopascals, and the precipitation probability is in percentage. These data are matched with the air quality data to provide input variables for the subsequent prediction model.
[0050] Perform time series decomposition on the interpolated distribution maps of the air quality of the entire region to separate the long-term trend and periodic change patterns of the pollutant concentration. The time series decomposition uses the classical additive model to decompose the original time series data into three parts: the trend term, the seasonal term, and the residual term. For the pollutant concentration data of each grid cell, the moving average method is used to extract the long-term trend, and the 24-hour (daily variation), 168-hour (weekly variation), and seasonal cycles are identified through Fourier transform. This decomposition can effectively identify the internal laws of the pollutant concentration changes and provide a basis for prediction.
[0051] According to the historical concurrent pollutant concentration database, determine the seasonal variation law of the pollutant concentration. Generating the basic prediction curve is to analyze the seasonal patterns in the historical concurrent data, calculate the daily average value, weekly average value, and standard deviation in the historical concurrent period, and construct a prediction curve based on the historical concurrent laws. The specific operation is to extract the hourly pollutant concentration data for the same period in the past three years for each grid cell, calculate the average value and the range of variation at the same time in the concurrent period, and form the basic prediction curve. This curve reflects the natural variation law of the pollutant concentration after excluding the influence of special weather.
[0052] Using a meteorological forecast dataset, the basic prediction curve is adjusted to form a meteorological factor correction value. This step is based on the correlation analysis between meteorological factors and pollutant concentrations, and a regression model is constructed to quantify the impact of meteorological conditions on pollutant concentrations. The model takes into account the impact of temperature on the rate of photochemical reactions, the impact of humidity on particulate matter deposition, the impact of wind speed on pollutant dispersion, and the impact of pressure systems on regional transport. Through multiple linear regression or random forest algorithms, the correction coefficients under future meteorological conditions are calculated to adjust the basic prediction curve.
[0053] Considering the impact of the pollution status in the upwind area, the spatial transmission correction of the meteorological factor correction value is carried out to generate a spatio-temporal dynamic prediction result. This step introduces the spatially lagged effect driven by the wind field. According to the wind direction and wind speed data, the pollutant transmission path and transmission time are determined, and the contribution of pollutants in the upwind area to the target grid cell is calculated. The specific implementation is to construct a spatial weight matrix to quantify the influence relationship between different grid cells. For example, when the dominant wind direction is northwest and the wind speed is 3 m / s, the pollutants in the grid cell located 10 km in the northwest direction will take about 1 hour to be transmitted to the target area. At this time, the current pollution concentration of this grid cell will be used as the pollution contribution factor for the target grid cell after 1 hour. Through the data processing of the above six steps, accurate prediction of the pollutant concentration of each grid cell within the next 72 hours can be achieved. For example, when applying this method, the PM2.5 prediction results in a certain city show that based on the analysis of historical data in the same period, the predicted value of the PM2.5 basic concentration in the first week of July is 35 μg / m³. Considering that the meteorological forecast for the next 72 hours shows a process of temperature drop and humidity increase, the meteorological factor correction reduces the predicted value to 28 μg / m³. Further considering the planned emission activities in the industrial area in the upwind direction, the spatial transmission correction increases the predicted value in a specific area to 42 μg / m³.
[0054] In a specific embodiment, the process of executing step S105 may specifically include the following steps: (1) Obtain the wind direction and wind speed data within seventy-two hours in the target area, divide the wind direction into sixteen sectors, divide the wind speed into six levels, and construct a wind field data grid; (2) Extract the location information, types of pollutants emitted, and their hourly average emission intensities of industrial parks, high-density traffic areas, heating areas, and surrounding city transmission areas from the pollution source list of the environmental protection department; (3) Based on the wind field data grid, use the Lagrangian particle method to regard each grid cell as a starting point, reverse-trace pollutant particles according to the wind direction, record the positions of pollutant particles every hour, and form a pollutant transmission path map for seventy-two hours; (4)Perform spatial overlay analysis on the pollutant transport path map, count the number of times each source area is crossed by particle paths, and divide by the total number of particles to obtain the source area influence frequency value; (5)Classify and label the source areas according to industrial sources, transportation sources, domestic sources, and regional transport sources, and sum up the source area influence frequency values within each type of source area to obtain the comprehensive influence frequency of each type of source area; (6)Multiply the comprehensive influence frequency of each type of source area by its emission intensity, and set an attenuation coefficient according to the distance of the source area from the target area to obtain the pollution contribution value of each type of source area to the target area.
[0055] Specifically, the implementation of multi-scenario pollution source tracing and contribution rate analysis begins with obtaining the wind direction and wind speed data within 72 hours of the target area. The wind direction is divided into 16 sectors, and the wind speed is divided into 6 levels to construct a wind field data grid. Specifically, the wind direction data obtained from the meteorological department is divided according to the angle value. The 360 degrees are evenly divided into 16 sectors, and each sector covers a range of 22.5 degrees. For example, the first sector is 348.75° - 11.25°, the second sector is 11.25° - 33.75°, and so on. The wind speed is divided into 6 levels according to the actual distribution: 0 - 1.5 m / s (calm or light breeze), 1.5 - 3.3 m / s (light wind), 3.3 - 5.4 m / s (gentle breeze), 5.4 - 7.9 m / s (fresh breeze), 7.9 - 10.7 m / s (strong wind), above 10.7 m / s (gale or storm). Each spatial grid unit has a corresponding wind direction sector number and wind speed level number at each time point, forming a four-dimensional wind field data grid (three spatial dimensions plus one time dimension).
[0056] Then, extract the location information, types of pollutants emitted, and their hourly average emission intensities from the pollution source list of the environmental protection department for industrial parks, high-density traffic areas, heating areas, and surrounding city transport areas. An industrial park refers to an area where industrial enterprises are concentrated. A high-density traffic area refers to traffic arteries and transportation hub areas with high traffic flow. A heating area refers to an area with centralized heating or decentralized coal or gas heating during the winter heating period. A surrounding city transport area refers to the built-up areas of cities in the direction where pollutants from surrounding cities may be transported to the target area. For each type of pollution source area, record its central coordinates and shape boundary (polygon or circle), as well as the main types of pollutants emitted and the hourly average emission intensity (in kilograms per hour).
[0057] Based on the wind field data grid, the Lagrangian particle method is used to regard each grid cell as a starting point, and pollutant particles are traced backward according to the wind direction. The positions of the pollutant particles are recorded every hour to form a 72-hour pollutant transport path map. The Lagrangian particle method is a numerical method for simulating the transport of atmospheric pollutants. Its basic principle is to release a large number of virtual particles in the air and trace the trajectories of these particles as they move with the airflow. In this method, for each grid cell, a certain number (such as 100) of virtual particles are released, and then the movement direction and speed of the particles are determined according to the wind field data. Since it is backward tracing, the movement direction of the particles is opposite to the wind direction, and the movement speed is proportional to the wind speed. The specific algorithm is as follows: for each particle, according to the wind direction and wind speed of the current grid, calculate its displacement vector per hour under the action of the wind field, then update the particle position and record this position. Repeat this process 72 times to record the complete trajectories of each particle within 72 hours and form a pollutant transport path map of the entire region.
[0058] Perform a spatial overlay analysis on the pollutant transport path map, count the number of times each source region is crossed by the particle paths, and divide by the total number of particles to obtain the influence frequency value of the source region. Spatial overlay analysis refers to overlaying the pollutant particle trajectory data with the spatial distribution data of the source regions and calculating the intersection situation of each particle trajectory with each source region. When a particle trajectory crosses a certain source region, the number of times the source region is crossed is incremented by 1. After the statistics are completed, divide the number of times each source region is crossed by the total number of particles (the number of grid cells multiplied by the number of particles released in each cell) to obtain the influence frequency value of the source region, which reflects the potential impact degree of the source region on the air quality of the target region.
[0059] Classify and label the source regions according to industrial sources, transportation sources, domestic sources, and regional transmission sources, and sum the influence frequency values of the source regions within each type of source area to obtain the comprehensive influence frequency of each type of source area. Industrial sources include various industrial parks and factories, transportation sources include main roads, highways, and transportation hubs, domestic sources include residential areas and heating areas, and regional transmission sources include surrounding cities and regions with long-distance transmission. Sum the influence frequency values of the same type of source regions to obtain the comprehensive influence frequency of this type of source area, which reflects the relative contribution degree of different types of pollution sources to the air quality of the target region.
[0060] Multiply the comprehensive influence frequency of each type of source area by its emission intensity, and set an attenuation coefficient according to the distance of the source area from the target region to obtain the pollution contribution value of each type of source area to the target region. Emission intensity refers to the pollutant emission amount of the source region, with the unit of kilograms per hour. The attenuation coefficient reflects the characteristic that the pollutant concentration decreases with the increase of the transmission distance, and can be set as the reciprocal of the distance or an exponential decay function. By considering the three factors of comprehensive influence frequency, emission intensity, and attenuation coefficient, calculate the pollution contribution value of each type of source area to the target region, providing a quantitative basis for scientific pollution control.
[0061] Taking the PM2.5 source apportionment of a certain city as an example, the data processing process of this method is described. First, obtain the wind direction and wind speed data within 72 hours in the city and its surrounding areas. After dividing them into 16 wind direction sectors and 6 wind speed levels, a wind field data grid is constructed. Extract the locations and emission data of 4 industrial parks, 3 high-density traffic areas, 2 large-scale residential heating areas, and 2 surrounding cities from the pollution source list obtained from the environmental protection department. The average PM2.5 emission intensity of each industrial park is 5 kg / hour, that of the high-density traffic area is 3 kg / hour, that of the heating area is 2 kg / hour, and that of the surrounding cities is 8 kg / hour. Using the Lagrangian particle method, 100 virtual particles are released at the city center point, and the pollutant sources are traced backward according to the 72-hour wind field data. The tracking results show that within 72 hours, 35 out of 100 particles passed through the northern industrial park, 25 passed through the eastern traffic artery, 15 passed through the southern heating area, 20 passed through the western surrounding cities, and the trajectories of 5 particles did not intersect with any known source areas. The calculated influence frequency of industrial sources is 0.35, that of traffic sources is 0.25, that of domestic sources is 0.15, and that of regional transmission sources is 0.20. Considering that the northern industrial park is 15 km away from the city center, the attenuation coefficient is set to 0.8; the eastern traffic artery is 10 km away, and the attenuation coefficient is 0.9; the southern heating area is 5 km away, and the attenuation coefficient is 0.95; the western surrounding cities are 30 km away, and the attenuation coefficient is 0.6. The calculated contribution rate of industrial sources is 0.35×5×0.8 = 1.4, that of traffic sources is 0.25×3×0.9 = 0.675, that of domestic sources is 0.15×2×0.95 = 0.285, and that of regional transmission sources is 0.20×8×0.6 = 0.96. According to the order of the contribution rates from large to small, the pollutant sources are industrial sources, regional transmission sources, traffic sources, and domestic sources.
[0062] In a specific embodiment, the process of executing step S106 may specifically include the following steps: (1) Obtain the permanent population density and floating population density of the target area from the census database, and construct a population distribution data map; (2) Extract the activity patterns of urban residents based on social survey data, divide a day into peak commuting periods, working hours, leisure hours, and night rest hours, and form a time period activity feature library; (3) Mark the residential areas of patients with respiratory diseases in the data of patients received from medical institutions, and generate a sensitive population distribution map; (4) Combine the pollution contribution value, population distribution data map, time period activity feature library, and sensitive population distribution map to calculate the population exposure amount of each grid cell at different time periods; (5)Set health risk thresholds according to the population exposure level, and divide the risk values into four health risk levels: low risk, medium risk, high risk, and extremely high risk. (6)Based on the health risk levels, generate classified early warning information and control measure suggestions, including the early warning period, affected area, protective measures, and control measure suggestions for the main pollution source areas.
[0063] Specifically, obtain the permanent population density and floating population density of the target area from the census database, and construct a population distribution data map. The permanent population density refers to the number of people permanently residing within a square kilometer, while the floating population density refers to the number of people who are not permanently residing but are frequently active within a square kilometer. Census data usually contains population statistics at the street or community level, and these data need to be further processed to the same grid scale as the air quality assessment. The specific processing method is to allocate the street or community-level population data to overlapping grid cells according to the area ratio, forming a population distribution data map with the same spatial resolution as the pollutant concentration distribution map. In the urban area, a fine grid of 500 meters × 500 meters is used, and each grid cell records the number of permanent residents and the number of floating population; in the suburban area, a larger grid of 2 kilometers × 2 kilometers is used, and the two types of population data are also recorded.
[0064] Extract the activity patterns of urban residents based on social survey data, divide a day into commuting peak periods, working hours, leisure hours, and night rest hours, and form a time period activity characteristic library. The commuting peak periods usually refer to 7:00 - 9:00 in the morning and 17:00 - 19:00 in the evening, and residents are mainly active on traffic arteries and public transportation hubs; the working hours refer to 9:00 - 17:00, and on weekdays, residents are mainly active in workplaces (office areas, industrial areas, commercial areas, etc.), and on non-working days, they are active in commercial areas and leisure places; the leisure hours refer to 19:00 - 23:00, and residents are mainly active in residential areas, commercial areas, and leisure places; the night rest hours refer to 23:00 - 7:00, and residents are mainly active in residential areas. Based on this time division, assign activity intensity coefficients to each grid cell at different time periods to reflect how many people in the grid are active during a specific time period. For example, the activity intensity in commercial areas is relatively high during working hours and leisure hours, and the activity intensity in residential areas is relatively high before and after the commuting peak periods and during the night rest hours.
[0065] For the data of patients with respiratory diseases received from medical institutions, mark the residential areas of the patients and generate a distribution map of sensitive populations. Patients with respiratory diseases include those with diseases such as asthma, chronic bronchitis, emphysema, pulmonary heart disease, and pulmonary fibrosis. These populations are particularly sensitive to air pollutants. The patient data provided by medical institutions contains information on the patient's residential address or affiliated community. Through geocoding technology, the address information is converted into geographical coordinates, and then aggregated into grid cells. Calculate the number of sensitive populations and their proportion in the total population in each grid cell. In addition to patients with respiratory diseases, the elderly over 65 years old and children under 14 years old are also marked as sensitive populations. The generated distribution map of sensitive populations is in units of grids, recording the number and proportion information of sensitive populations in each grid cell.
[0066] Combined with the pollution contribution value, the population distribution data map, the time period activity characteristic library, and the sensitive population distribution map, calculate the population exposure in each grid cell at different time periods. Population exposure is a comprehensive function of air pollutant concentration, population quantity, activity intensity, and sensitivity. For each grid cell at each time period, first obtain the predicted pollutant concentration value of the grid, then multiply it by the population quantity and the time period activity intensity coefficient of the grid, and then weight according to the proportion of sensitive populations to calculate the population exposure. The weighting factor for sensitive populations is usually set to 1.5 - 3 times, reflecting their higher sensitivity to pollutants. This calculation method takes into account the spatio-temporal dynamic characteristics of air quality, population distribution, and activity patterns, and can more accurately reflect the actual impact of pollutants on population health.
[0067] Set a health risk threshold based on the population exposure, and divide the risk values into four health risk levels: low risk, medium risk, high risk, and extremely high risk. The setting of the health risk threshold is based on epidemiological studies and health standards, taking into account the health impact characteristics of different pollutants and the exposure-response relationship curve. For example, for PM2.5, the risk threshold can be set according to the World Health Organization and national ambient air quality standards, combined with the characteristics of the local population. When the population exposure is lower than the first threshold, it is low risk; when it is higher than the first threshold but lower than the second threshold, it is medium risk; when it is higher than the second threshold but lower than the third threshold, it is high risk; when it is higher than the third threshold, it is extremely high risk. Different risk levels correspond to different degrees of severity of health impacts and levels of urgency for prevention and control.
[0068] The last step is to generate graded early warning information and control measure suggestions based on the health risk level, including the early warning period, the affected area, protective measures, and control measure suggestions for the main pollution source areas. The early warning period refers to the continuous time period during which the health risk exceeds a specific level; the affected area refers to the spatial area where the health risk exceeds a specific level; the protective measures are health protection suggestions for the general public and sensitive populations; the control measure suggestions are pollution reduction measures for the main pollution source areas. At the low risk level, mainly information release and protective suggestions for sensitive populations are provided; at the medium risk level, public protection suggestions and some pollution source control suggestions are added; at the high risk level, strict public protection measures and main pollution source emission limit measures are implemented; at the extremely high risk level, the emergency plan is activated and the strictest regional control measures are implemented.
[0069] Taking the early warning of heavy pollution weather in a certain city in winter as an example to illustrate the complete process of this method. First, obtain the population density data of each region of the city from the census database. The average density in the urban central area is 20,000 people per square kilometer, and in the far suburbs is 2,000 people per square kilometer. Based on the social survey data, determine that the activity intensity of the grid units on the main urban roads during the commuting peak hours (7:00 - 9:00 and 17:00 - 19:00) is 1.8, the activity intensity of the grid units in the commercial and industrial areas during the working hours (9:00 - 17:00) is 1.5, the activity intensity of the commercial and residential areas during the leisure hours (19:00 - 23:00) is 1.3, and the activity intensity of the residential areas during the night hours (23:00 - 7:00) is 1.0. The data obtained from medical institutions shows that about 8% of the population in the city belongs to patients with respiratory diseases, mainly concentrated around industrial areas and in the old city. Combining the pollution contribution values determined in the previous steps, it is found that the PM2.5 concentration in the northern region of the city will reach 150 μg / m³ within the next 48 hours, which exactly overlaps with the high-density residential areas and the concentrated areas of sensitive populations. It is calculated that during the commuting peak hours, the population exposure in a certain grid unit in the northern region reaches 2.5 times the risk assessment standard, belonging to the high risk level. Based on this, early warning information is generated: the early warning period is from 7:00 to 19:00 tomorrow, the affected area is six streets in the northern part of the city, it is recommended that sensitive populations stay indoors and use air purification equipment, the general public reduce their outdoor activity time, and at the same time, implement temporary control measures of reducing production by 30% for the northern industrial area and implement temporary traffic restrictions on vehicles entering and leaving this area.
[0070] The above describes the air quality assessment method based on multi-source data in the embodiments of the present application. Next, the air quality assessment system based on multi-source data in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the air quality assessment system based on multi-source data in the embodiments of the present application includes: Calibration module, which is used to collect pollutant concentration data of fixed air monitoring stations, micro environmental stations and portable sensors, and perform outlier identification and correction processing on the data to obtain a multi-source air quality original data set; Partition module, which is used to divide the target area into urban fine grids and suburban large-scale grids according to the multi-source air quality original data set, and combine terrain, land use, roads, traffic, population and pollution source information to form an environmental characteristic grid database; Weighting module, which is used to perform weighted calculation on the air quality data in the environmental characteristic grid database according to data accuracy, environmental similarity and spatial distance to generate an air quality interpolation distribution map for the whole region; Prediction module, which is used to calculate the predicted values of pollutant concentrations in future periods according to the air quality interpolation distribution map for the whole region, combined with historical data of the same period and meteorological forecast information, to form a spatio-temporal dynamic prediction result; Tracking module, which is used to utilize the spatio-temporal dynamic prediction result, combined with wind field data and pollution source distribution, to track the pollutant transmission path and determine the pollution contribution value of various source areas to the target area; Control module, which is used to calculate the health risk levels of different regions according to the pollution contribution value, combined with population distribution and activity patterns, to generate hierarchical early warning information and control measure suggestions.
[0071] Through the coordinated cooperation of the above components, by collecting multi-source pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors, and identifying and correcting outliers, the limitations of the traditional single data source have been broken through, the integrity and reliability of the data have been significantly improved, and a solid data foundation has been laid for air quality assessment; combining terrain, land use, roads, transportation, population and pollution source information to divide the target area into fine grids in urban areas and large-scale grids in suburbs to form an environmental characteristic grid database, achieving high-resolution spatial expression of environmental elements and fully considering the heterogeneous characteristics of the regional environment; by weighted calculation according to data accuracy, environmental similarity and spatial distance, an interpolation distribution map of air quality in the entire region is generated, solving the problem of uneven spatial distribution of monitoring points. It has achieved accurate reconstruction of the pollutant concentration field; based on the interpolation distribution map of air quality in the entire region, combined with historical data for the same period and meteorological forecast information, it has calculated the predicted pollutant concentration values for future time periods, thus realizing dynamic prediction of the spatiotemporal distribution of pollutants and providing a time window for early prevention and control; using the spatiotemporal dynamic prediction results, combined with wind field data and pollution source distribution, it has tracked the pollutant transmission path, determined the pollution contribution value of various source areas to the target area, realized the quantification and scientific nature of pollution source tracing, and provided direction for precise pollution control; based on the pollution contribution value, combined with population distribution and activity patterns, it has calculated the health risk levels of different regions, generated graded warning information and control measures, and realized a closed loop from air quality assessment to health risk management, providing precise support for public health protection and pollution prevention and control decisions. It is particularly worth emphasizing that this solution applies artificial intelligence algorithms in multiple links. For example, the improved geographically weighted regression algorithm used in multi-source data spatial fusion interpolation can adaptively learn the complex nonlinear relationship between environmental characteristics and pollutant concentrations; the random forest algorithm used in the spatiotemporal dynamic prediction model can handle the high-dimensional feature interaction between meteorological factors and pollutant concentrations; the Lagrangian particle model and source contribution matrix algorithm used in pollution source analysis can accurately simulate the complex transmission process of pollutants under different meteorological conditions; the multi-factor risk model used in health risk assessment can comprehensively consider population density, activity patterns and distribution characteristics of sensitive populations. The application of these algorithms enables this solution to process a large amount of heterogeneous data, capture complex spatiotemporal patterns and causal relationships, and realize the intelligent and precise assessment and management of air quality.
[0072] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3As shown in the figure. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0073] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0074] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0075] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database, or other media provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0076] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, systems, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0077] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0078] The above is the case. The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. An air quality assessment method based on multi-source data, characterized in that: The air quality assessment method based on multi-source data includes: Collect pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors, identify and correct outliers on the data, and obtain a multi-source air quality raw data set; Based on the multi-source air quality original data set, the target area is divided into fine grids in urban areas and large-scale grids in suburban areas, and the terrain, land use, road, traffic, population and pollution source information are combined to form an environmental characteristic grid database; The air quality data in the environmental characteristic grid database is weightedly calculated according to data accuracy, environmental similarity and spatial distance to generate an interpolation distribution map of air quality for the entire region; According to the interpolated distribution map of air quality in the whole area, combined with historical data of the same period and meteorological forecast information, the predicted value of pollutant concentration in the future period is calculated to form a spatiotemporal dynamic prediction result; Using the spatiotemporal dynamic prediction results, combined with wind field data and pollution source distribution, the pollutant transmission path is tracked to determine the pollution contribution value of various source areas to the target area; Based on the pollution contribution value, combined with population distribution and activity patterns, the health risk levels of different areas are calculated, and graded warning information and control measures recommendations are generated.
2. The air quality assessment method based on multi-source data according to claim 1, characterized in that: The pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors are collected, and outlier identification and correction processing are performed on the data to obtain a multi-source air quality original data set, including: Obtain the concentration data of fine particulate matter, inhalable particulate matter, sulfur dioxide, nitrogen dioxide, ozone, carbon monoxide and the geographical coordinate information of the equipment from monitoring equipment of different accuracy levels; Establish a monitoring equipment accuracy rating system, classifying fixed air monitoring stations as high-precision equipment, micro-environmental stations as medium-precision equipment, and portable sensors as low-precision equipment; Calculate the deviation of the collected pollutant concentration data from the seasonal historical average. When the deviation exceeds three times the standard deviation, it is marked as an abnormal data point in the time dimension. The deviation percentage between each monitoring point and the average value of the five adjacent monitoring points is calculated. When the percentage exceeds 75%, it is marked as an abnormal data point in the spatial dimension; Remove the data points marked as both time dimension anomalies and space dimension anomalies, and keep the remaining data points; Cross-validate the data of adjacently placed devices with different accuracy levels and establish data conversion equations between devices with different accuracy levels; The missing data of time series are filled by linear interpolation, and the missing data of spatial distribution are filled by inverse distance weighted interpolation to form a multi-source air quality original data set.
3. The air quality assessment method based on multi-source data according to claim 1, characterized in that: According to the multi-source air quality original data set, the target area is divided into fine grids in urban areas and large-scale grids in suburban areas, and combined with terrain, land use, roads, transportation, population and pollution source information to form an environmental feature grid database, including: The urban built-up areas in the target area are divided into fine urban grid cells of 500 meters by 500 meters, and the suburban areas in the target area are divided into large-scale suburban grid cells of 2,000 meters by 2,000 meters; Extract digital elevation data from the geographic information system, calculate the average altitude and terrain relief of each grid cell, and obtain terrain feature data; Extract the dominant land type and land type composition ratio of each grid unit from the land use classification data to form land use characteristic data; Obtain road network data from the traffic management department, calculate the road length density and the number of road intersections for each grid unit, and generate road feature data; Obtain traffic data from the traffic flow monitoring system, calculate the average vehicle flow and traffic congestion index of each grid unit, and obtain traffic characteristic data; Extract population distribution information from census data, calculate the resident population density and floating population density of each grid unit, and form population characteristic data; Obtain a list of fixed pollution sources from the environmental protection department, mark the number of industrial enterprises and emission levels in each grid unit, and generate pollution source characteristic data; Through spatial overlay analysis, the terrain feature data, land use feature data, road feature data, traffic feature data, population feature data and pollution source feature data are associated with grid units to form an environmental feature grid database.
4. The air quality assessment method based on multi-source data according to claim 1, characterized in that: The air quality data in the environmental feature grid database is weightedly calculated according to data accuracy, environmental similarity and spatial distance to generate an interpolation distribution map of air quality in the entire region, including: Assign accuracy weight coefficients to the data of each monitoring device, where the weight of fixed air monitoring station data is 1, the weight of micro-environmental station data is 0.6, and the weight of portable sensor data is 0.3; Calculating the environmental similarity between the grid unit to be interpolated and the known monitoring point, and obtaining the environmental similarity weight by calculating the vector cosine similarity of the terrain, land use, road, traffic, population and pollution source characteristics; Based on the road network data, the actual reachable distance between the grid unit to be interpolated and the monitoring point is calculated to form a spatial distance weight; For each grid cell to be interpolated, extract the data of all monitoring points within a radius of 5 kilometers; When the grid cell to be interpolated is located downwind of the monitoring point, the influence weight of the monitoring point is enhanced according to the wind field data; The accuracy weight coefficient, the environmental similarity weight and the spatial distance weight are normalized and multiplied to obtain a comprehensive weight coefficient; Using the comprehensive weight coefficient, weighted average calculation is performed on the pollutant concentration data of the monitoring points to obtain the pollutant concentration value of the grid unit to be interpolated, and the interpolation uncertainty index is calculated; According to the division of fine grids in urban areas and large-scale grids in suburbs, the pollutant concentration value of each grid unit is calculated one by one to form the spatial distribution data of pollutant concentration; The spatial distribution data of the pollutant concentration is visualized by the contour line method to generate an interpolation distribution map of the air quality in the entire region.
5. The air quality assessment method based on multi-source data according to claim 1, characterized in that: The method of calculating the predicted pollutant concentration value for the future period based on the interpolated distribution map of air quality in the whole region, combining historical data for the same period and meteorological forecast information, and forming a spatiotemporal dynamic prediction result includes: Extract the interpolation distribution map of the air quality in the target area over the same period in the past three years and build a database of pollutant concentrations in the same period; Obtain wind direction, wind speed, temperature, humidity, air pressure, and precipitation probability data for the next 72 hours from the meteorological department to form a weather forecast data set; Performing time series decomposition on the interpolated distribution map of air quality in the entire region to separate the long-term trend and periodic variation pattern of pollutant concentration; Determine the seasonal variation pattern of pollutant concentrations based on the historical pollutant concentration database and generate a basic prediction curve; Using the weather forecast data set, adjusting the basic prediction curve to form a weather factor correction value; Taking into account the impact of pollution conditions in the upwind area, the meteorological factor correction value is subjected to spatial transmission correction to generate a spatiotemporal dynamic prediction result.
6. The air quality assessment method based on multi-source data according to claim 1, characterized in that: The spatiotemporal dynamic prediction results are used in combination with wind field data and pollution source distribution to track pollutant transmission paths and determine the pollution contribution values of various source areas to the target area, including: Obtain wind direction and wind speed data in the target area within 72 hours, divide the wind direction into 16 sectors, divide the wind speed into six levels, and construct a wind field data grid; Extract the location information, types of pollutants emitted and their hourly average emission intensity of industrial parks, high-density traffic areas, heating areas, and surrounding urban transmission areas from the pollution source list of the environmental protection department; Based on the wind field data grid, each grid unit is regarded as a starting point using the Lagrangian particle method, and pollutant particles are tracked in reverse according to the wind direction, and the positions of pollutant particles are recorded every hour to form a 72-hour pollutant transmission path map; Performing spatial overlay analysis on the pollutant transmission path diagram, counting the number of times each source area is crossed by the particle path, dividing the number by the total number of particles, and obtaining the source area impact frequency value; The source areas are classified and marked according to industrial sources, traffic sources, living sources, and regional transmission sources, and the source area impact frequency values in each source area are summed to obtain the comprehensive impact frequency of each source area; The comprehensive impact frequency of each type of source area is multiplied by its emission intensity, and the attenuation coefficient is set according to the distance between the source area and the target area to obtain the pollution contribution value of each type of source area to the target area.
7. The air quality assessment method based on multi-source data according to claim 1, characterized in that: The health risk levels of different areas are calculated based on the pollution contribution value, combined with population distribution and activity patterns, and graded warning information and control measures are generated, including: Obtain the resident population density and floating population density of the target area from the census database and construct a population distribution data map; Based on social survey data, the activity patterns of urban residents are extracted, and a day is divided into commuting peak hours, working hours, leisure hours, and night rest hours to form a time-period activity feature library; For the respiratory disease patient data received from medical institutions, mark the patient's residential area and generate a sensitive population distribution map; Calculate the population exposure of each grid unit at different time periods by combining the pollution contribution value, population distribution data map, time period activity feature library and sensitive population distribution map; Set health risk thresholds based on the population exposure, and divide the risk values into four health risk levels: low risk, medium risk, high risk, and extremely high risk; Based on the health risk level, graded warning information and control measures recommendations are generated, including the warning period, impact range, protective measures and control measures recommendations for major pollution source areas.
8. An air quality assessment system based on multi-source data, used to implement the air quality assessment method based on multi-source data as described in any one of claims 1 to 7, characterized in that: The air quality assessment system based on multi-source data includes: The correction module is used to collect pollutant concentration data from fixed air monitoring stations, micro-environmental stations and portable sensors, and perform outlier identification and correction processing on the data to obtain a multi-source air quality original data set; A partitioning module is used to divide the target area into fine grids in urban areas and large-scale grids in suburban areas according to the multi-source air quality original data set, and to form an environmental characteristic grid database by combining terrain, land use, roads, transportation, population and pollution source information; A weighting module is used to perform weighted calculation on the air quality data in the environmental feature grid database according to data accuracy, environmental similarity and spatial distance to generate an interpolation distribution map of air quality for the entire region; A prediction module is used to calculate the predicted value of pollutant concentration in the future period according to the interpolation distribution map of air quality in the whole area, combined with historical data of the same period and meteorological forecast information, to form a spatiotemporal dynamic prediction result; A tracking module is used to use the spatiotemporal dynamic prediction results, combined with wind field data and pollution source distribution, to track the pollutant transmission path and determine the pollution contribution value of various source areas to the target area; The control module is used to calculate the health risk level of different areas based on the pollution contribution value, combined with population distribution and activity patterns, and generate graded warning information and control measure recommendations.
9. A computer device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and is characterized in that when the processor executes the computer program, the air quality assessment method based on multi-source data described in any one of claims 1 to 7 is implemented. 10 . A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is enabled to perform the air quality assessment method based on multi-source data according to any one of claims 1 to 7 .
Citation Information
Patent Citations
Method for estimating fine particulate matter concentration in real time based on temporal and spatial characteristics
CN105117610A
Multi-scale air quality space interpolation method, system, medium and equipment
CN109636719A
Site pollution feature analysis method and device, electronic equipment and storage medium
CN113111964A
Atmospheric pollutant concentration data prediction method and device, equipment and storage medium
CN118425423A
Industrial park environment quality monitoring system
CN118446513A
Cited By
Method, device and equipment for identifying target in farmland based on multi-source information fusion and medium
CN120429829A
Urban pollutant distribution prediction method based on multi-source data dynamic and static feature fusion
CN120496663A
Construction site pollution source analysis method and system
CN120635789A
Carbon emission monitoring method and system based on space-time distribution
CN121212577A
VOCs spatial and temporal distribution rule analysis system and method
CN121253702A