Pollutant concentration prediction method and device, storage medium and electronic equipment

By regionally dividing, clustering and merging the steel industrial plant area, determining the target prediction model, and combining historical data and meteorological data to predict, the problem of low accuracy of pollutant concentration prediction in the existing technology is solved, and higher prediction accuracy is achieved.

CN120069240AInactive Publication Date: 2025-05-30CENT RES INST OF BUILDING & CONSTR CO LTD MCC GRP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510542981.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing pollutant concentration prediction methods cannot accurately predict pollutant concentrations in various areas of the steel industrial plant area, resulting in a low prediction accuracy.

Method used

The target areas are divided based on the location information of multiple monitoring points, clustered and merged the first sub-regions to determine the target prediction model of each second sub-regions, and predicted in combination with historical pollutant data and meteorological data.

Benefits of technology

Accurate prediction of pollutant concentrations in each second sub-region is achieved, and the accuracy of pollutant concentration prediction results in the target area is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069240A_ABST
    Figure CN120069240A_ABST
Patent Text Reader

Abstract

The invention provides a pollutant concentration prediction method and device, a storage medium and electronic equipment, and belongs to the technical field of pollutant concentration prediction. According to the embodiment of the invention, the method comprises the steps: dividing a target region based on the position information of a plurality of monitoring points of the target region, obtaining a plurality of first sub-regions corresponding to the plurality of monitoring points in a one-to-one manner, carrying out the clustering of the plurality of first sub-regions based on the pollutant data of the plurality of monitoring points, obtaining a plurality of second sub-regions corresponding to a plurality of clusters in a one-to-one manner, and carrying out the clustering of the plurality of second sub-regions. And then the corresponding target prediction model is utilized to perform targeted prediction on each second sub-region, so that accurate prediction of the pollutant concentration of each second sub-region can be realized, and the accuracy of the pollutant concentration prediction result of the target region is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of pollutant concentration prediction, and particularly to a method, device, storage medium and electronic device for predicting pollutant concentration. Background Art

[0002] In traditional pollutant concentration prediction methods, historical pollutant concentration data collected at each monitoring point in an iron and steel industrial plant area is usually collected uniformly, and the future overall pollution situation in this area is directly predicted based on this historical pollutant concentration data.

[0003] However, since there are usually significant differences in pollutant emissions in different areas of an iron and steel industrial plant area, this method cannot accurately predict the pollutant concentration in each area of the iron and steel industrial plant area, and thus there is a problem of low accuracy in pollutant concentration prediction. Summary of the Invention

[0004] The present application provides a method, device, storage medium and electronic device for predicting pollutant concentration to solve the problem of low accuracy in current pollutant concentration prediction.

[0005] To solve the above problems, the present application adopts the following technical solutions: In a first aspect, an embodiment of the present application provides a method for predicting pollutant concentration, the method comprising: Dividing the target area based on the location information of multiple monitoring points in the target area to obtain multiple first sub-areas corresponding one-to-one to the multiple monitoring points; Clustering the multiple first sub-areas based on the pollutant data of the multiple monitoring points to obtain multiple clustering clusters; different clustering clusters correspond to different pollutant category characteristics; For any one of the clustering clusters, merging each first sub-area in the clustering cluster to obtain multiple second sub-areas corresponding one-to-one to the multiple clustering clusters; For any one of the second sub-areas, determining the target prediction model of the second sub-area in a preset prediction model library based on the pollutant category characteristic corresponding to the second sub-area; For any one of the second sub-areas, inputting the historical pollutant data of each monitoring point in the second sub-area and the meteorological data of the prediction time period into the target prediction model of the second sub-area, and outputting the pollutant concentration of the second sub-area in the prediction time period; Based on the pollutant concentrations of each second sub-area in the prediction time period, obtaining the pollutant concentration prediction result of the target area.

[0006] In an embodiment of the present application, based on the position information of multiple monitoring points in the target area, the target area is divided to obtain a plurality of first sub-areas corresponding one-to-one to the multiple monitoring points, including: Based on the position information of multiple monitoring points in the target area, the target area is divided by using the Delaunay triangulation algorithm to generate a Delaunay triangulation graph; Based on the Delaunay triangulation graph, Voronoi polygons are generated to obtain a plurality of first sub-areas corresponding one-to-one to the multiple monitoring points.

[0007] In an embodiment of the present application, based on the pollutant data of multiple monitoring points, the multiple first sub-areas are clustered to obtain a plurality of clustering clusters, including: The pollutant data of multiple monitoring points are preprocessed to obtain characteristic data corresponding to each of the multiple monitoring points; A clustering operation is performed on the characteristic data corresponding to each of the multiple monitoring points to divide the multiple first sub-areas into a plurality of clustering clusters.

[0008] In an embodiment of the present application, the pollutant data of multiple monitoring points are preprocessed to obtain characteristic data corresponding to each of the multiple monitoring points, including: For the pollutant data of any one of the monitoring points, the text data in the pollutant data is converted into numerical data to obtain the numericalized data corresponding to the pollutant data; Feature processing is performed on the numericalized data to obtain the characteristic data corresponding to the monitoring point.

[0009] In an embodiment of the present application, a clustering operation is performed on the characteristic data corresponding to each of the multiple monitoring points to divide the multiple first sub-areas into a plurality of clustering clusters, including: A preset clustering algorithm is used to perform a clustering operation on the characteristic data corresponding to each of the multiple monitoring points according to the target number of clustering clusters, so as to divide the multiple first sub-areas into a plurality of clustering clusters with the same number as the target number of clustering clusters.

[0010] In an embodiment of the present application, the prediction model library includes a plurality of prediction models; wherein, different prediction models are trained based on pollutant samples with different pollutant category characteristics, and different prediction models have different pollutant category characteristic labels; Based on the pollutant category characteristics corresponding to the second sub-area, in a preset prediction model library, the target prediction model of the second sub-area is determined, including: Determine the similarity between the pollutant category characteristics corresponding to the second sub-area and the pollutant category characteristic labels of multiple prediction models; Determine the prediction model corresponding to the pollutant category feature label with the highest similarity as the target prediction model for the second sub-region.

[0011] In an embodiment of the present application, based on the pollutant concentrations of each of the second sub-regions during the prediction time period, obtaining the pollutant concentration prediction result of the target region includes: Based on the concentration influence factors of each of the second sub-regions, determine the influence weights of each of the second sub-regions; the concentration influence factors include the regional area, topographic features, and the pollutant emission intensity and meteorological data during the prediction time period; Based on the influence weights of each of the second sub-regions and the pollutant concentrations during the prediction time period, obtain the pollutant concentration prediction result of the target region.

[0012] Second, based on the same inventive concept, an embodiment of the present application provides a pollutant concentration prediction device, and the device includes: A region division module, configured to divide the target region based on the position information of multiple monitoring points in the target region to obtain multiple first sub-regions corresponding one-to-one to the multiple monitoring points; A region clustering module, configured to cluster the multiple first sub-regions based on the pollutant data of the multiple monitoring points to obtain multiple clustering clusters; different clustering clusters correspond to different pollutant category features; A region merging module, configured to, for any one of the clustering clusters, merge each of the first sub-regions in the clustering cluster to obtain multiple second sub-regions corresponding one-to-one to the multiple clustering clusters; A model determination module, configured to, for any one of the second sub-regions, determine the target prediction model of the second sub-region in a preset prediction model library based on the pollutant category feature corresponding to the second sub-region; A concentration prediction module, configured to, for any one of the second sub-regions, input the historical pollutant data of each monitoring point in the second sub-region and the meteorological data during the prediction time period into the target prediction model of the second sub-region, and output the pollutant concentration of the second sub-region during the prediction time period; A concentration determination module, configured to obtain the pollutant concentration prediction result of the target region based on the pollutant concentrations of each of the second sub-regions during the prediction time period.

[0013] In an embodiment of the present application, the region division module includes: A region division sub-module, configured to divide the target region based on the position information of the multiple monitoring points in the target region through a Delaunay triangulation algorithm to generate a Delaunay triangulation map; The Thiessen polygon generation sub-module is used to generate Thiessen polygons based on the Delaunay triangulation graph to obtain a plurality of first sub-regions corresponding one by one to the plurality of monitoring points.

[0014] In an embodiment of the present application, the region clustering module includes: The preprocessing sub-module is used to preprocess the pollutant data of the plurality of monitoring points to obtain characteristic data corresponding to each of the plurality of monitoring points; The clustering sub-module is used to perform a clustering operation on the characteristic data corresponding to each of the plurality of monitoring points to divide the plurality of first sub-regions into a plurality of clustering clusters.

[0015] In an embodiment of the present application, the preprocessing sub-module includes: The data conversion unit is used to convert the text data in the pollutant data into numerical data for any of the pollutant data of the monitoring points to obtain the numerical data corresponding to the pollutant data; The feature processing unit is used to perform feature processing on the numerical data to obtain the characteristic data corresponding to the monitoring point.

[0016] In an embodiment of the present application, the clustering sub-module includes: The clustering unit is used to perform a clustering operation on the characteristic data corresponding to each of the plurality of monitoring points according to a preset clustering algorithm according to the target number of clustering clusters to divide the plurality of first sub-regions into a plurality of clustering clusters with the same number as the target number of clustering clusters.

[0017] In an embodiment of the present application, the prediction model library includes a plurality of prediction models; among them, different prediction models are trained based on pollutant samples with different pollutant category characteristics, and different prediction models have different pollutant category characteristic labels; the model determination module includes: The similarity determination sub-module is used to determine the similarity between the pollutant category characteristics corresponding to the second sub-region and the pollutant category characteristic labels of the plurality of prediction models; The model determination sub-module is used to determine the prediction model corresponding to the pollutant category characteristic label with the highest similarity as the target prediction model of the second sub-region.

[0018] In an embodiment of the present application, the concentration determination module includes: The influence weight determination sub-module is used to determine the influence weight of each of the second sub-regions based on the concentration influence factors of each of the second sub-regions; the concentration influence factors include the regional area, terrain characteristics, and the pollutant emission intensity and meteorological data of the prediction time period; A prediction result determination sub-module, configured to obtain a pollutant concentration prediction result of the target area based on the influence weights of each of the second sub-areas and the pollutant concentration in the prediction time period.

[0019] In a third aspect, based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, on which an executable program is stored, and when the executable program is executed by a processor, the pollutant concentration prediction method proposed in the first aspect of the present application is implemented.

[0020] In a fourth aspect, based on the same inventive concept, an embodiment of the present application provides an electronic device, including: A memory, configured to store an executable program; A processor; When the executable program is executed by the processor, the pollutant concentration prediction method proposed in the first aspect of the present application is implemented.

[0021] Compared with the prior art, the present application has the following advantages: A pollutant concentration prediction method provided by an embodiment of the present application first divides a target area based on the position information of multiple monitoring points in the target area to obtain multiple first sub-areas corresponding one-to-one to the multiple monitoring points; then clusters the multiple first sub-areas based on the pollutant data of the multiple monitoring points to obtain multiple clustering clusters; and merges each first sub-area in the clustering clusters to obtain multiple second sub-areas corresponding one-to-one to the multiple clustering clusters; subsequently, based on the pollutant category characteristics corresponding to each second sub-area, a target prediction model of the second sub-area is determined in a preset prediction model library, and the historical pollutant data of each monitoring point in the second sub-area and the meteorological data in the prediction time period are input into the target prediction model of the second sub-area, and the pollutant concentration of the second sub-area in the prediction time period is output; finally, based on the pollutant concentration of each second sub-area in the prediction time period, a pollutant concentration prediction result of the target area is obtained. By merging the first sub-areas with similar pollutant category characteristics and using the corresponding target prediction model to perform targeted prediction on the merged second sub-areas, the embodiment of the present application can accurately predict the pollutant concentration of each second sub-area, thereby improving the accuracy of the pollutant concentration prediction result of the target area. Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0023] Figure 1 It is a flowchart of the steps of a method for predicting pollutant concentration in an embodiment of the present application.

[0024] Figure 2 It is a schematic diagram of the modules of a device for predicting pollutant concentration in an embodiment of the present application.

[0025] Figure 3 It is a schematic structural diagram of an electronic device in an embodiment of the present application. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that as an important basic industry, the steel industry involves the extraction, processing, and smelting of a large number of raw materials in its production process. Therefore, various pollutants will be generated in each production link, including but not limited to sulfur dioxide, nitrogen oxides, dust, and heavy metals. The emission situations of these pollutants usually vary significantly in different areas of the plant area, mainly affected by the following factors: 1. Different production processes: Different stages in the steel production process (such as ironmaking, steelmaking, and rolling, etc.) use different equipment and technologies, and the types and quantities of pollutants generated by them are different. For example, the blast furnace emissions in the ironmaking process will generate a large amount of sulfur dioxide and particulate matter, while the steelmaking process may release more nitrogen oxides.

[0028] 2. Equipment layout and area division: Steel plants usually divide the plant area into different areas according to the production process, such as raw material area, smelting area, waste gas treatment area, etc. The environmental characteristics, equipment configuration, and operation methods in different areas will also lead to differences in the generation and diffusion patterns of pollutants. For example, in the areas close to the furnaces or smelting equipment, the pollutant concentration is usually higher, while in the areas far from the production core area, it is relatively lower.

[0029] 3. Climate and meteorological conditions: Meteorological factors such as wind direction, temperature, humidity, and rainfall will all have important impacts on the diffusion and settlement of pollutants. Under certain meteorological conditions, pollutants may accumulate in specific areas, resulting in increased local concentration; while under other conditions, pollutants may be effectively diluted or migrate to other areas.

[0030] 4. Use of emission control equipment: Different regions may be equipped with different types of pollutant treatment facilities, such as flue gas desulfurization, denitrification, and dust removal equipment. The effectiveness and operating conditions of these facilities will directly affect the final emission level of pollutants, resulting in differences between regions.

[0031] 5. Management and supervision measures: The strictness of plant management and the application of monitoring technologies will also affect pollutant emissions. For example, in some regions, more cleaner production technologies may be implemented due to stricter supervision measures, resulting in a significant reduction in pollutant emissions.

[0032] Due to the complex interaction of the above various factors, it is often difficult to accurately reflect the actual pollutant concentration levels in each region of the iron and steel industrial plant area by using a single prediction model for unified prediction of the entire plant area.

[0033] Aiming at the problem that the pollutant concentration in each region of the iron and steel industrial plant area cannot be accurately predicted in the related technologies, the present application aims to provide a pollutant concentration prediction method. By merging the first sub-regions with similar pollutant category characteristics and using the corresponding target prediction model to specifically predict the merged second sub-regions, it is possible to accurately predict the pollutant concentration in each second sub-region, thereby improving the accuracy of the pollutant concentration prediction result in the target area.

[0034] Refer to Figure 1 , which shows a pollutant concentration prediction method of the present application. The method may include the following steps: S101: Based on the position information of multiple monitoring points in the target area, divide the target area to obtain multiple first sub-regions corresponding one by one to the multiple monitoring points.

[0035] In this embodiment, the target area may be an iron and steel industrial plant area to be monitored. The iron and steel industrial plant area is provided with multiple monitoring points, and different monitoring points are used to monitor pollutant data generated in different production links or different production areas.

[0036] In this embodiment, by dividing the target area based on the position information of multiple monitoring points, the effectiveness and scientific nature of data analysis can be significantly improved. Since multiple monitoring points correspond one by one to multiple first sub-regions, the pollutant data collected by each monitoring point can accurately reflect the environmental characteristics of the corresponding first sub-region, making the pollutant sources, meteorological conditions, and pollutant concentration distribution in each first sub-region relatively consistent, which is more convenient for refined monitoring and data analysis.

[0037] S102: Cluster the multiple first sub-regions based on the pollutant data of the multiple monitoring points to obtain multiple clusters.

[0038] In this embodiment, different clustering clusters correspond to different pollutant category characteristics, and each first sub-region under the same clustering cluster has the same pollutant category characteristics. For example, each first sub-region under the same clustering cluster is used to produce the same or similar industrial products and generates the same or similar pollutants.

[0039] In this embodiment, by clustering multiple first sub-regions according to the pollutant data of multiple monitoring points, the first sub-regions with the same pollutant category characteristics can be divided into the same clustering cluster for unified analysis, thereby improving the prediction efficiency of pollutant concentration while ensuring the prediction accuracy.

[0040] S103: For any clustering cluster, merge each first sub-region in the clustering cluster to obtain multiple second sub-regions corresponding one by one to the multiple clustering clusters.

[0041] In this embodiment, each clustering cluster includes one or more first sub-regions. By merging each first sub-region in each clustering cluster, multiple second sub-regions can be obtained, and different second sub-regions correspond to different pollutant category characteristics.

[0042] S104: For any second sub-region, based on the pollutant category characteristics corresponding to the second sub-region, determine the target prediction model of the second sub-region in the preset prediction model library.

[0043] In this embodiment, the prediction model library is used to store multiple pre-trained prediction models. Among them, different prediction models are trained based on pollutant samples with different pollutant category characteristics.

[0044] In this embodiment, considering that pollutant data is usually time series data, the prediction model can be constructed based on the LSTM (Long Short-Term Memory Neural Network) neural network. It should be noted that the LSTM neural network is highly sensitive to time series data, is an optimized model based on the RNN (Recurrent Neural Network), is more suitable for the analysis and prediction of time series data, and as an optimized model based on the RNN, the LSTM neural network can effectively avoid the phenomena of gradient explosion and gradient disappearance.

[0045] In this embodiment, by training and storing corresponding prediction models for different pollutant category characteristics, the appropriate target prediction model can be quickly matched in the prediction model library based on the pollutant category characteristics corresponding to the second sub-region for pollutant concentration prediction, thereby improving the accuracy of concentration prediction.

[0046] S105: For any second sub-region, input the historical pollutant data of each monitoring point in the second sub-region and the meteorological data of the prediction time period into the target prediction model of the second sub-region, and output the pollutant concentration of the second sub-region in the prediction time period.

[0047] In this embodiment, the meteorological data may include wind speed data, wind direction data, temperature data, humidity data, etc. of the prediction time period. These data are obtained through the prediction data provided by meteorological agencies or meteorological service platforms, or can also be output by a pre-trained meteorological prediction model according to historical meteorological data.

[0048] In this embodiment, since meteorological factors have a direct impact on the diffusion and transformation of pollutants, therefore, by integrating historical pollutant data and meteorological data of the prediction time period, the target prediction model can more accurately predict the pollutant concentration in the prediction time period.

[0049] In this embodiment, the pollutant concentration of the second sub-region in the prediction time period includes the concentration conditions of different types of pollutants in the second sub-region in the prediction time period, and thus can accurately reflect the change trend of the pollutant concentration.

[0050] S106: Based on the pollutant concentrations of each second sub-region in the prediction time period, obtain the pollutant concentration prediction result of the target region.

[0051] In this embodiment, by integrating the pollutant concentrations of each second sub-region, the prediction results of different second sub-regions can be integrated into the overall prediction value of the target region, and thus the pollutant concentration prediction result of the target region can be obtained.

[0052] In this embodiment, after obtaining the pollutant concentration prediction result of the target region, it can be presented in the form of a report, graph, or data visualization, etc., to facilitate users to intuitively and comprehensively understand the air quality status of the target region.

[0053] In this embodiment, corresponding concentration thresholds can also be set for different pollutants, and then the pollutant concentration in the prediction time period can be warned. For example, three concentration thresholds can be set to divide the warning level into 3 levels, with level 1 being the highest level. When the concentration prediction result of any pollutant is greater than any concentration threshold, the warning measures corresponding to the concentration threshold will be triggered to achieve the purpose of timely and accurate warning.

[0054] In this embodiment, by merging the first sub-regions with similar pollutant category characteristics and using the corresponding target prediction model to conduct targeted prediction on the merged second sub-regions, the accurate prediction of the pollutant concentrations of each second sub-region can be realized, and thus the accuracy of the pollutant concentration prediction result of the target region can be improved.

[0055] In a feasible implementation, S101 may specifically include the following sub-steps: S101-1: Based on the position information of multiple monitoring points in the target area, divide the target area through the Delaunay triangulation algorithm to generate a Delaunay triangulation graph.

[0056] In this implementation, to ensure that each monitoring point has a corresponding first sub-region, the target area will be divided based on the Thiessen polygon.

[0057] It should be noted that the Thiessen polygon is also called the Voronoi diagram. Hereinafter, it is simply referred to as the Voronoi diagram. The Voronoi diagram divides the two-dimensional plane into n Voronoi polygons according to a set of given points, and each Voronoi polygon is a Voronoi cell. For each point on the plane, find the nearest seed point, and the area composed of all these points is the Voronoi cell of this seed point. Therefore, each seed point has a corresponding area, and all points within this area are the closest to this seed point. It should be noted that common algorithms for constructing the Thiessen polygon include the sweep line algorithm, the divide-and-conquer algorithm, the incremental method, etc.

[0058] In this implementation, the most efficient Delaunay triangulation algorithm will be used to construct the Thiessen polygon. In specific implementation, the position information of multiple monitoring points in the target area is used as the given points to obtain a set of given points, and then the Delaunay triangulation algorithm is used to triangulate the set of given points to generate a set of Delaunay triangles to obtain a Delaunay triangulation graph.

[0059] S101-2: Based on the Delaunay triangulation graph, generate the Thiessen polygon to obtain multiple first sub-regions corresponding one-to-one to multiple monitoring points.

[0060] It should be noted that the Thiessen polygon and the Delaunay triangulation graph are dual geometric structures. Specifically, each Delaunay triangle corresponds to a Voronoi vertex (that is, the seed point of the Voronoi cell) in the Voronoi diagram, and each edge of the Delaunay triangulation corresponds to the edge of the Voronoi polygon in the Voronoi diagram. Therefore, as long as the Delaunay triangulation graph is obtained, the corresponding Thiessen polygon can be transformed.

[0061] In this implementation, each monitoring point corresponds to a Voronoi cell, and this Voronoi cell is the first sub-region corresponding to this monitoring point.

[0062] In this embodiment, by generating Thiessen polygons based on the spatial positions of each monitoring point, the target area is divided into a plurality of first sub-regions with the same number as the monitoring points, which can clarify the coverage range of each monitoring point, ensure that the distance from the points in each area to this monitoring point is the shortest, thus avoiding data duplication or omission; at the same time, it is convenient to identify the main pollution sources in each area and facilitate the targeted monitoring and treatment of pollution sources.

[0063] In a feasible embodiment, S102 may specifically include the following sub-steps: S102-1: Preprocess the pollutant data of multiple monitoring points to obtain the characteristic data corresponding to each monitoring point.

[0064] In this embodiment, the preprocessing may specifically include data cleaning operations and data conversion operations. Among them, the data cleaning operations include handling missing values and outlier processing, etc. Specifically, the missing values can be filled by interpolation method, mean filling, etc., and the outlier processing can specifically use statistical methods (such as Z-score) or data mining techniques to identify and process outliers. The data conversion operations may specifically include standardization or normalization processing to unify the data scale, making the pollutant data of each monitoring point more comparable.

[0065] In this embodiment, considering that in addition to numerical data, the pollutant data also includes text data, such as the names of various pollutants and the description information of pollution sources, etc., therefore, to improve the clustering effect, for the pollutant data of any monitoring point, the text data in the pollutant data can be converted into numerical data to obtain the numerical data corresponding to the pollutant data; perform feature processing on the numerical data to obtain the characteristic data corresponding to the monitoring point.

[0066] In specific implementation, methods such as label encoding and one-hot encoding can be used to convert the text data in the pollutant data into numerical data.

[0067] In this embodiment, since the pollutant data has been converted into numerical data that is convenient for the model to identify, therefore, by performing feature processing on the numerical data, the characteristic data corresponding to each pollutant data can be obtained.

[0068] In specific implementation, by performing normalization processing on the numerical data, such as linear normalization processing, a linear transformation of various types of numerical data can be realized, and then the numerical data is mapped between [0,1]. By mapping various types of numerical data to the same value range, it is more conducive to data processing, so as to improve the efficiency and clustering effect of data processing when performing clustering operations later.

[0069] S102-2: Perform a clustering operation on the characteristic data corresponding to each of the multiple monitoring points to divide the multiple first sub-regions into multiple clusters.

[0070] In this embodiment, by analyzing the characteristic data, the similarity between the pollutant data of each monitoring point can be obtained. Then, the first sub-regions with the same or similar pollutant data are clustered together to obtain multiple clusters.

[0071] In a specific implementation, a preset clustering algorithm can be used to perform a clustering operation on the characteristic data corresponding to each of the multiple monitoring points according to the target number of clusters, so as to divide the multiple first sub-regions into multiple clusters with the same number as the target number of clusters. Among them, the clustering algorithm can adopt the fuzzy c-means algorithm (FCMA) or the K-means clustering algorithm.

[0072] In this embodiment, the target number of clusters can be determined based on the types of pollutants that can be generated in the iron and steel industrial plant area. For example, when the types of pollutants that can be generated in the iron and steel industrial plant area include three types of pollutants: gaseous pollutants (such as sulfur dioxide, nitrogen oxides, carbon monoxide, ozone, etc.), particulate matter (such as PM2.5, PM10), and organic pollutants (such as volatile organic compounds, polycyclic aromatic hydrocarbons), the target number of clusters can be determined as 3. In this way, by selecting an appropriate target number of clusters, the clustering effect can be further improved.

[0073] In a feasible embodiment, the prediction model library includes multiple prediction models, and different prediction models have different pollutant category characteristic labels; S104 can specifically include the following sub-steps: S104-1: Determine the similarity between the pollutant category characteristics corresponding to the second sub-region and the pollutant category characteristic labels of the multiple prediction models.

[0074] In this embodiment, considering that pollutants with different pollutant category characteristics usually have different diffusion patterns, therefore, corresponding prediction models are respectively trained using pollutant samples with different pollutant category characteristics, so that the multiple prediction models can fully learn the diffusion patterns of various pollutants, and then achieve accurate prediction of the concentrations of various pollutants.

[0075] In this embodiment, different prediction models have different pollutant category characteristic labels, and the pollutant category characteristic label represents that the prediction model is used to predict the concentration of pollutants with corresponding pollutant category characteristics.

[0076] In this embodiment, after determining the pollutant category characteristics corresponding to any second sub-region, by comparing the pollutant category characteristics with the pollutant category characteristic labels of each prediction model one by one, the similarity with each pollutant category characteristic label can be calculated.

[0077] S104-2: Determine the prediction model corresponding to the pollutant category characteristic label with the highest similarity as the target prediction model for the second sub-region.

[0078] In this embodiment, by calculating the similarity, a suitable target prediction model can be matched for the second sub-region among multiple prediction models for targeted prediction, thereby significantly improving the accuracy of the prediction.

[0079] In a feasible embodiment, S106 may specifically include the following sub-steps: S106-1: Determine the influence weights of each second sub-region based on the concentration influence factors of each second sub-region.

[0080] In this embodiment, the concentration influence factors include the regional area, terrain features, and pollutant emission intensity and meteorological data during the prediction period. Among them, the pollutant emission intensity during the prediction period can be determined based on the historical pollutant emission intensity of the second sub-region.

[0081] It should be noted that a second sub-region with a large area usually contains more pollution sources, and its total pollutant emissions may be larger, thus having a greater impact on the pollutant concentration in the target region. In addition, a sub-region with a large area may have stronger diffusion ability and a stronger dilution effect on pollutants.

[0082] It should be noted that the terrain features of the second sub-region, such as altitude, valleys, and plains, will significantly affect the diffusion and transmission of pollutants. For example, valley terrain may hinder the diffusion of pollutants, resulting in the accumulation of pollutants at the bottom of the valley, while plain areas are conducive to the diffusion and dilution of pollutants.

[0083] It should be noted that a second sub-region with a high emission intensity will release more pollutants during the prediction period, having a greater impact on the pollutant concentration in the target region.

[0084] It should be noted that meteorological conditions such as wind speed, wind direction, temperature, and humidity directly affect the diffusion and transmission of pollutants. In the case of high wind speed, pollutants diffuse faster and the concentration distribution is more uniform; in the case of low wind speed, pollutants are prone to accumulate, resulting in an increase in local concentration. Temperature and humidity also affect the chemical reaction rate and physical state of pollutants, thereby affecting their concentration distribution.

[0085] In specific implementation, the influence weight of the second sub-region can be determined according to the following formula: Wi = α×w1 + β×w2 + γ×w3 + δ×w4; Wherein, Wi represents the influence weight of the i-th second sub-region; w1 represents the influence weight of the regional area of the second sub-region; w2 represents the influence weight of the regional area of the second sub-region; w3 represents the influence weight of the pollutant emission intensity of the second sub-region during the prediction period; w4 represents the influence weight of the meteorological data of the second sub-region during the prediction period; α, β, γ, and δ respectively represent the weight coefficients of w1, w2, w3, and w4.

[0086] In this embodiment, α, β, γ, and δ can be optimized by historical data regression to improve the accuracy of each weight coefficient.

[0087] S106-2: Obtain the pollutant concentration prediction result of the target region based on the influence weights of each second sub-region and the pollutant concentration during the prediction period.

[0088] In this embodiment, the weighted average method can be used to multiply the pollutant concentration of each second sub-region during the prediction period by its corresponding influence weight, and then sum to obtain the overall pollutant concentration prediction result of the target region.

[0089] In this embodiment, by comprehensively considering factors such as regional area, terrain features, pollutant emission intensity, and meteorological data, the influence weights of each second sub-region can be determined more accurately, thereby improving the accuracy and reliability of the pollutant concentration prediction of the target region.

[0090] Second, referring to Figure 2 , the embodiment of the present application provides a pollutant concentration prediction device 200, and the pollutant concentration prediction device 200 includes: A region division module 201, configured to divide the target region based on the position information of multiple monitoring points in the target region to obtain multiple first sub-regions corresponding one-to-one to the multiple monitoring points; A region clustering module 202, configured to cluster the multiple first sub-regions based on the pollutant data of the multiple monitoring points to obtain multiple clustering clusters; different clustering clusters correspond to different pollutant category characteristics; A region merging module 203, configured to merge each first sub-region in the clustering cluster for any one clustering cluster to obtain multiple second sub-regions corresponding one-to-one to the multiple clustering clusters; A model determination module 204, configured to determine the target prediction model of the second sub-region in the preset prediction model library based on the pollutant category characteristics corresponding to the second sub-region for any one second sub-region; A concentration prediction module 205, configured to input historical pollutant data of each monitoring point in a second sub-region and meteorological data of a prediction time period into a target prediction model of the second sub-region for any second sub-region, and output the pollutant concentration of the second sub-region in the prediction time period; A concentration determination module 206, configured to obtain a pollutant concentration prediction result of the target region based on the pollutant concentrations of each second sub-region in the prediction time period.

[0091] In an embodiment of the present application, the region division module includes: A region division sub-module, configured to divide the target region based on the position information of multiple monitoring points in the target region through a Delaunay triangulation algorithm to generate a Delaunay triangulation graph; A Thiessen polygon generation sub-module, configured to generate Thiessen polygons based on the Delaunay triangulation graph to obtain multiple first sub-regions corresponding one by one to the multiple monitoring points.

[0092] In an embodiment of the present application, the region clustering module includes: A preprocessing sub-module, configured to preprocess the pollutant data of multiple monitoring points to obtain characteristic data corresponding to each of the multiple monitoring points; A clustering sub-module, configured to perform a clustering operation on the characteristic data corresponding to each of the multiple monitoring points to divide the multiple first sub-regions into multiple clustering clusters.

[0093] In an embodiment of the present application, the preprocessing sub-module includes: A data conversion unit, configured to convert the text data in the pollutant data into numerical data for the pollutant data of any monitoring point to obtain numerical data corresponding to the pollutant data; A feature processing unit, configured to perform feature processing on the numerical data to obtain characteristic data corresponding to the monitoring point.

[0094] In an embodiment of the present application, the clustering sub-module includes: A clustering unit, configured to perform a clustering operation on the characteristic data corresponding to each of the multiple monitoring points according to a preset clustering algorithm according to the number of target clustering clusters to divide the multiple first sub-regions into multiple clustering clusters with the same number as the number of target clustering clusters.

[0095] In an embodiment of the present application, the prediction model library includes multiple prediction models; among them, different prediction models are trained based on pollutant samples with different pollutant category characteristics, and different prediction models have different pollutant category characteristic labels; the model determination module includes: A similarity determination sub-module, configured to determine the similarity between the pollutant category features corresponding to the second sub-region and the pollutant category feature labels of multiple prediction models; A model determination sub-module, configured to determine the prediction model corresponding to the pollutant category feature label with the highest similarity as the target prediction model of the second sub-region.

[0096] In an embodiment of the present application, the concentration determination module includes: An influence weight determination sub-module, configured to determine the influence weight of each second sub-region based on the concentration influence factors of each second sub-region; the concentration influence factors include the regional area, topographic features, and pollutant emission intensity and meteorological data during the prediction period; A prediction result determination sub-module, configured to obtain the pollutant concentration prediction result of the target region based on the influence weight of each second sub-region and the pollutant concentration during the prediction period.

[0097] It should be noted that the specific implementation manner of the pollutant concentration prediction device 200 in the embodiment of the present application refers to the specific implementation manner of the pollutant concentration prediction method proposed in the first aspect of the embodiment of the present application, and will not be elaborated here.

[0098] In a third aspect, based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, on which an executable program is stored, and when the executable program is executed by a processor, it implements the pollutant concentration prediction method proposed in the first aspect of the present application.

[0099] It should be noted that the specific implementation manner of the computer-readable storage medium in the embodiment of the present application refers to the specific implementation manner of the pollutant concentration prediction method proposed in the first aspect of the embodiment of the present application, and will not be elaborated here.

[0100] In a fourth aspect, referring to Figure 3 , based on the same inventive concept, an embodiment of the present application provides an electronic device 300, including: A memory 301, configured to store an executable program; A processor 302; When the executable program is executed by the processor 302, it implements the pollutant concentration prediction method proposed in the first aspect of the present application.

[0101] It should be noted that the specific implementation manner of the electronic device 300 in the embodiment of the present application refers to the specific implementation manner of the pollutant concentration prediction method proposed in the first aspect of the embodiment of the present application, and will not be elaborated here.

[0102] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, apparatuses, or computer program products. Therefore, the embodiments of the present invention can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0103] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0106] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0107] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0108] The above has introduced in detail a method, device, storage medium and electronic device for predicting pollutant concentration provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for predicting pollutant concentration, characterized in that: The method comprises: Based on the location information of the plurality of monitoring points in the target area, the target area is divided to obtain a plurality of first sub-areas corresponding to the plurality of monitoring points one by one; Based on the pollutant data of the plurality of monitoring points, clustering the plurality of the first sub-areas to obtain a plurality of clusters; different clusters correspond to different pollutant category characteristics; For any of the clusters, merging the first sub-regions in the cluster to obtain a plurality of second sub-regions corresponding to the plurality of clusters one by one; For any second sub-region, based on the pollutant category characteristics corresponding to the second sub-region, determine a target prediction model for the second sub-region in a preset prediction model library; For any second sub-region, historical pollutant data of each monitoring point in the second sub-region and meteorological data of the forecast time period are input into the target forecast model of the second sub-region, and the pollutant concentration of the second sub-region in the forecast time period is output; Based on the pollutant concentrations of each of the second sub-areas in the prediction time period, a pollutant concentration prediction result of the target area is obtained.

2. A pollutant concentration prediction method according to claim 1, characterized in that: Based on the location information of multiple monitoring points in the target area, the target area is divided to obtain multiple first sub-areas corresponding to the multiple monitoring points one by one, including: Based on the location information of multiple monitoring points in the target area, the target area is divided by a Delaunay triangulation algorithm to generate a Delaunay triangulation graph; Based on the Delaunay triangulation graph, Thiessen polygons are generated to obtain a plurality of first sub-areas corresponding one-to-one to the plurality of monitoring points.

3. A pollutant concentration prediction method according to claim 1, characterized in that: Based on the pollutant data of the plurality of monitoring points, the plurality of the first sub-areas are clustered to obtain a plurality of cluster clusters, including: Preprocessing the pollutant data of the plurality of monitoring points to obtain characteristic data corresponding to each of the plurality of monitoring points; A clustering operation is performed on the characteristic data corresponding to each of the plurality of monitoring points to divide the plurality of first sub-areas into a plurality of cluster clusters.

4. A pollutant concentration prediction method according to claim 3, characterized in that: Preprocessing the pollutant data of the plurality of monitoring points to obtain characteristic data corresponding to each of the plurality of monitoring points includes: For the pollutant data of any of the monitoring points, converting the text data in the pollutant data into numerical data to obtain numerical data corresponding to the pollutant data; Perform feature processing on the numerical data to obtain feature data corresponding to the monitoring point.

5. A pollutant concentration prediction method according to claim 3, characterized in that: Performing a clustering operation on the characteristic data corresponding to each of the plurality of monitoring points to divide the plurality of first sub-areas into a plurality of cluster clusters includes: A preset clustering algorithm is used to perform a clustering operation on the feature data corresponding to each of the plurality of monitoring points according to the target number of clustering clusters, so as to divide the plurality of first sub-areas into a plurality of clustering clusters whose number is the same as the target number of clustering clusters.

6. A pollutant concentration prediction method according to claim 1, characterized in that: The prediction model library includes multiple prediction models; wherein different prediction models are obtained by training pollutant samples based on different pollutant category characteristics, and different prediction models have different pollutant category characteristic labels; Based on the pollutant category characteristics corresponding to the second sub-region, determining a target prediction model for the second sub-region in a preset prediction model library includes: Determine the similarity between the pollutant category feature corresponding to the second sub-region and the pollutant category feature labels of multiple prediction models; The prediction model corresponding to the pollutant category feature label with the highest similarity is determined as the target prediction model for the second sub-region.

7. A pollutant concentration prediction method according to claim 1, characterized in that: Obtaining a pollutant concentration prediction result of the target area based on the pollutant concentration of each of the second sub-areas in the prediction time period, including: Determining the influence weight of each second sub-region based on the concentration influence factor of each second sub-region; the concentration influence factor includes the area of ​​the region, the terrain characteristics, and the pollutant emission intensity and meteorological data in the predicted time period; Based on the influence weights of each of the second sub-areas and the pollutant concentration in the predicted time period, a pollutant concentration prediction result of the target area is obtained.

8. A pollutant concentration prediction device, characterized in that: The device comprises: A region division module, used for dividing the target region based on the location information of the plurality of monitoring points in the target region to obtain a plurality of first sub-regions corresponding to the plurality of monitoring points one by one; A regional clustering module, used for clustering the plurality of first sub-regions based on the pollutant data of the plurality of monitoring points to obtain a plurality of clusters; different clusters correspond to different pollutant category characteristics; A region merging module, configured to merge the first sub-regions in any of the clusters to obtain a plurality of second sub-regions corresponding to the plurality of clusters; A model determination module, for determining, for any second sub-region, a target prediction model for the second sub-region in a preset prediction model library based on the pollutant category characteristics corresponding to the second sub-region; A concentration prediction module, for inputting historical pollutant data of each monitoring point in the second sub-region and meteorological data of a prediction period into a target prediction model of the second sub-region for any second sub-region, and outputting the pollutant concentration of the second sub-region in the prediction period; The concentration determination module is used to obtain the pollutant concentration prediction result of the target area based on the pollutant concentration of each of the second sub-areas in the prediction time period.

9. A computer-readable storage medium having an executable program stored thereon, characterized in that: When the executable program is executed by a processor, the pollutant concentration prediction method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: A memory for storing an executable program; processor; When the executable program is executed by the processor, the pollutant concentration prediction method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Air quality prediction method and device based on transfer learning

    CN111461410A

  • Large-range land subsidence space-time prediction method and system based on deep learning

    CN112446559A

  • Air quality forecasting method and system

    CN112465243A

  • Pollutant early warning method and system and storage medium

    CN114118756A

  • Air quality and meteorological condition prediction method and device, equipment and storage medium

    CN115983329A