Air pollution early warning method and system
By combining meteorological data and city information, using K-Means clustering and deep convolutional neural networks to build an air pollution prediction model, the problem of inaccurate early warning results in the existing technology is solved, and more accurate and flexible early warning of air pollution is achieved.
Patent Information
- Application Number
- CN202510397803.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing air pollution warning methods fail to fully consider changes in meteorological data and urban information, resulting in the warning results that are inconsistent with the actual pollution situation and insufficient accuracy.
By obtaining the meteorological data and urban information of the cities to be tested, using a pre-trained urban pollution prediction model, combining the K-Means clustering algorithm and deep convolutional neural network, a city pollution prediction model is constructed to predict and early warning of the air pollution index.
It improves the accuracy and flexibility of air pollution prediction, can adapt to changes in air quality in different cities and meteorological conditions, and provides personalized early warning services.
Smart Images

Figure CN120275576A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air pollution detection, and particularly to an air pollution warning method and system. Background Art
[0002] Air pollution refers to harmful substances present in the air that exceed the tolerance limit of the natural environment, seriously endangering human health, ecosystems, and the global climate. Air pollution is not only one of the core global environmental problems but also has a significant impact on social and economic development, public health, and the ecological environment. Air pollution warning is to collect and analyze air quality data, predict the changing trend of air pollution, and issue a warning in advance when the pollution reaches a certain concentration. Air pollution warning is crucial for aspects such as public health, environmental protection, emergency response, and government decision-making.
[0003] In an existing technology, this technology measures the concentrations of pollutants such as PM2.5, PM10, NOx, SO2, etc. through air quality monitoring stations, and through statistical analysis of historical data, identifies the changing trend of pollution, classifies the severity of pollution in combination with the air quality index, and issues a warning message when the monitored data is about to exceed the set warning threshold. This technology obtains real-time air quality data and predicts based on simple physical models or historical data, only considering air quality while ignoring the impact of urban pollution sources on air quality, or only collecting real-time air quality-related data within the city for prediction while ignoring the impact of factors outside the city on the air quality within the city, such as wind direction, wind force, humidity, etc.
[0004] In summary, the existing technology cannot fully consider the changes in meteorological data and urban information. Especially in the case of large geographical location differences, the warning result does not match the actual pollution situation, resulting in inaccurate warning results. Summary of the Invention
[0005] The present invention provides an air pollution warning method and system to solve the problem of inaccurate warning results caused by the mismatch between the warning result and the actual pollution situation.
[0006] In the first aspect, to solve the above technical problem, the present invention provides an air pollution warning method, including:
[0007] Obtain the meteorological data and urban information of the city to be measured, and input them into a pre-trained urban pollution prediction model to obtain an air pollution index;
[0008] According to the air pollution index, compare and judge with the air comprehensive pollution rule to obtain the air pollution level and issue a warning;
[0009] Among them, the training process of the urban pollution prediction model includes:
[0010] Obtain the historical meteorological data and historical air quality grades of each city;
[0011] According to the historical meteorological data, perform normalization processing on non-linear data to obtain linear meteorological data;
[0012] According to the linear meteorological data, use the K-Means clustering algorithm to extract features to obtain clustered meteorological data;
[0013] Based on the historical air quality grades and the clustered meteorological data as input data, construct a city pollution prediction model based on a deep convolutional neural network, train the city pollution prediction model, and determine that the training is completed after detecting that the loss of the model meets the conditions to obtain a trained city pollution prediction model;
[0014] Input the meteorological data and city information of the city to be measured into the trained city pollution prediction model to obtain the air pollution index.
[0015] In an alternative embodiment, the obtaining of the historical meteorological data and historical air quality grades of each city includes:
[0016] Download the historical meteorological data of each city through the city meteorological data platform;
[0017] Download and obtain the historical air quality grades through the environmental monitoring general station or the environmental monitoring platform.
[0018] In an alternative embodiment, the historical meteorological data includes temperature, humidity, wind direction, wind speed, and air pressure.
[0019] In an alternative embodiment, the performing of normalization processing on non-linear data according to the historical meteorological data to obtain linear meteorological data includes:
[0020] According to the historical meteorological data, calculate the maximum value and the minimum value;
[0021] According to the maximum value, the minimum value, and the historical meteorological data, perform min-max normalization processing to obtain linear meteorological data;
[0022] Among them, the calculation formula of the linear meteorological data is as follows:
[0023]
[0024] In the formula, X is the historical meteorological data, X min is the minimum value, X max is the maximum value, and X norm is the linear meteorological data.
[0025] In an alternative embodiment, extracting features from the linear meteorological data using the K-Means clustering algorithm to obtain clustered meteorological data includes:
[0026] Based on the linear meteorological data, using the silhouette coefficient method to obtain the optimal K value;
[0027] Based on the optimal K value, setting the number of clusters of K-Means to K, and randomly selecting K data as the initial cluster centers to obtain the initial cluster centers;
[0028] Based on the linear meteorological data and the initial cluster centers, calculating the distance from each data point to the cluster center to obtain the Euclidean distance;
[0029] Based on the initial cluster centers and the Euclidean distance, assigning each data point to the nearest cluster to obtain the initial cluster data;
[0030] Based on the initial cluster data, calculating the mean of all data points within each cluster data and updating the cluster centers to obtain the second cluster centers;
[0031] Based on the second cluster centers and the linear meteorological data, recalculating the Euclidean distance and assigning data to obtain the second cluster data;
[0032] Continue the next iteration. When the number of iterations is greater than or equal to the preset maximum number of iterations, it is determined that the iteration is completed to obtain the optimal cluster centers;
[0033] Based on the optimal cluster centers and the linear meteorological data, assigning data points according to the Euclidean distance to obtain the clustered meteorological data.
[0034] In an alternative embodiment, using the historical air quality grades and the clustered meteorological data as input data, constructing an urban pollution prediction model based on a deep convolutional neural network, training the urban pollution prediction model, and determining that the training is completed after the loss of the model meets the conditions to obtain the trained urban pollution prediction model, including:
[0035] Initializing the parameters of the deep convolutional neural network, where the parameters of the deep convolutional neural network include the convolutional layer size, the convolutional kernel size, the pooling strategy, and the activation function;
[0036] Based on the clustered meteorological data, dividing the data into training data and test data to obtain a training set and a test set;
[0037] Based on the training set, autonomously learning local features through the convolutional layer to obtain data features;
[0038] Based on the data features, reducing the feature space size by sampling in the pooling layer to obtain the main features;
[0039] Perform non - linear matching in the activation layer according to the described main features and the historical air quality levels to obtain a non - linear relationship;
[0040] According to the non - linear relationship and the test set, use the activation function to perform non - linear transformation on the test data to obtain the test pollution index;
[0041] According to the test pollution index and the historical air quality levels, perform test result judgment to determine whether the difference between the result obtained by the prediction model based on the meteorological data of the test set and the actual air quality level is less than a preset loss threshold, and use the gradient descent method to adjust and update the network weights and biases, optimize the model parameters, update and train the prediction model;
[0042] When the calculated loss value is less than the preset threshold, determine that the training is completed and obtain the urban pollution prediction model.
[0043] In an alternative embodiment, the obtaining of the optimal K value according to the linear meteorological data by using the silhouette coefficient method includes:
[0044] Initialize the value range of K and select a series of different numbers of clusters K;
[0045] According to the linear meteorological data and the different numbers of clusters K, calculate the silhouette coefficients under different K values to obtain the meteorological data silhouette coefficients under different K values;
[0046] According to the meteorological data silhouette coefficients, perform average value calculation to obtain the overall silhouette coefficients under different K values;
[0047] According to the overall silhouette coefficients, make a comparison and select the K value corresponding to the maximum overall silhouette coefficient as the optimal K value;
[0048] Among them, the formula for the silhouette coefficient is as follows:
[0049]
[0050] In the formula, s(i) is the silhouette coefficient of data point i, b(i) is the average distance from data point i to all points in the nearest other cluster, representing the inter - cluster separation, and a(i) is the average distance from data point i to all other data points within the same cluster, representing the intra - cluster compactness.
[0051] In a second aspect, the present invention provides an air pollution early warning system, including:
[0052] A data acquisition module for obtaining the meteorological data and urban information of the city to be measured;
[0053] A historical data module for obtaining the historical meteorological data and historical air quality levels of each city;
[0054] A model construction module, configured to construct a model based on the historical meteorological data and historical air quality levels to obtain an urban pollution prediction model;
[0055] A pollution calculation module, configured to input the meteorological data and urban information of the city to be measured into a pre-trained urban pollution prediction model to obtain an air pollution index;
[0056] A pollution warning module, configured to compare and judge according to the air pollution index with the air comprehensive pollution rules to obtain an air pollution level and display a warning.
[0057] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the air pollution warning method described in any one of the above is implemented.
[0058] In a fourth aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the air pollution warning method described in any one of the above.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] The present invention discloses an air pollution warning method, including obtaining the meteorological data and urban information of the city to be measured, inputting them into a pre-trained urban pollution prediction model to obtain an air pollution index; comparing and judging according to the air pollution index with the air comprehensive pollution rules to obtain an air pollution level and give a warning; wherein, the training process of the urban pollution prediction model includes obtaining the historical meteorological data and historical air quality levels of each city; performing normalization processing on non-linear data according to the historical meteorological data to obtain linear meteorological data; extracting features using the K-Means clustering algorithm according to the linear meteorological data to obtain clustered meteorological data; constructing an urban pollution prediction model based on the deep convolutional neural network with the historical air quality level and the clustered meteorological data as input data, training the urban pollution prediction model, and determining that the training is completed after detecting that the loss of the model meets the conditions to obtain a trained urban pollution prediction model; inputting the meteorological data and urban information of the city to be measured into the trained urban pollution prediction model to obtain an air pollution index.
[0061] This method transforms meteorological data from a non-linear state to a linear state through non-linear normalization of historical meteorological data, facilitating processing by deep learning models. Using the K-Means clustering algorithm to preprocess meteorological data can effectively reduce data noise and redundant information, extracting more representative and discriminative features. In addition, the training process of the deep convolutional neural network can gradually optimize model parameters, ensuring that the trained model has strong generalization ability and can adapt to air quality prediction under different cities and different meteorological conditions. By obtaining historical meteorological data and historical air quality levels of each city and combining deep learning and clustering algorithms, an efficient urban pollution prediction model is constructed. In particular, the K-Means clustering algorithm is used in the method to extract features from meteorological data, which can effectively identify air pollution patterns under different meteorological conditions. Then, through training with a deep convolutional neural network (DCNN), complex non-linear relationships can be automatically learned from meteorological data. In this way, the air pollution index can be predicted more accurately, especially in response to pollution changes under different meteorological conditions, avoiding the dependence on linear relationships in traditional methods. This improves the adaptability of the prediction model to pollution situations under various complex weather conditions and thus enhances the accuracy of air pollution index prediction.
[0062] Since this method takes the meteorological data and historical air quality levels of each city as inputs and combines clustering and deep learning models for training, it can provide customized air pollution predictions according to the specific climate conditions and pollution characteristics of different cities. The meteorological conditions in different cities may vary greatly, such as the differences between temperate and tropical climates, local weather factors, etc. Therefore, this urban pollution prediction model trained based on historical data can effectively adapt to these differences and improve the flexibility of early warnings. Thus, this method can perform personalized air pollution predictions for different cities without being restricted by regional differences. In this way, the early warning system can be precisely adjusted according to the meteorological characteristics and historical pollution situations of each city, improving the flexibility and wide applicability of the system. Brief Description of the Drawings
[0063] Figure 1 is a schematic flowchart of the air pollution early warning method provided by the first embodiment of the present invention;
[0064] Figure 2 is a schematic diagram of the training of the air pollution early warning model provided by the embodiment of the present invention;
[0065] Figure 3 is a schematic structural diagram of an air pollution early warning system provided by the second embodiment of the present invention. Detailed Embodiments
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0067] Referring to Figure 1 , the first embodiment of the present invention provides an air pollution warning method, including the following steps:
[0068] S1. Obtain the meteorological data and urban information of the city to be measured, input them into the pre-trained urban pollution prediction model, and obtain the air pollution index;
[0069] S2. Compare and judge according to the air pollution index with the air comprehensive pollution rule to obtain the air pollution level and give a warning.
[0070] In step S1, the meteorological data and urban information of the city to be measured are obtained, input into the pre-trained urban pollution prediction model, and the air pollution index is obtained.
[0071] It should be noted that for provincial capital cities, meteorological data can be obtained through meteorological monitoring stations, and there are multiple meteorological monitoring stations in provincial capital cities to provide real-time meteorological data. For cities without meteorological stations, remote sensing technology can be used to obtain meteorological data through satellite data. For example, satellite sensors can obtain data such as temperature, humidity, and precipitation, and provide real-time meteorological conditions through data processing. The meteorological data includes temperature, humidity, air pressure, precipitation, atmospheric stability, wind speed, and wind direction.
[0072] Temperature directly affects the density of air and air flow patterns. High temperature can cause an increase in the air flow velocity, which will lead to an accelerated diffusion rate of pollutants. In cold weather, the air is relatively stable, and pollutants are likely to accumulate. Humidity is the content of water vapor in the air, which directly affects the formation of aerosols and the diffusion of air pollutants. High humidity will increase the sedimentation of particulate matter in the air, affecting air quality; while low humidity will cause particulate matter to remain in the air, especially PM2.5. Wind speed affects the diffusion rate of pollutants in the air. Strong winds can help disperse pollutants and reduce the local pollution concentration; while when the wind speed is low, pollutants are more likely to accumulate. Wind direction determines the flow direction of pollutants, which has an important impact on the diffusion path and area of air pollutants. Air pressure has an important impact on air flow and meteorological systems. Lower air pressure will cause the air to rise, driving pollutants to diffuse upward; while high air pressure means stable air, and pollutants are difficult to diffuse, which will lead to the accumulation of pollutants. Precipitation, such as rain, snow, etc. can help remove pollutants in the air, especially particulate matter. Rainwater can reduce the concentration of pollutants such as PM2.5 and PM10 through the "washing effect". Atmospheric stability determines the vertical mixing ability of the air. When the atmosphere is stable, pollutants stay near the ground and are not easily dispersed, resulting in an increase in pollution concentration, while unstable atmospheric conditions contribute to the diffusion of pollutants.
[0073] Meteorological data plays an important role in the air pollution warning system. By obtaining and analyzing the meteorological data of the city to be measured in real time, it can help predict the diffusion trend of pollutants, issue warnings in a timely manner, and reduce the impact of air pollution on public health. Meteorological data not only provides a direct basis for predicting pollution, but also helps determine the timeliness, accuracy, and protective measures of air quality changes.
[0074] As Figure 2 shown, the step of obtaining the meteorological data and city information of the city to be measured in step S1 and inputting them into the pre-trained city pollution prediction model to obtain the air pollution index includes:
[0075] S11, obtaining the historical meteorological data and historical air quality grades of each city;
[0076] S12, performing normalization processing on the non-linear data according to the historical meteorological data to obtain linear meteorological data;
[0077] S13, extracting features according to the linear meteorological data by using the K-Means clustering algorithm to obtain clustered meteorological data;
[0078] S14, based on the historical air quality grades and the clustered meteorological data as input data, constructing a city pollution prediction model based on a deep convolutional neural network, training the city pollution prediction model, and determining that the training is completed after detecting that the loss of the model meets the conditions to obtain the trained city pollution prediction model;
[0079] S15, inputting the meteorological data and city information of the city to be tested into the trained city pollution prediction model to obtain the air pollution index.
[0080] In step S11, historical meteorological data and historical air quality levels of each city are obtained.
[0081] In one implementation, historical meteorological data of each city is downloaded through a city meteorological data platform;
[0082] Obtain historical air quality levels by downloading from the Environmental Monitoring Center or the Environmental Monitoring Platform.
[0083] It should be noted that historical meteorological data include temperature, humidity, air pressure, precipitation, sunshine duration, wind speed and wind direction. Historical air quality levels include air quality index, pollutant concentration and pollution level. Among them, the air quality index AQI is a standard for measuring air quality, including the concentration of pollutants such as PM2.5, PM10, NO2, SO2, CO and O3. Pollution level refers to the six levels of 0-50 (I), 51-100 (II), 101-150 (III), 151-200 (IV), 201-250 (V), and 251-300 (VI) according to the comprehensive air pollution index, which are represented by blue, yellow, orange, red, purple and brown, respectively, which are excellent, good, light pollution, moderate pollution, heavy pollution and severe pollution.
[0084] These historical data come from monitoring data from the past five years. The purpose of using historical data to train the air pollution prediction model is to allow the model to learn the complex relationship and changing rules between pollution indices. During the training process, the model will try to find the best mapping relationship between input data and output. In this embodiment, the mapping relationship between meteorological data and air quality levels is found. By combining historical meteorological data and historical air quality level data, the air pollution prediction model can make pollution predictions more accurately. Meteorological data helps understand the impact of factors such as air flow and humidity changes on the diffusion of pollutants, while historical air quality data provides historical trends in pollutant concentrations. The combination of the two can more accurately predict changes in air quality.
[0085] In step S12, nonlinear data is normalized based on the historical meteorological data to obtain linear meteorological data.
[0086] In one implementation, the historical meteorological data is sorted by size to obtain the maximum value and the minimum value;
[0087] Based on the maximum value, minimum value, and the historical meteorological data, perform min-max normalization to scale the range of the normalized data to a preset range, obtaining linear meteorological data;
[0088] Among them, the calculation formula for the linear meteorological data is as follows:
[0089]
[0090] In the formula, X is the historical meteorological data, X min is the minimum value, X max is the maximum value, and X norm is the linear meteorological data.
[0091] It should be noted that min-max normalization is a common data preprocessing method, especially in machine learning and data analysis, used to transform the original data into a unified range. For the air pollution warning system, historical meteorological data has different dimensions and ranges. For example, the temperature is between 0 - 40 °C, while the wind speed is between 0 - 10 m / s, which can cause certain features to dominate in model training. Min-max normalization effectively solves this problem by scaling the data to a unified range, such as [0, 1] or [-1, 1], making different features have similar scales and avoiding the problem that the large difference in the range of feature values affects the training effect of the model.
[0092] Min-max normalization in the air pollution warning system can unify the feature scales, ensuring that each meteorological feature has the same influence during model training. Accelerate model convergence, avoiding training instability caused by large differences in feature scales. Improve prediction accuracy, enabling the model to better capture the combined effects of multiple meteorological factors on air pollution. Optimize multi-source data fusion, ensuring that features from different data sources can effectively contribute to model prediction under a unified scale.
[0093] In step S13, based on the linear meteorological data, use the K-Means clustering algorithm to extract features, obtaining clustered meteorological data.
[0094] In one implementation, based on the linear meteorological data, use the silhouette coefficient method to obtain the optimal K value;
[0095] According to the optimal K value, set the number of clusters of K-Means to K, randomly select K data as the initial cluster centers, obtaining the initial cluster centers;
[0096] Based on the linear meteorological data and the initial cluster centers, calculate the distance from each data point to the cluster center, obtaining the Euclidean distance;
[0097] According to the initial cluster centers and the Euclidean distances, each data point is assigned to the nearest cluster to obtain initial cluster data;
[0098] According to the initial cluster data, calculate the mean value of all data points within each cluster data, update the cluster centers, and obtain the second cluster centers;
[0099] According to the second cluster centers and the linear meteorological data, recalculate the Euclidean distances and assign the data to obtain second cluster data;
[0100] Continue the next iteration. When the number of iterations is greater than or equal to the preset maximum number of iterations, it is determined that the iteration is completed, and the optimal cluster centers are obtained;
[0101] According to the optimal cluster centers and the linear meteorological data, assign data points according to the Euclidean distances to obtain clustered meteorological data.
[0102] It should be noted that in the embodiments of the present invention, the maximum number of iterations is set to 300 times, and it needs to be adjusted according to specific situations in actual applications. The K-Means clustering algorithm is a classic unsupervised learning algorithm, which is widely used in data mining and machine learning. Its main purpose is to divide a data set into several clusters, so that the data points within the same cluster are as similar as possible, while the data points between different clusters are quite different. In meteorological data analysis, the K-Means algorithm can be used to extract features from a large amount of meteorological data and perform clustering, so as to help us discover the potential structure of the data and conduct effective classification and prediction.
[0103] The K-Means algorithm can divide data into different clusters according to the similarity of meteorological data. In this way, we can reveal the potential patterns of meteorological data. For example, different weather conditions, such as sunny days, cloudy days, rainy days, etc., can be grouped to help us understand the distribution and diffusion characteristics of pollutants under different meteorological conditions. Meteorological data is high-dimensional and contains various features, such as temperature, humidity, wind speed, etc. K-Means clustering can simplify the complexity of these complex meteorological data by reducing the dimensions and summarizing them into several main cluster centers. Each cluster center represents a typical meteorological pattern and can effectively transform complex meteorological data into information that is easier to process. By clustering meteorological data, the K-Means algorithm can extract different weather conditions and climate characteristics and use them as features to input into an air pollution warning model.
[0104] Among them, obtaining the optimal K value according to the linear meteorological data by using the silhouette coefficient method includes:
[0105] Initialize the value range of K and select a series of different numbers of clusters K;
[0106] Calculate the silhouette coefficients for different values of K based on the linear meteorological data and the different number of clusters K, and obtain the meteorological data silhouette coefficients for different values of K.
[0107] Perform an average calculation based on the meteorological data silhouette coefficients to obtain the overall silhouette coefficients for different values of K.
[0108] Based on the overall silhouette coefficients, make a comparison and select the value of K corresponding to the maximum overall silhouette coefficient as the optimal K value.
[0109] Among them, the formula for calculating the silhouette coefficient is as follows:
[0110]
[0111] In the formula, s(i) is the silhouette coefficient of data point i, b(i) is the average distance from data point i to all points in the nearest other cluster, representing the separation between clusters, and a(i) is the average distance from data point i to all other data points within the same cluster, representing the compactness within the cluster.
[0112] It should be noted that the silhouette coefficient method is a commonly used method for evaluating the quality of clustering, which is used to quantify the compactness within the cluster and the separation between clusters of data points. Its value ranges between [-1, 1], and the larger the value, the better the clustering effect. The silhouette coefficient method is used to help select the optimal number of clustering clusters or evaluate the quality of the clustering result. The compactness within the cluster refers to the average distance between a data point and other data points within the same cluster. The smaller this value, the higher the similarity between the data point and other points within the cluster, and the tighter the internal structure of the cluster. The separation between clusters refers to the average distance from a data point to all points in the nearest other cluster. The larger this value, the greater the dissimilarity between the data point and other clusters, and the better the separation between clusters.
[0113] The silhouette coefficient method provides a quantitative evaluation criterion for the clustering result. By calculating the silhouette coefficient of each data point and the overall silhouette coefficient, the silhouette coefficient method can measure the compactness of the data within the cluster and the separation between clusters, helping us judge whether the clustering is reasonable. In K-Means clustering, selecting the number of clusters is an important decision. Using the silhouette coefficient method, we can select the optimal one by comparing the overall silhouette coefficients for different numbers of clusters, that is, select the number of clusters that can maximize the overall silhouette coefficient. Avoid too many or too few clusters. Too many clusters will lead to overfitting, and too few clusters will lead to over-simplifying the data structure. The silhouette coefficient method can help find a balance point to ensure that the clustering result has sufficient separation and compactness.
[0114] In air pollution prediction, K-Means clustering can be used to group air quality data under different meteorological conditions. By calculating the silhouette coefficient for different numbers of clusters, we can select the most suitable number of clusters, thereby effectively classifying and predicting air quality. Different meteorological conditions have different impacts on air pollution. Through clustering, we can group according to meteorological conditions and implement targeted pollution prediction and health warnings for different clusters. The silhouette coefficient method can help select the optimal number of clusters, ensuring that the data within each meteorological cluster has similar air quality characteristics, while the pollution characteristics between different clusters are quite different.
[0115] In step S14, using the historical air quality grades and the clustered meteorological data as input data, a city pollution prediction model is constructed based on a deep convolutional neural network, the city pollution prediction model is trained, and after detecting that the loss of the model meets the conditions, it is determined that the training is completed, and a trained city pollution prediction model is obtained.
[0116] In one implementation, the parameters of the deep convolutional neural network are initialized, where the parameters of the deep convolutional neural network include the size of the convolutional layer, the size of the convolutional kernel, the pooling strategy, and the activation function;
[0117] According to the clustered meteorological data, the data is divided into training data and test data to obtain a training set and a test set;
[0118] According to the training set, local features are autonomously learned through the convolutional layer to obtain data features;
[0119] According to the data features, the size of the feature space is reduced by sampling in the pooling layer to obtain main features;
[0120] According to the main features and the historical air quality grades, non-linear matching is performed in the activation layer to obtain a non-linear relationship;
[0121] According to the non-linear relationship and the test set, the test data is non-linearly transformed using the activation function to obtain a test pollution index;
[0122] According to the test pollution index and the historical air quality grades, a test result judgment is made to determine whether the difference between the result obtained by the prediction model based on the meteorological data of the test set and the actual air quality grade is less than a preset loss threshold, and the gradient descent method is used to adjust and update the network weights and biases, optimize the model parameters, update and train the prediction model;
[0123] When the calculated loss value is less than the preset threshold, it is determined that the training is completed, and a city pollution prediction model is obtained.
[0124] It should be noted that in the embodiments of the present invention, there is no requirement for the threshold value of the loss value. In practical applications, the setting of the threshold value needs to be adjusted and optimized according to specific circumstances. The deep convolutional neural network (DCNN) is a deep learning model widely used in computer vision, natural language processing, and other fields. Its structure mainly includes multiple convolutional layers, pooling layers, fully connected layers, and activation functions, etc., which can automatically extract important features from the input data and perform non-linear mapping. Through layer-by-layer progressive feature extraction and information processing, DCNN can learn complex patterns from the original data and is particularly suitable for processing high-dimensional data such as images and speech.
[0125] The role of the deep convolutional neural network (DCNN) in the air pollution warning model is very crucial, especially in the prediction and analysis of meteorological data and air quality data. The deep convolutional neural network can automatically extract effective features from the input data. For example, meteorological data includes multiple factors such as temperature, humidity, wind speed, etc. DCNN can automatically discover the complex relationships and patterns among these features, reducing human intervention and improving the accuracy and efficiency of the model. The convolutional layer can automatically learn the local features in the data and is suitable for the local impact of different meteorological factors on air pollution in meteorological data. Through the deep network structure, DCNN can integrate local features and learn the global patterns of the data, improving the prediction accuracy.
[0126] The relationship between air pollution and meteorological data is non-linear, and it is difficult for traditional linear regression models to accurately capture these complex relationships. The deep convolutional neural network introduces non-linearity through non-linear activation functions and can fit the complex, non-linear relationship between meteorological data and air pollution. The activation function helps the neural network perform complex non-linear mapping and accurately represent the impact of meteorological conditions on pollutant concentration. DCNN can perform trend prediction based on the non-linear relationship, providing accurate information for future air quality warnings.
[0127] By stacking multiple convolutional layers and pooling layers, DCNN can process feature information at multiple levels. Each factor in meteorological data, such as temperature, humidity, wind speed, etc., has different levels of impact on air quality. DCNN can effectively extract important information step by step from low-level features to high-level features, thereby helping to improve the accuracy of air pollution prediction. By combining meteorological data with historical air quality data, DCNN can learn complex patterns and provide accurate air quality predictions. For example, when certain meteorological factors, such as high temperature and low wind speed, cause pollutant accumulation, the model can identify and issue warnings in a timely manner. By learning from historical meteorological data and air quality data, DCNN can predict air pollution conditions in the next few hours or days. By learning the non-linear relationship between air quality levels and meteorological factors, DCNN can make more accurate predictions of pollution indices.
[0128] The role of the deep convolutional neural network (DCNN) in the air pollution warning system is crucial. Through automatic feature extraction, nonlinear modeling, and multi-level feature learning, DCNN can effectively capture the complex relationship between meteorological data and air pollution, improving the accuracy and timeliness of air pollution prediction. Based on meteorological data, DCNN can not only make short-term pollution predictions, but also enhance the intelligence and adaptive capabilities of the warning system, helping the government and the public better cope with air pollution problems and protect public health.
[0129] In summary, the present invention discloses an air pollution warning method, which includes obtaining meteorological data and city information of a city to be measured, inputting them into a pre-trained city pollution prediction model to obtain an air pollution index; comparing and judging the air pollution index with the air comprehensive pollution rule to obtain an air pollution level and issue a warning; wherein, the training process of the city pollution prediction model includes obtaining historical meteorological data and historical air quality levels of each city; performing normalization processing on nonlinear data according to the historical meteorological data to obtain linear meteorological data; extracting features using the K-Means clustering algorithm according to the linear meteorological data to obtain clustered meteorological data; constructing a city pollution prediction model based on the deep convolutional neural network with the historical air quality level and the clustered meteorological data as input data, training the city pollution prediction model, and determining that the training is completed after detecting that the loss of the model meets the conditions to obtain a trained city pollution prediction model; inputting the meteorological data and city information of the city to be measured into the trained city pollution prediction model to obtain an air pollution index.
[0130] This method transforms meteorological data from a non-linear state to a linear state through non-linear normalization of historical meteorological data, facilitating processing by deep learning models. Using the K-Means clustering algorithm to preprocess meteorological data can effectively reduce data noise and redundant information, extracting more representative and distinguishable features. In addition, the training process of the deep convolutional neural network can gradually optimize model parameters, ensuring that the trained model has strong generalization ability and can adapt to air quality prediction under different cities and meteorological conditions. By obtaining the historical meteorological data and historical air quality levels of each city and combining deep learning and clustering algorithms, an efficient urban pollution prediction model is constructed. In particular, the K-Means clustering algorithm is used in the method to extract features from meteorological data, which can effectively identify air pollution patterns under different meteorological conditions. Then, through training with a deep convolutional neural network (DCNN), complex non-linear relationships can be automatically learned from meteorological data. In this way, the air pollution index can be predicted more accurately, especially in response to pollution changes under different meteorological conditions, avoiding the dependence on linear relationships in traditional methods. Thus, the adaptability of the prediction model to pollution situations under various complex weather conditions is improved, and the accuracy of air pollution index prediction is further enhanced.
[0131] Since this method takes the meteorological data and historical air quality levels of each city as inputs and combines clustering and deep learning models for training, it can provide customized air pollution predictions according to the specific climate conditions and pollution characteristics of different cities. The meteorological conditions in different cities may vary greatly, such as the differences between temperate and tropical climates, local weather factors, etc. Therefore, this urban pollution prediction model trained based on historical data can effectively adapt to these differences and improve the flexibility of early warnings. Hence, this method can perform personalized air pollution predictions for different cities without being restricted by regional differences. In this way, the early warning system can be precisely adjusted according to the meteorological characteristics and historical pollution situations of each city, improving the flexibility and wide applicability of the system.
[0132] Refer to Figure 2 , the second embodiment of the present invention provides an air pollution early warning system, including:
[0133] A data acquisition module for obtaining the meteorological data and city information of the city to be measured;
[0134] A historical data module for obtaining the historical meteorological data and historical air quality levels of each city;
[0135] A model construction module for constructing a model based on the historical meteorological data and historical air quality levels to obtain an urban pollution prediction model;
[0136] A pollution calculation module, configured to input the meteorological data and city information of the city to be measured into a pre-trained urban pollution prediction model to obtain an air pollution index;
[0137] A pollution warning module, configured to compare and judge the air pollution index with the air comprehensive pollution rules to obtain an air pollution level and display a warning.
[0138] It should be noted that an air pollution warning device provided by an embodiment of the present invention is used to execute all process steps of an air pollution warning method in the above embodiment. The working principles and beneficial effects of the two correspond one by one, so they will not be elaborated here.
[0139] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as an air pollution warning program. When the processor executes the computer program, the steps in the above embodiments of various air pollution warning methods are implemented, such as Figure 1 the step S11 shown. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above device embodiments are implemented, such as the model construction module.
[0140] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.
[0141] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a smart tablet. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above components are only examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than the above, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0142] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines.
[0143] The memory can be used to store the computer program and / or module. The processor realizes various functions of the electronic device by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0144] Among them, if the modules / units integrated in the electronic device are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0145] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative effort.
[0146] In the above specific embodiments, the purpose, technical solution, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An air pollution warning method, characterized in that, Executed by a computer, including: Obtain the meteorological data and city information of the city to be measured, input them into the pre-trained urban pollution prediction model, and obtain the air pollution index; According to the air pollution index, compare and judge with the air comprehensive pollution rule to obtain the air pollution level and give an early warning; Among them, the training process of the urban pollution prediction model includes: Obtain the historical meteorological data and historical air quality levels of each city; According to the historical meteorological data, perform normalization processing on the non-linear data to obtain linear meteorological data; According to the linear meteorological data, use the K-Means clustering algorithm to extract features to obtain clustered meteorological data; Based on the historical air quality level and the clustered meteorological data as input data, build an urban pollution prediction model based on a deep convolutional neural network, train the urban pollution prediction model, and determine that the training is completed after detecting that the loss of the model meets the conditions, and obtain the trained urban pollution prediction model; Input the meteorological data and city information of the city to be measured into the trained urban pollution prediction model to obtain the air pollution index.
2. The air pollution warning method according to claim 1, wherein The obtaining of the historical meteorological data and historical air quality levels of each city includes: Download the historical meteorological data of each city through the urban meteorological data platform; Download and obtain the historical air quality levels through the national environmental monitoring center or the environmental monitoring platform.
3. The air pollution warning method according to claim 1, characterized in that, The historical meteorological data includes temperature, humidity, wind direction, wind speed, and air pressure.
4. The air pollution warning method according to claim 1, wherein According to the historical meteorological data, performing normalization processing on the non-linear data to obtain linear meteorological data includes: According to the historical meteorological data, perform size sorting to obtain the maximum value and the minimum value; According to the maximum value, the minimum value, and the historical meteorological data, perform min-max normalization processing to scale the value range of the normalized data to a preset range to obtain linear meteorological data; Among them, the calculation formula of the linear meteorological data is as follows: Wherein, X is historical meteorological data, X min is the minimum value, X max is the maximum value, X norm is linear meteorological data.
5. The air pollution warning method according to claim 1, characterized in that According to the linear meteorological data, using the K-Means clustering algorithm to extract features to obtain clustered meteorological data includes: According to the linear meteorological data, use the silhouette coefficient method to obtain the optimal K value; According to the optimal K value, set the number of clusters of K-Means to K, and randomly select K data as the initial cluster centers to obtain the initial cluster centers; According to the linear meteorological data and the initial cluster centers, calculate the distance from each data point to the cluster center to obtain the Euclidean distance; According to the initial cluster centers and the Euclidean distance, assign each data point to the nearest cluster to obtain the initial cluster data; According to the initial cluster data, calculate the mean value of all data points within each cluster data, update the cluster center, and obtain the second cluster center; According to the second cluster center and the linear meteorological data, recalculate the Euclidean distance and assign data to obtain the second cluster data; Continue the next iteration. When the number of iterations is greater than or equal to the preset maximum number of iterations, it is determined that the iteration is completed, and the optimal cluster center is obtained; According to the optimal cluster center and the linear meteorological data, allocate data points according to the Euclidean distance to obtain clustered meteorological data.
6. The air pollution warning method according to claim 1, characterized in that, Taking the historical air quality grades and the clustered meteorological data as input data, constructing an urban pollution prediction model based on a deep convolutional neural network, training the urban pollution prediction model, and determining that the training is completed after detecting that the loss of the model meets the conditions, to obtain a trained urban pollution prediction model, including: Initializing the parameters of the deep convolutional neural network, where the parameters of the deep convolutional neural network include the convolutional layer size, the convolutional kernel size, the pooling strategy, and the activation function; According to the clustered meteorological data, dividing the data into training data and test data to obtain a training set and a test set; According to the training set, autonomously learning local features through the convolutional layer to obtain data features; According to the data features, reducing the feature space size by sampling in the pooling layer to obtain main features; According to the main features and the historical air quality grades, performing non-linear matching in the activation layer to obtain a non-linear relationship; According to the non-linear relationship and the test set, performing non-linear transformation on the test data using the activation function to obtain a test pollution index; According to the test pollution index and the historical air quality grades, making a test result judgment, determining whether the difference between the result obtained by the prediction model according to the meteorological data of the test set and the actual air quality grade is less than a preset loss threshold, and using the gradient descent method to adjust and update the network weights and biases, optimizing the model parameters, updating and training the prediction model; When the calculated loss value is less than the preset threshold, determining that the training is completed to obtain an urban pollution prediction model.
7. The air pollution warning method according to claim 5, characterized in that According to the linear meteorological data, using the silhouette coefficient method to obtain the optimal K value, including: Initializing the value range of K and selecting a series of different numbers of clusters K; According to the linear meteorological data and the different numbers of clusters K, calculating the silhouette coefficients under different K values to obtain the meteorological data silhouette coefficients under different K values; According to the meteorological data silhouette coefficients, performing an average value calculation to obtain the overall silhouette coefficients under different K values; According to the overall silhouette coefficients, making a comparison and selecting the K value corresponding to the maximum overall silhouette coefficient as the optimal K value; Among them, the silhouette coefficient calculation formula is as follows: In the formula, s(i) is the silhouette coefficient of data point i, b(i) is the average distance from data point i to all points in the nearest other cluster, representing the inter-cluster separation degree, and a(i) is the average distance from data point i to all other data points within the same cluster, representing the intra-cluster compactness.
8. An air pollution warning system, characterized in that, Including: A data acquisition module for obtaining the meteorological data and urban information of the city to be measured; A historical data module for obtaining the historical meteorological data and historical air quality grades of each city; A model construction module for constructing a model according to the historical meteorological data and historical air quality grades to obtain an urban pollution prediction model; A pollution calculation module for inputting the meteorological data and urban information of the city to be measured into the pre-trained urban pollution prediction model to obtain an air pollution index; A pollution warning module for comparing and judging according to the air pollution index and the air comprehensive pollution rules to obtain the air pollution level and display a warning.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the air pollution warning method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the air pollution warning method according to any one of claims 1 to 7.