Industrial park weather typing method and device integrating self-organizing mapping neural network and k-means algorithm

By integrating self-organized mapping neural network and k-means algorithm, the meteorological data of industrial parks are dimensionality reduction and cluster analysis, and the weather classification of high concentrations of ozone is identified, which solves the problem of low prediction accuracy of ozone pollution in industrial parks and achieves rapid prediction of small-scale ozone pollution.

CN120217859APending Publication Date: 2025-06-27ZHEJIANG UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510295669.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In industrial parks, accurate prediction of ozone pollution still faces the problems of high data requirements and low prediction accuracy, especially in the absence of small-scale ozone pollution simulation models.

Method used

The integrated self-organized mapping neural network (SOM) and k-means algorithm are used to perform dimensionality reduction and cluster analysis on the meteorological data of industrial parks to identify weather classification with high concentrations of ozone.

Benefits of technology

It realizes rapid prediction of small-scale ozone pollution in industrial parks, and improves the accuracy and efficiency of ozone pollution prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217859A_ABST
    Figure CN120217859A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial park weather typing method and device integrating a self-organizing mapping neural network and a k-means algorithm, and the method comprises the steps: firstly, constructing an initial data set through air quality monitoring data and ERA5 meteorological reanalysis data, converting the data into a numerical value type, and carrying out the standardization, so as to guarantee that the feature data has the distribution with the mean value being 0 and the standard deviation being 1; thirdly, initializing a self-organizing mapping (SOM) model, and performing dimension reduction on the data by setting parameters such as grid size, sigma and learning rate so as to obtain two-dimensional feature representation; then, initializing a k-means clustering model, setting a clustering number and a random seed, and performing clustering analysis on the feature data after dimension reduction by using a fitpredict () method so as to obtain a clustering label of each sample; in the evaluation and visualization stage, evaluation indexes such as SSE, MSE, RMSE, a contour coefficient and a Clinski-Harabasz index are calculated so as to evaluate the clustering effect and the separation degree. And finally, adding a clustering label into the original data, and visually analyzing the distribution of each feature under different clusters by using a violin chart, thereby providing a visual basis for further analysis. According to the method, through integration of the SOM and the k-means model, the precision of ozone concentration weather typing of the small-scale industrial park is remarkably improved, and the high-concentration ozone weather is divided more finely and more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of prevention and control of air pollution, and particularly to a method and device for classifying weather types in industrial parks by integrating a self-organizing mapping neural network and a k-means algorithm. Background Art

[0002] In China, the ozone pollution shows a trend of fluctuating increase, which endangers people's health and ecological security. A series of documents such as the "Action Plan for the Prevention and Control of Ozone Pollution" and the "Guiding Opinions on Further Optimizing the Response Mechanism for Severe Pollution Weather" issued by the Ministry of Ecology and Environment have put forward higher requirements for the prevention and control of air pollution. The accurate prediction and forecasting of ozone pollution are important means to actively respond to ozone pollution and have increasingly attracted national attention. However, the formation of ozone pollution is jointly affected by meteorology and pollutants, resulting in problems such as high data requirements and low prediction accuracy in ozone pollution concentration prediction. Especially in industrial areas with high pollutant emissions, the ozone formation mechanism is not yet clear. In the absence of a small-scale ozone pollution simulation model, accurate and efficient ozone prediction has become the focus of ozone pollution prevention and control in industrial parks.

[0003] The application of air quality model simulation ozone pollution simulation technology is extensive. For example, CMAQ (Community Multiscale Air Quality) simulates the ambient ozone concentration through the atmospheric chemical process of pollutant emissions under the meteorological field; WRF-Chem (Weather Research and Forecasting model coupled with Chemistry) can simulate the emissions, transportation, mixing and chemical transformation of trace gases and aerosols. Since weather conditions directly affect the transmission and diffusion of ozone precursors VOCs and NOx after being emitted into the atmosphere, and at the same time affect the photochemical reaction and indirectly affect the formation and distribution of ozone, meteorological elements make a significant contribution to the formation of ozone pollution. The accuracy of air quality models in prediction is still insufficient. In particular, limited by the accuracy of emission inventories and the limited understanding of atmospheric chemical processes, it is difficult to simulate the formation process of air pollution in small-scale areas.

[0004] The statistical model method usually uses machine learning algorithms to achieve high-performance air pollution prediction. O3 pollution is significantly affected by large-scale meteorological conditions. The self-organizing mapping (SOM) method has been applied to the classification and visualization of meteorological data, and k-means clustering performs well in finding patterns and structures in data without predefined labels. The integrated method of SOM and k-means performs well in predicting ozone pollution in small-scale industrial parks. Summary of the Invention

[0005] The present invention mainly solves the problem of the lack of a method for classifying ozone pollution weather types in industrial parks, and provides a clustering recognition method and device based on the integration of SOM and k-means.

[0006] The present invention uses the air quality monitoring stations in industrial parks and the meteorological data supplemented by ECMWF Reanalysis v5 (ERA5) as input files; preprocesses the meteorological data set, and then reduces the dimension of the preprocessed data set through the SOM algorithm to obtain SOM features; then, inputs the reduced data set into the k-means algorithm for clustering, and identifies the weather types of high-concentration ozone through the ozone concentration distribution under the clusters.

[0007] The first aspect of the present invention relates to a method for classifying the weather types in industrial parks by integrating a self-organizing mapping neural network and the k-means algorithm, including the following steps:

[0008] S1, construction and preprocessing of the initial data set. Using the continuous monitoring data of the air monitoring stations in industrial parks and the meteorological data of the same period on the official website of ERA5 meteorological data, construct an initial data set with meteorological elements as variables; after missing value processing, standardize each column of variable data to construct a standardized input data set.

[0009] S1.1, data acquisition. Extract the continuous monitoring data during the simulation period from the air quality monitoring stations in industrial parks, including air temperature, atmospheric pressure, wind direction, relative humidity, wind speed, and ozone concentration; the supplementary data of ERA5 during the same simulation period is searched and downloaded through the official website of ERA5 meteorological data (https: / / cds.climate.copernicus.eu / ), including the east-west wind speed component u10 at a height of 10 meters, the north-south wind speed component v10 at a height of 10 meters, surface solar radiation SSRD, boundary layer height BLH, high cloud cover HCC, low cloud cover LCC, and middle cloud cover MCC.

[0010] S1.2, construction of the initial data set. Using meteorological elements as variables, arranging the values of meteorological elements at the same moment in order as a row, and arranging each row as a sample in time order to form an initial data set;

[0011] S1.3, data preprocessing. Delete the rows with missing values in the initial data set to ensure data integrity; standardize the values of each column of variables so that each variable has a distribution with a mean of 0 and a standard deviation of 1, and construct a standardized input data set.

[0012] In particular, the simulation period described in step S1.1 should be no less than 3 consecutive months;

[0013] In particular, the values of the meteorological elements described in step S1.2 are arranged in order, and the order of the meteorological elements should be kept consistent.

[0014] S2. Dimensionality reduction of the dataset based on the SOM model. Using the standardized input dataset as the model input, train and optimize the SOM dimensionality reduction model to obtain the feature dataset.

[0015] S2.1. SOM model initialization. Set the grid size and other parameters, and set the randomly initialized weights.

[0016] S2.2. SOM model training. Use the standardized input data to train the SOM model, and through iterative adjustment of the model weights, realize the mapping of the data.

[0017] S2.3. Obtain SOM features. Calculate the winning nodes of each data sample in the SOM grid and output them as new features to form the feature dataset.

[0018] In particular, for the SOM model initialization described in step S2.1, set the grid size of the SOM to 4×4, that is, a two-dimensional grid of 16 nodes. Other parameters include the input_len parameter, sigma parameter, learning_rate parameter, and random_seed parameter of MiniSom. Among them, the input_len parameter is set to the number of variables of the input data, that is, the number of meteorological elements; the sigma parameter and learning_rate parameter are set to default values; the random_seed parameter is set to 42 to ensure the repeatability of the model.

[0019] In particular, for the randomly initialized weights described in step S2.1, use the random_weights_init() method to set the randomly initialized weights of the SOM model.

[0020] In particular, for the SOM model training described in step S2.2, use the train_random() method to train the model, and set the number of training iterations to no less than 50000.

[0021] In particular, for the SOM features described in step S2.3, use the winner() method to obtain the coordinates of the winning nodes corresponding to each data sample, and convert the coordinates of the winning nodes into a two-dimensional feature array for subsequent k-means clustering.

[0022] S3. Simulation and optimization of ozone concentration by the k-means model. Using the feature dataset as the input, train and optimize the k-means clustering model to obtain the clustering labels, and evaluate the effectiveness of the clustering results and visualize their distribution.

[0023] S3.1, Initialize the k-means model. Initialize the clustering model using the 'k-means' class. It is recommended to preset the number of clusters n_clusters = 9, indicating that the data will be divided into 9 clusters; set random_state = 42 to ensure the reproducibility of the results.

[0024] S3.2, Train the k-means model and obtain the clustering labels. The k-means algorithm clusters by minimizing the distance from the samples to their respective cluster centers; use the fit_predict() method to cluster the som_features to obtain the clustering labels for each sample.

[0025] S3.3, Evaluate the clustering results. Calculate evaluation metrics such as the Sum of Squared Errors (SSE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE) to evaluate the clustering results; use silhouette_score() to calculate the silhouette coefficient to evaluate the separation of the clusters; use calinski_harabasz_score() to calculate the Calinski-Harabasz index to evaluate the effectiveness of the clustering.

[0026] S3.4, Analyze the clustering results. Plot violin plots of each meteorological element under each cluster respectively to view the distribution of each variable under different clusters; use sns.violinplot() to plot the distribution map of ozone concentration and group according to different clustering labels.

[0027] In particular, the number of clusters described in step S3.1 needs to be adjusted according to the data characteristics. The recommended adjustment range is 6 - 10 classes, and the most suitable number of clusters is selected according to the clustering result evaluation metrics in S3.3.

[0028] The second aspect of the present invention relates to an industrial park weather classification prediction device integrating a self-organizing map neural network and the k-means algorithm, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the industrial park weather classification method of the integrated self-organizing map neural network and the k-means algorithm of the present invention.

[0029] The advantages of the present invention are: By integrating the advantages of the SOM and k-means algorithms in dimensionality reduction and clustering analysis respectively, a meteorological classification method for ozone pollution prediction is established, realizing the rapid prediction of small-scale ozone pollution in industrial parks. Brief Description of the Drawings

[0030] Figure 1 This is the flowchart of the method of the present invention.

[0031] Figure 2 This is a schematic diagram of the initial data set of the method of the present invention.

[0032] Figure 3 This is a schematic diagram of the standardized input data set of the method of the present invention.

[0033] Figure 4 This is an example diagram of the evaluation index results of the clustering results of the method of the present invention.

[0034] Figure 5 This is an example of a weather classification violin plot of the method of the present invention, where Figure 5 part (a) shows the distribution of 9 weather classifications and the corresponding ozone concentration ranges, Figure 5 part (b) shows the temperature distribution, Figure 5 part (c) shows the air pressure distribution, Figure 5 part (d) shows the wind direction distribution, Figure 5 part (e) shows the humidity distribution, Figure 5 part (f) shows the wind speed distribution, Figure 5 part (g) shows the solar radiation distribution, Figure 5 part (h) shows the boundary layer height distribution, Figure 5 part (i) shows the north-south wind speed component distribution at a height of 10 meters, Figure 5 part (j) shows the east-west wind speed component distribution at a height of 10 meters, Figure 5 part (k) shows the low cloud cover distribution, Figure 5 part (l) shows the high cloud cover distribution, Figure 5 part (m) shows the middle cloud cover distribution.

[0035] Figure 6 This is a schematic diagram of the weather classification category output of the method of the present invention.

[0036] Figure 7 This is the device diagram of the present invention. Detailed implementation manners

[0037] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0038] Example 1

[0039] Refer to Figures 1-6 , in this example, a clustering method integrating SOM and k-means is used to analyze the weather classification of the meteorological data of an industrial park throughout 2023, and a method for weather classification of the industrial park integrating a self-organizing mapping neural network and the k-means algorithm is provided. The specific steps are as follows:

[0040] 1. Data preparation and preprocessing

[0041] Extract the data of the air quality monitoring stations in the case park in 2023, including air temperature, atmospheric pressure, wind direction, relative humidity, wind speed and ozone concentration; at the same time, search and download the corresponding meteorological data in 2023 on the ERA5 meteorological data official website (https: / / cds.climate.copernicus.eu / ), including the east-west wind speed component u10 at 10 meters height, the north-south wind speed component v10 at 10 meters height, surface solar radiation SSRD, boundary layer height BLH, high cloud cover HCC, low cloud cover LCC, and medium cloud cover MCC. Then, perform numerical conversion and standardization on the data to ensure that all feature data has a distribution with a mean of 0 and a standard deviation of 1, and construct a standardized input data set. This step is crucial for data preprocessing because standardization can eliminate the dimensional differences between different features, enabling subsequent dimensionality reduction algorithms to process data more effectively. When specifically implemented, the pd.to_numeric() method can be used to convert the data into numeric type, and the StandardScaler() is used to standardize the data.

[0042] 2. SOM Dimensionality Reduction

[0043] Use the SOM model to perform dimensionality reduction on the data. To initialize the SOM model, a series of parameters need to be set, including grid size, sigma, and learning rate, etc. In this embodiment, the grid size is set to 4×4, that is, a two-dimensional grid of 16 nodes, and the default sigma (the parameter affecting the learning range) and learning_rate (learning rate) are set. In addition, to ensure the reproducibility of the model, the random seed random_seed = 42 can be set. After initialization, use the random_weights_init() method to randomly initialize the weights of the SOM model to ensure that the model can explore a larger data space during training. Finally, train the model by randomly sampling samples, and use the train_random() method for model training, and the number of training iterations is set to 50000. The model continuously adjusts the weights of the grid nodes during training, thus achieving an effective mapping from high-dimensional data to a low-dimensional grid. Through these steps, a two-dimensional feature representation after dimensionality reduction can be obtained, laying a foundation for subsequent clustering analysis.

[0044] 3. k-means Clustering

[0045] After completing data dimensionality reduction, the k-means algorithm is used to perform clustering analysis on the dimensionality-reduced feature data. First, the k-means class is used to initialize the clustering model, setting the number of clusters and the random seed. In this embodiment, n_clusters = 9 is set, indicating that the data is divided into 9 clusters; random_state = 42 is set to ensure the reproducibility of the results. After initialization, the fit_predict() method is used to perform clustering analysis on the dimensionality-reduced feature data som_features. This method will return the cluster labels for each sample, and these labels indicate the cluster to which each sample belongs. In this way, the data is successfully assigned to different clusters, preparing for subsequent evaluation and visualization analysis.

[0046] 4. Evaluation and Visualization

[0047] After clustering is completed, it is necessary to evaluate and visualize the clustering results. First, calculate the evaluation metrics of the clustering results. The evaluation metrics include SSE, MSE, RMSE, silhouette coefficient, and Calinski-Harabasz index. Calculating SSE is a common metric for evaluating the effect of k-means clustering. The smaller the SSE value, the tighter the clustering. Then, MSE and RMSE are calculated through the total number of samples. RMSE is the square root of the clustering error, providing a more intuitive measure of the error. The silhouette_score() is used to calculate the silhouette coefficient to evaluate the separation of the clustering; the calinski_harabasz_score() is used to calculate the Calinski-Harabasz index to evaluate the effectiveness of the clustering. Higher silhouette coefficients and Calinski-Harabasz indices indicate good clustering results. Next, the cluster labels are added to the original data for convenient visualization analysis. Finally, the sns.violinplot() is used to draw a violin plot to show the distribution of each feature under different clusters. In this embodiment, the first and third meteorological classification types in the clustering results belong to the types prone to ozone pollution. By generating violin plots of all features, the clustering results of the data can be comprehensively analyzed and visualized, providing important reference information for environmental monitoring and governance.

[0048] Embodiment 2

[0049] Refer to Figure 7 , this embodiment relates to an industrial park weather classification device integrating a self-organizing mapping neural network and the k-means algorithm, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the industrial park weather classification method of the integrated self-organizing mapping neural network and the k-means algorithm in Embodiment 1.

[0050] Embodiment 3

[0051] This embodiment relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the industrial park weather classification method integrating the self-organizing mapping neural network and the k-means algorithm of Embodiment 1.

[0052] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept.

Claims

1. The industrial park weather classification method integrating self-organizing map neural network and k-means algorithm is characterized by: The steps include: S1, initial data set construction and preprocessing; using the continuous monitoring data of the air monitoring station in the industrial park and the meteorological data of the same period from the ERA5 meteorological data official website, the initial data set was constructed with meteorological elements as variables; After missing value processing, each column of variable data is standardized to construct a standardized input data set; S2, dataset dimensionality reduction based on the SOM model; using the standardized input dataset as the model input, training and optimizing the SOM dimensionality reduction model to obtain the feature dataset; S3, simulation and optimization of ozone concentration using the k-means model; using the feature data set as input, training and optimizing the k-means clustering model, obtaining clustering labels, and evaluating the effectiveness of the clustering results and their distribution visualization.

2. The industrial park weather classification method integrating self-organizing map neural network and k-means algorithm according to claim 1 is characterized in that: The step S1 specifically includes: S1.1, Data acquisition: The continuous monitoring data during the simulation period were extracted from the air quality monitoring station in the industrial park, including wind direction, wind speed, temperature, atmospheric pressure, relative humidity and ozone concentration; the ERA5 supplementary data during the same simulation period were downloaded from the ERA5 meteorological data official website, including the east-west wind speed component u10 at a height of 10 meters, the north-south wind speed component v10 at a height of 10 meters, the surface solar radiation SSRD, the boundary layer height BLH, the high cloud cover HCC, the low cloud cover LCC, and the medium cloud cover MCC; S1.2, initial data set construction: taking meteorological elements as variables, arranging the values ​​of meteorological elements at the same time in order into a row, and arranging each row as a sample in chronological order to form the initial data set; S1.3, data preprocessing; delete rows with missing values ​​in the initial data set to ensure data integrity; standardize the values ​​of each column of variables so that each variable has a distribution with a mean of 0 and a standard deviation of 1, and construct a standardized input data set.

3. The industrial park weather classification method integrating self-organizing map neural network and k-means algorithm according to claim 1 is characterized in that: The step S2 specifically includes: S2.1, SOM model initialization; set the grid size and other parameters, and set the random initialization weights; S2.2, SOM model training: Use standardized input data to train the SOM model, and iteratively adjust the model weights to achieve data mapping; S2.3, obtain SOM features; calculate the winning node of each data sample in the SOM grid, and output it as a new feature as a feature data set.

4. The industrial park weather classification method integrating self-organizing map neural network and k-means algorithm according to claim 1 is characterized in that: The step S3 specifically includes: S3.1, k-means model initialization; Use the 'k-means' class to initialize the clustering model. It is recommended to preset the number of clusters n_clusters = 9, which means that the data is divided into 9 clusters; set random_state = 42 to ensure the repeatability of the results; S3.2, train the k-means model and obtain cluster labels; the k-means algorithm clusters by minimizing the distance from the sample to the center of the cluster to which it belongs; use the fit_predict() method to cluster som_features and obtain the cluster label of each sample; S3.3, clustering result evaluation; calculate the sum of squared error SSE, mean square error MSE and root mean square error RMSE evaluation indicators to evaluate the clustering results; use silhouette_score() to calculate the silhouette coefficient to evaluate the separation of clusters; use calinski_harabasz_score() to calculate the Calinski-Harabasz index to evaluate the effectiveness of clustering; S3.4, clustering result analysis; draw violin plots of each meteorological element under each cluster to view the distribution of each variable under different clusters; use sns.violinplot() to draw the distribution map of ozone concentration and group them according to different cluster labels.

5. The industrial park weather classification method integrating self-organizing map neural network and k-means algorithm as claimed in claim 2, characterized in that: The simulation period described in step S1.1 should be no less than 3 consecutive months; The values ​​of the meteorological elements described in step S1.2 are arranged in order, and the order of the meteorological elements should be kept consistent.

6. The industrial park weather classification method integrating self-organizing map neural network and k-means algorithm as claimed in claim 3, characterized in that: The random initialization weights described in step S2.1 use the random_weights_init() method to randomly initialize the weight settings of the SOM model; Initialize the SOM model described in step S2.1; set the grid size of the SOM to 4×4, i.e., a 16-node two-dimensional grid; other parameters include the input_len parameter, sigma parameter, learning_rate parameter, and random_seed parameter of MiniSom, where the input_len parameter is set to the number of variables of the input data, i.e., the number of meteorological elements; the sigma parameter and the learning_rate parameter are set to the default values; and the random_seed parameter is set to 42 to ensure the repeatability of the model; The SOM model training described in step S2.2 uses the train_random() method to perform model training, and the number of training iterations is set to no less than 50,000; The SOM features described in step S2.3 use the winner() method to obtain the coordinates of the winning node corresponding to each data sample, and convert the coordinates of the winning node into a two-dimensional feature array for subsequent k-means clustering.

7. The industrial park weather classification method integrating self-organizing map neural network and k-means algorithm as claimed in claim 1, characterized in that: The number of clusters described in step S3.1 needs to be adjusted according to the data characteristics. The recommended adjustment range is 6-10 categories. The most appropriate number of clusters is selected based on the clustering result evaluation index in S3.

3.

8. An industrial park weather classification device integrating a self-organizing map neural network and a k-means algorithm, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the industrial park weather classification method integrating the self-organizing map neural network and the k-means algorithm as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Polluted weather typing method and device, electronic equipment and storage medium

    CN112990355A

  • Photovoltaic prediction method and system based on refined weather typing

    CN117175569A

  • Meteorological dominant-based multi-factor future ozone prediction multi-target super integrated learning method

    CN117933430A

  • Photovoltaic power station refined weather typing method in micro-meteorological environment

    CN118211084A

  • Distribution board

    KR1020240162304A