Atmospheric Pollution Three-Dimensional Monitoring System Based on Big Data Analysis
By introducing threshold values and determination distances, core data points, contour coefficients and individual movement strategies into the atmospheric pollution monitoring system, the clustering process is optimized, and the problems of large calculations, unstable results and low accuracy in the atmospheric pollution monitoring system are solved, and efficient and accurate monitoring effects are achieved.
Patent Information
- Application Number
- CN202411709303.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The existing air pollution monitoring system has too much calculation due to excessive data sets, poor adaptability during clustering, improper selection of clustering centers leads to unstable results, improper parameter selection leads to low accuracy, and weak monitoring performance.
The threshold value and the determination distance setting are introduced to form a point-to-point similarity matrix, select the data points, select the best clustering center based on the core data points and contour coefficients, establish a search space and formulate an individual movement strategy, re-initialize individuals with the lowest fitness value, and optimize the clustering process.
Reduce data storage space and calculation amount, accelerate convergence effect, improve the stability and accuracy of clustering results, enhance monitoring performance, avoid local optimal solutions, and improve global search capabilities.
Smart Images

Figure CN119207646B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air pollution monitoring, and specifically refers to a three-dimensional air pollution monitoring system based on big data analysis. Background Art
[0002] An air pollution monitoring system is a system used to continuously monitor, analyze, and report the concentrations of pollutants in the atmospheric environment. It provides important information support for environmental protection, public health, and policy-making by collecting and processing environmental data. The main purpose of this system is to evaluate air quality, identify pollution sources, warn of pollution events, and help formulate and evaluate air quality improvement measures. However, in general air pollution monitoring systems, there are problems such as excessive computational amount due to too large a dataset, poor adaptability to the dataset during clustering resulting in poor clustering effect, and low accuracy of air pollution monitoring results; there is a problem that the selection of clustering centers is inappropriate during clustering, resulting in unstable clustering results and ultimately poor detection effects; there is a problem that the selection of parameters is inappropriate during clustering, resulting in low accuracy of clustering results and thus weak monitoring performance. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides a three-dimensional air pollution monitoring system based on big data analysis. Aiming at the problems in general air pollution monitoring systems, such as excessive computational amount due to too large a dataset, poor adaptability to the dataset during clustering resulting in poor clustering effect, and low accuracy of air pollution monitoring results, this solution introduces the setting of thresholds and determination distances, selectively forms a point-to-point similarity matrix, and ensures that only one nearest point is selected each time when allocating data points; to reduce data storage space and computational amount; and to accelerate the convergence effect; aiming at the problem in general air pollution monitoring systems that the selection of clustering centers is inappropriate during clustering, resulting in unstable clustering results and ultimately poor detection effects, this solution introduces core data points; selects the optimal clustering center based on the core data points and the silhouette coefficient; reduces the data computational amount, and makes the clustering results more reasonable and highly flexible; aiming at the problem in general air pollution monitoring systems that the selection of parameters is inappropriate during clustering, resulting in low accuracy of clustering results and thus weak monitoring performance, this solution establishes a search space based on clustering hyperparameters, formulates different movement strategies for individuals based on the individual position and the preferred individual position, improves the convergence speed and search efficiency of the algorithm; and re-initializes the three individuals with the lowest fitness values; can maintain the diversity of the population during the search process, avoid falling into local optimal solutions, and improve the global search ability of the algorithm; realizes air pollution monitoring based on clustering optimization.
[0004] The technical solution adopted by the present invention is as follows: The three-dimensional air pollution monitoring system based on big data analysis provided by the present invention includes a data acquisition module, a data preprocessing module, a sub-cluster generation module, an optimal number of cluster centers selection module, a clustering optimization module, and a three-dimensional air pollution monitoring module;
[0005] The data acquisition module acquires historical monitoring data and real-time monitoring data;
[0006] The data preprocessing module performs data cleaning, data transformation, and standardization processing on the acquired data;
[0007] The sub-cluster generation module introduces the setting of a threshold and a determination distance, selectively forms a point pair similarity matrix, and ensures that only one nearest point is selected each time when allocating data points; thus obtaining sub-clusters;
[0008] The optimal number of cluster centers selection module introduces core data points; selects the optimal cluster centers based on the core data points and the silhouette coefficient;
[0009] The clustering optimization module establishes a search space based on clustering hyperparameters, formulates different movement strategies for individuals based on the individual position and the preferred individual position; and re-initializes the three individuals with the lowest fitness values; to achieve clustering optimization;
[0010] The three-dimensional air pollution monitoring module performs three-dimensional air pollution monitoring on the real-time monitoring data based on the optimized clustering results.
[0011] Further, in the data acquisition module, both the historical monitoring data and the real-time monitoring data include pollutant concentration data, meteorological data, auxiliary data, and geographic information data; the pollutant concentration data includes particulate matter concentration data, air pollutant concentration data, and organic matter concentration data; the meteorological data includes temperature, humidity, wind speed, wind direction, and atmospheric pressure; the auxiliary data includes traffic flow data and industrial production data; the geographic information data includes population density, terrain, and land use; the historical monitoring data also includes the air pollution level.
[0012] Further, in the data preprocessing module, the data cleaning is to process missing values, duplicate values, and outliers; the data transformation is to convert the data into a vector form; the standardization processing is to standardize the data based on the maximum-minimum normalization to obtain the original data set.
[0013] Further, the sub-cluster generation module specifically includes the following:
[0014] An initial point pair similarity matrix formation unit, which is pre-set with a distance threshold ; for each pair of data points p i and p j, calculate the similarity d of data points based on the Euclidean distance ij ; if d ij is not greater than , then add the point pair and the corresponding similarity to the point pair similarity matrix;
[0015] Data clustering unit, select the maximum distance from the initial point pair similarity matrix as the decision distance, and randomly initialize the first clustering center;
[0016] Data point assignment unit, find the point that is closest to the current clustering center and the distance is not greater than , add it to the current cluster; ensure that only one nearest point is selected each time; when there is no qualified point, go to the update clustering center unit;
[0017] Update clustering center unit, if there are unassigned data points, then select the data points whose distance from the clustering center is greater than as the new generation of clustering centers, and go to the data point assignment unit; otherwise go to the iteration unit;
[0018] Iteration unit, if the maximum number of iterations is not reached or the clustering has not converged, then update the point pair similarity matrix with it as the distance threshold, select the center point of the current cluster as the new clustering center, and go to the data point assignment unit; if the maximum number of iterations is reached, re-initialize the first clustering center; otherwise the sub-cluster generation ends.
[0019] Furthermore, the optimal number of clustering centers selection module specifically includes the following:
[0020] Core dataset generation unit, initialize the core dataset as empty; for each sub-cluster, select the data point closest to the center point of the sub-cluster, the point closest to other sub-clusters in the sub-cluster, and the point with the maximum density in the sub-cluster as the core data points; add the core data points of all sub-clusters to the core dataset;
[0021] Core similarity matrix generation unit, calculate the Euclidean distance between all data points in the core dataset to obtain the core similarity matrix;
[0022] Multi-clustering unit, select different numbers of clustering centers and the corresponding decision distances to perform clustering processing based on the core similarity matrix, randomly initialize the first clustering center, and the clustering process is the same as that from the data point assignment unit to the iteration unit; when the clustering converges, evaluate the clustering results of different numbers of clustering centers based on the silhouette coefficient, and select the number of clustering centers corresponding to the clustering result with the largest evaluation value as the optimal number of clustering centers;
[0023] The final clustering unit performs clustering on the original data set based on the optimal number of clustering centers. Specifically, a data point is randomly selected from the core data set as the first initialized clustering center, and each time a data point with the farthest distance from the current clustering center is selected as the clustering center until the optimal number of clustering centers is reached. The data points are allocated based on the data point allocation unit. When updating the clustering center, the center point of the cluster is selected as the new generation of clustering center. If the clustering converges, the final clustering is completed. If the maximum number of iterations is reached, the first clustering center is re-initialized. Otherwise, the clustering is continued iteratively.
[0024] Furthermore, the clustering optimization module specifically includes the following:
[0025] The initialization unit establishes a search space based on the distance threshold, the initial first clustering center of the sub-cluster, the determination distances corresponding to different numbers of clustering centers in multi-clustering, and the initial first clustering center; randomly initializes the search population position, and takes the average silhouette coefficient of the cluster after clustering iteratively k times based on the individual position as the individual fitness value.
[0026] The position update unit selects the three individuals with the highest fitness values as the preferred individuals, denoted by a, b, and c respectively; updates the positions of the individuals, and the formula used is as follows:
[0027] ;
[0028] ;
[0029] In the formula, X i (·) is the position of the i-th individual except the preferred individuals, and t represents the number of iterations; X a (·), X b (·), and X c (·) are the positions of the preferred individuals, Fa, Fb, and Fc are the fitness values corresponding to the preferred individuals; i ∈ a, i ∈ b, and i ∈ c respectively represent that the i-th individual is the closest to a, b, and c; Y(·) is the position of the preferred individual; T is the maximum number of iterations; is the average position of the preferred individuals; is the smoothing term;
[0030] The re-initialization unit initializes the three individuals with the lowest fitness values, and the formula used is as follows:
[0031] ;
[0032] In the formula, newX d (·) and X d (·) are the positions of the d-th dimension of the three individuals with the lowest fitness values after and before initialization respectively, L d and U dThey are the upper and lower limits of the search space in the d-th dimension respectively;
[0033] The search determination unit is preset with a fitness threshold. When there is an individual fitness value higher than the fitness threshold, clustering processing is performed on the data set based on the individual position to obtain the optimized final clustering result; if the maximum number of iterations is reached, the population position is re-initialized; otherwise, a preferred individual is re-selected to continue iterative search.
[0034] Furthermore, based on the optimized final clustering result, the air pollution three-dimensional monitoring module selects the label with the largest number of historical data as the cluster label; and outputs the cluster label to which the real-time data belongs as the final monitoring result of the real-time data.
[0035] The beneficial effects achieved by the present invention using the above solution are as follows:
[0036] (1) Aiming at the problems of the general air pollution monitoring system that the data set is too large resulting in excessive calculation amount and poor adaptability to the data set during clustering, resulting in poor clustering effect and low accuracy of air pollution monitoring results, this solution introduces the setting of a threshold and a determination distance, selectively forms a point pair similarity matrix, and ensures that only one nearest point is selected each time when allocating data points; to reduce data storage space and calculation amount; and accelerate the convergence effect.
[0037] (2) Aiming at the problem that the general air pollution monitoring system has an improper selection of the clustering center during clustering, resulting in unstable clustering results and ultimately poor detection effects, this solution introduces core data points; selects the best clustering center based on the core data points and the silhouette coefficient; reduces the data calculation amount, and makes the clustering result more reasonable and highly flexible.
[0038] (3) Aiming at the problem that the general air pollution monitoring system has an improper selection of parameters during clustering, resulting in low accuracy of the clustering result and thus weak monitoring performance, this solution establishes a search space based on clustering hyperparameters, formulates different movement strategies for individuals based on the individual position and the preferred individual position, improves the convergence speed and search efficiency of the algorithm; and re-initializes the three individuals with the lowest fitness values; can maintain the diversity of the population during the search process, avoid falling into local optimal solutions, and improve the global search ability of the algorithm; realizes air pollution monitoring based on clustering optimization. Description of the Drawings
[0039] Figure 1 It is a schematic flow chart of the air pollution three-dimensional monitoring system based on big data analysis provided by the present invention;
[0040] Figure 2 It is a schematic flow chart of the sub-clustering generation module;
[0041] Figure 3It is a schematic flow diagram of the optimal number of clustering centers selection module.
[0042] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. Detailed implementation manners
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0045] Example 1, refer to Figure 1 , the three-dimensional air pollution monitoring system based on big data analysis provided by the present invention includes a data acquisition module, a data preprocessing module, a sub-clustering generation module, an optimal number of clustering centers selection module, a clustering optimization module, and a three-dimensional air pollution monitoring module;
[0046] The data acquisition module collects historical monitoring data and real-time monitoring data and sends the data to the data preprocessing module;
[0047] The data preprocessing module receives the data sent by the data acquisition module, performs data cleaning, data conversion, and standardization processing on the collected data, and sends the data to the sub-clustering generation module;
[0048] The sub-clustering generation module receives the data sent by the data preprocessing module, introduces the setting of thresholds and determination distances, selectively forms a point pair similarity matrix, and ensures that only one nearest point is selected each time when allocating data points, thereby obtaining sub-clusters, and sending the data to the optimal number of clustering centers selection module;
[0049] The optimal number of clustering centers selection module receives the data sent by the sub-clustering generation module, introduces core data points, selects the optimal clustering centers based on the core data points and the silhouette coefficient, and sends the data to the clustering optimization module;
[0050] The clustering optimization module receives the data sent by the optimal number of clustering centers selection module; establishes a search space based on clustering hyperparameters, formulates different movement strategies for individuals based on individual positions and preferred individual positions; and re-initializes the three individuals with the lowest fitness values; to achieve clustering optimization; and sends the data to the three-dimensional air pollution monitoring module;
[0051] The three-dimensional air pollution monitoring module receives the data sent by the clustering optimization module; and performs three-dimensional air pollution monitoring on the real-time monitoring data based on the optimized clustering results.
[0052] Embodiment 2, refer to Figure 1 , this embodiment is based on the above embodiment. In the data acquisition module, both historical monitoring data and real-time monitoring data include pollutant concentration data, meteorological data, auxiliary data, and geographic information data; the pollutant concentration data includes particulate matter concentration data, air pollutant concentration data, and organic matter concentration data; the meteorological data includes temperature, humidity, wind speed, wind direction, and atmospheric pressure; the auxiliary data includes traffic flow data and industrial production data; the geographic information data includes population density, terrain, and land use; the historical monitoring data also includes the air pollution level.
[0053] Embodiment 3, refer to Figure 1 , this embodiment is based on the above embodiment. In the data preprocessing module, data cleaning is to process missing values, duplicate values, and outliers; the data conversion is to convert the data into a vector form; the standardization process is to standardize the data based on maximum-minimum normalization to obtain the original data set.
[0054] Embodiment 4, refer to Figure 1 and Figure 2 , this embodiment is based on the above embodiment. The sub-clustering generation module specifically includes the following:
[0055] Initial point pair similarity matrix formation unit, with a distance threshold preset ; for each pair of data points p i and p j in the data set, calculate the data point similarity d ij based on the Euclidean distance; if d ij is not greater than , then add the point pair and the corresponding similarity to the point pair similarity matrix;
[0056] Data clustering unit, select the maximum distance from the initial point pair similarity matrix as the determination distance, and randomly initialize the first clustering center;
[0057] Data point allocation unit, find the data point closest to the current clustering center and with a distance not greater than Points that meet the criteria are added to the current cluster. Ensure that only one nearest point is selected each time. When there are no eligible points, go to the cluster center update unit.
[0058] Update the cluster center unit. If there are unassigned data points, select data points whose distance from the cluster center is greater than as the new generation of cluster centers and go to the data point assignment unit; otherwise, go to the iteration unit.
[0059] Iteration unit. If the maximum number of iterations has not been reached or the clustering has not converged, then is used as the distance threshold to update the point pair similarity matrix. Select the center point of the current cluster as the new cluster center and go to the data point assignment unit. If the maximum number of iterations has been reached, re-initialize the first cluster center; otherwise, the sub-cluster generation ends.
[0060] By performing the above operations, for the problems existing in the general air pollution monitoring system, such as excessive calculation due to too large a dataset and poor adaptability of the dataset during clustering resulting in poor clustering effect and low accuracy of air pollution monitoring results, this solution introduces the setting of thresholds and determination distances, selectively forms a point pair similarity matrix, and ensures that only one nearest point is selected each time when allocating data points; to reduce data storage space and calculation amount; and to accelerate the convergence effect.
[0061] Example 5, refer to Figure 1 and Figure 3 . Based on the above example, the best cluster center number selection module specifically includes the following:
[0062] Core dataset generation unit, initialize the core dataset as empty; for each sub-cluster, select the data point closest to the center point of the sub-cluster, the point closest to other sub-clusters in the sub-cluster, and the point with the maximum density in the sub-cluster as the core data points; add the core data points of all sub-clusters to the core dataset.
[0063] Core similarity matrix generation unit, calculate the Euclidean distance between all data points in the core dataset to obtain the core similarity matrix.
[0064] Multi-cluster unit, select different numbers of cluster centers and corresponding determination distances to perform clustering processing based on the core similarity matrix, randomly initialize the first cluster center, and the clustering process is the same as that from the data point assignment unit to the iteration unit; when the clustering converges, evaluate the clustering results of different numbers of cluster centers based on the silhouette coefficient, and select the number of cluster centers corresponding to the clustering result with the largest evaluation value as the best number of cluster centers.
[0065] The final clustering unit performs clustering on the original data set based on the optimal number of clustering centers. Specifically, a data point is randomly selected from the core data set as the first initialized clustering center, and each time a data point with the farthest distance from the current clustering center is selected as the clustering center until the optimal number of clustering centers is reached. The data points are assigned based on the data point assignment unit. When updating the clustering center, the center point of the cluster is selected as the new generation of clustering center. If the clustering converges, the final clustering is completed. If the maximum number of iterations is reached, the first clustering center is re-initialized. Otherwise, the clustering is continued iteratively.
[0066] By performing the above operations, for the problem that the general air pollution monitoring system has unstable clustering results due to improper selection of clustering centers during clustering, which ultimately leads to poor detection effects, this solution introduces core data points. The optimal clustering center is selected based on the core data points and the silhouette coefficient. The data calculation amount is reduced, and the clustering result is made more reasonable and highly flexible.
[0067] Example six, refer to Figure 1 , based on the above example, the clustering optimization module specifically includes the following content:
[0068] The initialization unit establishes a search space based on the distance threshold, the initial first clustering center of the sub-clustering, the determination distances corresponding to different numbers of clustering centers in multi-clustering, and the initial first clustering center. The search population position is randomly initialized, and the average silhouette coefficient of the cluster after clustering iteration k times based on the individual position is used as the individual fitness value.
[0069] The position update unit selects the three individuals with the highest fitness values as the preferred individuals, denoted as a, b, and c respectively. The positions of the individuals are updated using the following formula:
[0070] ;
[0071] ;
[0072] In the formula, X i (·) is the position of the i-th individual except the preferred individuals, and t represents the number of iterations; X a (·), X b (·) and X c (·) are the positions of the preferred individuals, Fa, Fb, and Fc are the fitness values corresponding to the preferred individuals; i ∈ a, i ∈ b, and i ∈ c respectively represent that the i-th individual is the closest to a, b, and c; Y(·) is the position of the preferred individual; T is the maximum number of iterations; is the average position of the preferred individuals; is the smoothing term;
[0073] The re-initialization unit initializes the three individuals with the lowest fitness values using the following formula:
[0074] ;
[0075] In the formula, newX d (·) and X d (·) are respectively the positions of the three individuals with the lowest fitness after and before initialization in the d -th dimension, L d and U d are respectively the upper and lower limits of the search space in the d -th dimension;
[0076] The search decision unit is preset with a fitness threshold. When there is an individual fitness value higher than the fitness threshold, clustering processing is performed on the data set based on the individual positions to obtain the optimized final clustering result; if the maximum number of iterations is reached, the population positions are re - initialized; otherwise, preferred individuals are re - selected to continue iterative search.
[0077] By performing the above operations, aiming at the problem that the clustering result has low accuracy due to improper parameter selection in the general air pollution monitoring system, this solution establishes a search space based on clustering hyper - parameters, formulates different movement strategies for individuals based on the individual positions and the positions of preferred individuals, improves the convergence speed and search efficiency of the algorithm; and re - initializes the three individuals with the lowest fitness values; can maintain the diversity of the population during the search process, avoid falling into local optimal solutions, and improve the global search ability of the algorithm; realizes air pollution monitoring based on clustering optimization.
[0078] Example Seven, refer to Figure 1 , based on the above - mentioned example, the three - dimensional air pollution monitoring module selects the label with the largest number of historical data as the cluster label based on the optimized final clustering result; outputs the cluster label to which the real - time data belongs as the final monitoring result of the real - time data.
[0079] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device.
[0080] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.
[0081] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design, without creative efforts, structural manners and embodiments similar to the technical solution without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. An atmospheric pollution three-dimensional monitoring system based on big data analysis, characterized in that: The system includes a data acquisition module, a data preprocessing module, a sub-cluster generation module, an optimal number of cluster centers selection module, a clustering optimization module, and a three-dimensional air pollution monitoring module; The data acquisition module acquires historical monitoring data and real-time monitoring data; The data preprocessing module performs data cleaning, data transformation, and standardization processing on the acquired data; The sub-cluster generation module introduces the setting of a threshold and a determination distance, selectively forms a pairwise similarity matrix, and ensures that only one nearest point is selected each time when allocating data points; thereby obtaining sub-clusters; The optimal number of cluster centers selection module introduces core data points; selects the optimal cluster centers based on the core data points and the silhouette coefficient; The clustering optimization module establishes a search space based on clustering hyperparameters, formulates different movement strategies for individuals based on the individual positions and the preferred individual positions; and re-initializes the three individuals with the lowest fitness values; to achieve clustering optimization; The three-dimensional air pollution monitoring module performs three-dimensional air pollution monitoring on real-time monitoring data based on the optimized clustering results; The sub-cluster generation module specifically includes the following: Initial point pair similarity matrix forming unit, with a distance threshold preset ; For each pair of data points p i and p j in the dataset, calculate the data point similarity d ij based on the Euclidean distance; if d ij is not greater than , then add the point pair and the corresponding similarity to the point pair similarity matrix; Data clustering unit, selects the maximum distance from the initial point pair similarity matrix As the determination distance, randomly initialize the first clustering center; The data point allocation unit searches for the point that is closest to the current cluster center and whose distance is not greater than and adds it to the current cluster. Ensure that only one nearest point is selected each time; when there is no qualified point, go to the update cluster center unit; Update the clustering center unit. If there are data points that have not been assigned, select the data points whose distance from the clustering center is greater than as the new generation of clustering centers, and transfer to the data point assignment unit; otherwise, transfer to the iteration unit; The iterative unit, if the maximum number of iterations is not reached or the clustering has not converged, then is used as the distance threshold to update the point pair similarity matrix, the center point of the current cluster is selected as the new cluster center, and it goes to the data point allocation unit; If the maximum number of iterations is reached, re-initialize the first cluster center; otherwise, the sub-cluster generation ends; It is characterized in that: the clustering optimization module specifically includes the following: The initialization unit establishes a search space based on the distance threshold, the initial first cluster center of the sub-cluster, the determination distances corresponding to different numbers of cluster centers in multi-clustering, and the initial first cluster center; randomly initializes the search population positions, and takes the average silhouette coefficient of the clusters obtained by clustering the individuals based on their positions for k iterations as the individual fitness value; The position update unit selects the three individuals with the highest fitness values as the preferred individuals, denoted by a, b, and c respectively; updates the positions of the individuals, and the formula used is as follows: ; ; In the formula, X i (·) is the position of the i-th individual except for the optimal individuals, and t represents the number of iterations; X a (·), X b (·) and X c (·) are the positions of the optimal individuals. Fa, Fb, and Fc are the fitness values corresponding to the optimal individuals; i ∈ a, i ∈ b, and i ∈ c respectively indicate that the i-th individual is the closest to a, b, and c; Y(·) is the position of the optimal individual; T is the maximum number of iterations; is the average position of the optimal individuals; is the smoothing term; The re-initialization unit initializes the three individuals with the lowest fitness values, and the formula used is as follows: ; where newX d (·) and X d (·) are the positions of the d-th dimension of the three individuals with the lowest fitness after and before initialization respectively, L d and U d are the upper and lower limits of the search space of the d-th dimension respectively; The search determination unit is preset with a fitness threshold. When there are individual fitness values higher than the fitness threshold, clustering processing is performed on the data set based on the individual positions to obtain the optimized final clustering result; if the maximum number of iterations is reached, re-initialize the population positions; otherwise, re-select the preferred individuals and continue iterative search; The optimal number of cluster centers selection module specifically includes the following: The core data set generation unit initializes the core data set to be empty; for each sub-cluster, selects the data point closest to the center point of the sub-cluster, the point closest to other sub-clusters in the sub-cluster, and the point with the maximum density in the sub-cluster as the core data points; adds the core data points of all sub-clusters to the core data set; The core similarity matrix generation unit calculates the Euclidean distances between all data points in the core data set to obtain the core similarity matrix; Multiple clustering units select different numbers of clustering centers and corresponding determination distances to perform clustering processing based on the core similarity matrix. The first clustering center is randomly initialized, and the clustering process is the same as that of the data point allocation unit to the iteration unit. When the clustering converges, the clustering results of different numbers of clustering centers are evaluated based on the silhouette coefficient, and the number of clustering centers corresponding to the clustering result with the largest evaluation value is selected as the optimal number of clustering centers. The final clustering unit performs clustering processing on the original data set based on the optimal number of clustering centers. Specifically, a data point is randomly selected from the core data set as the first initialized clustering center, and each time a data point with the farthest distance from the current clustering center is selected as the clustering center until the optimal number of clustering centers is reached. The data points are allocated based on the data point allocation unit. When updating the clustering center, the center point of the cluster is selected as the new generation of clustering center. If the clustering converges, the final clustering is completed. If the maximum number of iterations is reached, the first clustering center is re-initialized. Otherwise, the clustering is continued iteratively. In the data acquisition module, both the historical monitoring data and the real-time monitoring data include pollutant concentration data, meteorological data, auxiliary data, and geographic information data. The pollutant concentration data includes particulate matter concentration data, air pollutant concentration data, and organic matter concentration data. The meteorological data includes temperature, humidity, wind speed, wind direction, and atmospheric pressure. The auxiliary data includes traffic flow data and industrial production data. The geographic information data includes population density, terrain, and land use. The historical monitoring data also includes the air pollution level. The three-dimensional air pollution monitoring module selects the label with the largest number of historical data as the cluster label based on the optimized final clustering result, and outputs the cluster label to which the real-time data belongs as the final monitoring result of the real-time data.
2. The atmospheric pollution three-dimensional monitoring system based on big data analysis according to claim 1, wherein: In the data preprocessing module, data cleaning processes missing values, duplicate values, and outliers. Data transformation converts the data into vector form. The normalization process normalizes the data based on the maximum-minimum normalization to obtain the original data set.
Citation Information
Patent Citations
Knowledge base data information clustering method and system based on MapReduce model
CN115687539A
Electronic commerce market trend prediction system based on big data analysis
CN118212001A