Atmospheric pollutant concentration clustering method based on time sequence shape features
By combining the O-SBD distance estimated by Theil-Sen with the shape extraction strategy constrained by KL divergence, the problem of measuring sequence distance in air pollution monitoring data is solved, and more efficient cluster analysis and pollution prevention and control support are achieved.
Patent Information
- Application Number
- CN202511442120.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies struggle to effectively measure the distance between sequences in air pollution monitoring data. Traditional methods are unsuitable for handling outliers, non-stationarity, and noise, resulting in insufficient accuracy and robustness of clustering algorithms.
The maximum triangle three-segment downsampling algorithm is used to downsample the time series of atmospheric pollutant concentrations. The normalized cross-correlation coefficient estimated by Theil-Sen is constructed, and the O-SBD distance is used. The cluster center sequence is updated by the shape extraction strategy of KL divergence constraint and cluster analysis is performed.
It significantly improves the accuracy and stability of atmospheric pollutant concentration clustering, effectively handles outliers, preserves key shape features, and provides more reliable pollution pattern identification and prevention and control decision support.
Smart Images

Figure CN121256409A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of atmospheric pollution prevention and control engineering and pattern recognition technology, and particularly relates to a clustering method for atmospheric pollutant concentration based on time series shape features. BACKGROUND
[0002] Atmospheric pollution prevention and control engineering is a key technical field in environmental engineering, and its core task is to prevent and control atmospheric pollution caused by human activities through engineering technical measures to improve air quality. Therefore, developing effective atmospheric pollution control technology and monitoring methods is crucial for regional air quality management and pollution source control. Time series clustering is a technique that groups time series based on their characteristics, and it has a wide range of applications in atmospheric monitoring data. The difficulty of atmospheric sequence data clustering lies in how to measure the distance (similarity) between two sequences. Traditional distance measurement methods such as Euclidean distance are not suitable for sequence data. Shape-based clustering algorithms have been proven to be effective in handling raw sequence data sets. The core advantage of this type of algorithm is that it analyzes the overall shape characteristics of the sequence, rather than relying solely on the original numerical values of the data points. This method can reveal the global patterns and long-term trends of sequence data, rather than just focusing on local fluctuations or transient changes. In atmospheric monitoring data, data often contains outliers, non-stationarity and noise, which pose challenges to the accuracy and robustness of clustering algorithms. Therefore, the present application provides a clustering method for atmospheric pollutant concentration based on time series shape features. SUMMARY
[0003] To solve the above technical problems, the present application provides a clustering method for atmospheric pollutant concentration based on time series shape features to solve the problems existing in the prior art.
[0004] To achieve the above purpose, the present application provides a clustering method for atmospheric pollutant concentration based on time series shape features, comprising the following steps: Performing down-sampling processing on the original atmospheric pollutant concentration time series based on the maximum triangle three-section down-sampling algorithm to obtain a down-sampled sequence; Based on the down-sampled sequence, constructing a normalized cross-correlation coefficient using the Theil-Sen estimation method, and obtaining an O-SBD distance based on the normalized cross-correlation coefficient; Based on the O-SBD distance, updating the cluster center sequence using a KL divergence constrained shape extraction strategy to obtain the final center sequence; Performing clustering analysis on the atmospheric pollutant concentration time series based on the center sequence until convergence is achieved to obtain the clustering result.
[0005] Optionally, the process of obtaining the down-sampled sequence comprises: The maximum triangle three-segment algorithm is used to segment the original atmospheric pollutant concentration time series to obtain several data segments; Representative data points are selected from each data segment using an importance-based selection method; The downsampling sequence is obtained by combining the representative data points.
[0006] Optionally, the expression for constructing the normalized cross-correlation coefficient is as follows: ; In the formula, This represents the normalized cross-correlation coefficient, where X and Y represent two sequences. and This indicates that a fast Fourier transform was performed on two atmospheric pollutant concentration sequences. Indicates to Perform inverse Fourier transform Indicates to Perform Theil-Sen estimation storage.
[0007] Optionally, the expression for calculating the O-SBD distance is: ; In the formula, Let O-SBD be the distance between sequences X and Y.
[0008] Optionally, based on the O-SBD distance, the process of updating the cluster center sequence using a KL divergence constraint shape extraction strategy to obtain the final center sequence includes: Each atmospheric sequence is assigned to a corresponding cluster based on the O-SBD distance; For each cluster, an optimization objective is constructed to balance the shape similarity and distribution consistency among the fused sequences, where KL divergence constraints are used to ensure that the probability distribution of the center sequence is consistent with the overall distribution of the sequences within the cluster. By solving the eigenvalue problem corresponding to the optimization objective, the eigenvector corresponding to the largest eigenvalue is used as the updated center sequence; The atmospheric sequence allocation and central sequence update steps are executed iteratively to obtain the final central sequence.
[0009] Optionally, the expression for calculating the center sequence is: ; In the formula, The optimal solution represents the cluster center sequence. Represents the cluster center sequence, This represents a sequence in the air pollution concentration dataset. Let M be the matrix. Let KL be the Lagrange multiplier, and KL denote the KL divergence. For sequence The probability distribution, Represents the cluster center sequence The probability distribution, This represents a cluster containing a sequence of atmospheric pollutant concentrations.
[0010] Optionally, the KL divergence constraint ensures consistent distributions by constructing a matrix M, the calculation expression of which is: ; ; ; In the formula, This represents the sum of the outer products of all sequences in cluster C. It is a matrix in which all elements are 1. The number of sequences in the cluster. It is the identity matrix. for The transpose of .
[0011] Compared with the prior art, the present invention has the following advantages and technical effects: This invention introduces a maximum triangle three-segment downsampling algorithm to effectively compress data size while preserving key shape features during the preprocessing stage, and simultaneously removes outliers, significantly improving subsequent processing efficiency. Employing an O-SBD distance metric based on Theil-Sen estimation enhances robustness to outliers and enables more accurate assessment of shape similarity between sequences. Furthermore, a center sequence extraction strategy constrained by KL divergence ensures that cluster centers accurately reflect the overall characteristics of the data within each cluster in both morphology and distribution. This method demonstrates higher accuracy and stability in atmospheric pollutant concentration clustering, providing reliable technical support for regional pollution pattern recognition and prevention and control decisions. Attached Figure Description
[0012] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of the clustering method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of atmospheric monitoring data according to an embodiment of the present invention; Figure 3 This is a clustering result diagram of the present invention applied to the Jiangsu, Zhejiang and Shanghai region; Figure 4 This is a clustering result diagram of the application of the present invention in the Shaanxi-Gansu-Ningxia region. Detailed Implementation
[0013] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0014] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0015] Example 1.
[0016] This invention provides O-Shape, an atmospheric pollutant concentration clustering method based on time-series shape features. This method is of great significance in the fields of air pollution control engineering and time-series clustering. This method specifically considers the high-latitude characteristics, structural complexity, unique distribution, and outliers of atmospheric monitoring data. These problems may be caused by various factors, including industrial emissions, traffic pollution, and natural factors. The O-Shape algorithm first downsamples the complex atmospheric sequence monitoring data using the maximum triangle three-segment downsampling method, thereby preserving the true sequence shape while maximizing the acquisition of sequence data features. Then, the algorithm proposes a novel shape-based distance that is robust to outliers and can effectively eliminate outlier interference. Finally, O-Shape proposes an advanced shape extraction strategy that can accurately identify the central trend of atmospheric sequences and promote the effective clustering of similar sequences. In air pollution control, the application of this method can reveal the potential correlations and differences in air quality in different regions, which is crucial for the precise management of regional air pollution. For example, by analyzing the time-series air quality data between geographically adjacent and non-adjacent cities, the propagation paths and impact ranges of pollution sources can be identified. Furthermore, this method can improve the accuracy and reliability of cluster analysis, enabling more stable and reliable clustering results on sequence datasets containing outliers.
[0017] Figure 2This is a schematic diagram of atmospheric monitoring data according to an embodiment of the present invention. First, the present invention uses a downsampling algorithm to process the original atmospheric monitoring sequence to eliminate redundant data and improve the efficiency of atmospheric data processing. Second, a distance metric O-Shape based on Theil-Sen robust estimation is defined. Furthermore, the present invention proposes a novel center sequence determination strategy based on shape extraction and constrained by KL divergence. These modules work together to enable O-Shape to effectively identify and process outliers in the data while maintaining sensitivity to the shape of atmospheric time series data, reducing the interference of outliers on clustering results. Therefore, this algorithm not only improves the accuracy and reliability of atmospheric data clustering analysis, but also provides more stable and reliable clustering results on atmospheric time series datasets containing outliers, especially in the field of atmospheric monitoring, providing strong data support for air pollution prevention and control.
[0018] like Figure 1 As shown in the figure, this embodiment provides a clustering method for atmospheric pollutant concentrations based on time series shape features, including the following steps: The original atmospheric pollutant concentration time series is downsampled using the maximum triangle three-segment downsampling algorithm to obtain a downsampled sequence; Based on the downsampled sequence, a normalized cross-correlation coefficient is constructed using the Theil-Sen estimation method, and the O-SBD distance is obtained based on the normalized cross-correlation coefficient; Based on the O-SBD distance, a shape extraction strategy with KL divergence constraints is used to update the cluster center sequence to obtain the final center sequence; Cluster analysis is performed on the atmospheric pollutant concentration time series based on the center sequence until convergence, yielding the clustering result.
[0019] S1: Atmospheric sequence downsampling processing.
[0020] First, the maximum triangle triad algorithm is used to reduce the amount of atmospheric sequence data while preserving shape features and removing outliers from the sequence. The specific steps are as follows: The maximum triangle triad algorithm is used to segment the original atmospheric pollutant concentration time series to obtain several data segments; the importance screening method is used to select representative data points from each data segment; and the representative data points are combined to obtain the downsampled sequence.
[0021] 1) The original sequences of the obtained atmospheric sequence dataset were divided into roughly equal-sized segments based on the number of sampled data. The first and last segments contained only the first and last data points, respectively. 2) Iterate from the first segment to the last segment, and select the most important sequence from each segment; 3) Combine the most important sequences identified in each segment into a new sequence to represent the original data, thereby achieving downsampling.
[0022] S2: New distance construction.
[0023] The normalized cross-correlation coefficient is estimated using Theil-Sen, and then a novel shape-based distance O-SBD is employed, the specific derivation of which is as follows: For atmospheric sequence datasets Two atmospheric sequences were detected. and : , , ,have The definition is as follows: (1) in, express the median of 0 and 1 represent uncertain and certain estimates, respectively. It can effectively handle outliers.
[0024] Sequences estimated based on Theil-Sen and The new cross-correlation coefficients between them are shown below: (2) in, , ,when When the expression is equal to 1, it indicates that the two sequences are completely dissimilar, while 1 indicates that the two sequences are completely identical. Normalization is performed: (3) A distance term called O-SBD is defined to evaluate the shape similarity between two sequences, and its derivation is as follows: (4) in, ,and The closer the value is to 0, the stronger the correlation between the two sequences.
[0025] S3: A new strategy for determining the cluster center sequence.
[0026] A strategy based on shape extraction using KL divergence constraints to determine the center sequence is proposed. Specifically, atmospheric sequences are assigned to their corresponding clusters based on O-SBD distance. For each cluster, an optimization objective is constructed that fuses the shape similarity and distribution consistency between sequences, where KL divergence constraints ensure that the probability distribution of the center sequence is consistent with the overall distribution of sequences within the cluster. By solving the eigenvalue problem corresponding to the optimization objective, the eigenvector corresponding to the largest eigenvalue is used as the updated center sequence. The atmospheric sequence assignment and center sequence update steps are iteratively executed to obtain the final center sequence.
[0027] 1) KSE is first initialized by randomly selecting an initial center sequence from the atmospheric sequence data. For a cluster C containing atmospheric sequences, the center sequence is... Determined by equation (5), here and Let M represent the probability distributions of the time series and the center series, respectively. Matrix M is defined by equation (6), the center matrix Q is defined by equation (7), and matrix S is defined by equation (8). (5) (6) (7) (8) In the formula, The optimal solution represents the cluster center sequence. Represents the cluster center sequence, This represents a sequence in the air pollution concentration dataset. Let M be the matrix. These are Lagrange multipliers used to balance the weights of two functions, and KL represents the KL divergence, used to measure the difference between two probability distributions. For sequence The probability distribution, Represents the cluster center sequence The probability distribution, This represents a cluster containing a sequence of atmospheric pollutant concentrations. In the formula, This represents the sum of the outer products of all sequences in cluster C. It is a matrix in which all elements are 1. The number of sequences in the cluster. It is the identity matrix. for The transpose of .
[0028] 2) Subsequently, KSE assigns each sequence to a cluster associated with the initially selected center sequence, and then the center sequence of each cluster... The eigenvectors corresponding to the largest eigenvalue of matrix M are refined and updated. The eigenvectors corresponding to the largest eigenvalue of matrix M are derived from all sequences in the cluster. 3) Repeat the above steps until the central sequence becomes stable or the maximum number of iterations is reached; S4: Cluster analysis based on atmospheric monitoring data.
[0029] The clustering results are obtained by iterating through the determined center sequence until convergence.
[0030] This method can effectively handle outliers in atmospheric monitoring data and accurately assess potential connections and differences in air quality between cities.
[0031] Furthermore, the time complexity of this algorithm is as follows: Let The length of the original atmospheric sequence. The number of original atmospheric sequences. The number of processed sequences. The number of clusters. The time complexity of LTTB data processing. To measure the similarity between sequences, the algorithm first uses O-SBD to calculate... Item sequence and The distance between clusters, the time complexity of this step is O(n). During clustering, for each cluster, O-Shape calculates the cluster center matrix. The time complexity is The time complexity of eigenvalue decomposition is O(n). Therefore, the time complexity of this process is O(n). In summary, the overall time complexity of this algorithm is O(n log n). .
[0032] Differences and similarities with other methods: The O-Shape method proposed in this invention, applicable to atmospheric monitoring sequence data, differs from traditional shape-based clustering methods. k -Different shapes, etc. k O-Shape assesses the similarity between sequences and determines the center sequence by normalizing the cross-correlation coefficient. In this algorithm, O-Shape employs the LTTB downsampling algorithm to remove obvious outliers while preserving the original shape characteristics of the atmospheric sequence data, ensuring that key information is not lost. When processing urban air pollutant data, LTTB can effectively remove outliers caused by sudden industrial emissions or extreme weather conditions, preserving the overall trend of data changes and providing a more accurate basis for subsequent analysis. Furthermore, this invention enhances the robustness of distance to outliers through Theil-Sen estimation, proposing a new O-SBD distance to more accurately capture sequence shape characteristics. In atmospheric monitoring data, pollutant concentration changes are often influenced by multiple factors, such as meteorological conditions and industrial activities, and the data may contain a large number of outliers. The new O-SBD distance can effectively resist the interference of these outliers, more accurately reflecting the shape characteristics of the atmospheric pollutant concentration sequence, thereby improving the accuracy of cluster analysis. Secondly, this invention also proposes a new shape extraction strategy constrained by KL divergence to capture unique change patterns in atmospheric monitoring time series data, ensuring that the obtained cluster center sequences are consistent with the distribution characteristics of the original atmospheric monitoring data.
[0033] and kCompared to partition-based clustering methods such as Means, the O-Shape of this invention, when processing air pollution monitoring data, not only considers the correlation between sequences, scale expansion, and advection relationships, but also specifically considers the potential impact of outliers. For example, when analyzing changes in air pollutant concentrations in different cities, O-Shape can identify the correlation between pollutant concentration changes between cities, while considering the differences in pollutant concentration levels (scale expansion) and the transport of pollutants within the region (advection relationships), and effectively handles outliers in the data, providing more comprehensive clustering analysis results. Under the constraint of KL divergence, by decomposing the distance matrix obtained from O-SBD to determine the center sequence, this invention can preserve the shape characteristics of the sequence, which is crucial for air pollution prevention and control. For example, when formulating regional air pollution prevention and control strategies, O-Shape can classify cities with similar pollution change trends into the same category based on the shape characteristics of air pollutant concentration sequences, providing a scientific basis for targeted pollution prevention and control measures. Integrating these features, this invention provides a more comprehensive understanding of air pollution monitoring data, improves the efficiency of cluster centers, and is of great significance for formulating effective pollution prevention and control strategies. For example, in practical applications, O-Shape can help environmental protection departments quickly identify areas with serious air pollution problems and similar pollution characteristics, thereby concentrating resources on key areas for treatment and improving the effectiveness of air pollution prevention and control.
[0034] In summary, O-Shape, as an excellent clustering method for atmospheric monitoring sequence data, not only performs well in terms of complexity, occlusion, and uniform scaling, but also provides strong robustness and accuracy when processing atmospheric pollution monitoring data, making it a powerful tool in atmospheric pollution prevention and control.
[0035] Experimental Analysis: The atmospheric monitoring sequence data used in this invention comes from the China National Environmental Monitoring Centre, covering air pollution source monitoring stations in 10 adjacent cities in the Jiangsu, Zhejiang, and Shanghai regions, spanning from January to June 2023, and recording daily Air Quality Index (AQI). AQI, as a dimensionless index, is used to quantify air quality. Its calculation method is based on the maximum AQI value of the concentrations of six major pollutants (SO2, NO2, CO, PM10, PM2.5, and O3) at the same time. The AQI index for each pollutant is determined through a piecewise linear function. Although AQI can comprehensively reflect the degree of air pollution and air quality level in a specific region, its original classification method (divided into six levels) may lead to insufficient data for some levels, affecting the effectiveness of the analysis. Therefore, this section simplifies the levels into three main categories: 0-150 (Excellent), 151-300 (Good), and above 300 (Heavy Pollution) to ensure sufficient data for analysis in each category. This simplified method not only preserves the original AQI classification information but also makes the data distribution more uniform, making it suitable for data mining tasks such as cluster analysis.
[0036] Figure 3 This paper presents the clustering results of AQI sequences from 10 adjacent cities in the Yangtze River Delta region (Jiangsu, Zhejiang, and Shanghai), with different colored curves representing different clusters. From January to June 2023, the air quality in the selected cities was generally good, with AQI values mostly below 300, falling into the "good" category. This phenomenon may be related to the fact that these cities are located in the eastern coastal region of China, including economically developed cities such as Shanghai, Hangzhou, and Ningbo. The relatively good air quality in these areas may be attributed to local environmental protection measures and a high level of environmental awareness. Furthermore, the AQI indices of each city exhibit distinct temporal characteristics: a flat period in spring and a fluctuating period in summer. This indicates that air quality is closely related to climate factors, with rising temperatures having a significant impact on air quality. If extreme values are not considered... Figure 3 The data shows that the fluctuation range of AQI values in various cities has decreased, and the AQI itself has also shown a certain downward trend. This may be related to the ecological environment protection, energy conservation and emission reduction measures implemented in my country.
[0037] In contrast, Quzhou, Shaoxing, and Ningbo ranked last among the selected cities in terms of air quality during the first half of 2023. This may be related to the following factors: The geographical location of Quzhou City may lead to the influence of maritime climate, such as sea fog, which may aggravate the concentration of air pollutants.
[0038] Shaoxing has experienced rapid industrial and manufacturing development in recent years, which may lead to high pollutant emissions.
[0039] As an important port city, Ningbo faces significant air pollution from ship emissions and logistics transportation. Furthermore, Ningbo has a strong industrial base, particularly in chemical and machinery manufacturing, which may exert considerable environmental pollution pressure.
[0040] This study selected 27 geographically non-adjacent cities in Shaanxi, Gansu, and Ningxia, and used the O-Shape algorithm to perform cluster analysis on their atmospheric monitoring data. The O-Shape algorithm has significant advantages in handling outliers and is suitable for analyzing data with outliers caused by differences in air quality over a large area. The data in this study came from the China National Environmental Monitoring Centre and covered the average concentrations of six pollutants in the first half of 2023. By calculating the mean of the air quality index corresponding to the hourly average concentration, the pollutant concentrations were converted into AQI values, and then cluster analysis was performed. Table 1 shows the pollutant concentrations in major cities in the Shaanxi, Gansu, and Ningxia regions in 2023 (unit: μg / m³). 3 CO is mg / m³ 3 ).
[0041] Table 1
[0042] Figure 4 The clustering results of the O-Shape algorithm are visually demonstrated, with curves of different colors representing different clusters, and curves of the same color indicating cities belonging to the same cluster. Specifically: Cluster 1 (black curve) includes Zhongwei, Tianshui, Ankang, Hanzhong, Jiuquan, and Tongchuan. The AQI index of these cities ranges from 0 to 150, indicating that the air quality is excellent.
[0043] Cluster 2 (red curve) covers Wuzhong, Xianyang, Shangluo, Jiayuguan, Dingxi, Pingliang, Qingyang, Yan'an, Zhangye, Yulin, Wuwei, Weinan, Baiyin, Shizuishan, Xi'an, Jinchang, Yinchuan, and Longnan. The AQI index of these cities is generally greater than 300, indicating that the air quality is in a state of heavy pollution.
[0044] Cluster 3 (blue curve) includes Lanzhou, Guyuan, and Baoji, whose AQI index ranges from 150 to 300, representing good air quality.
[0045] Table 2 shows the cluster center results analysis. Further analysis of the concentration data of six major pollutants in the three types of clusters leads to the following conclusions: The first cluster includes Zhongwei City in Ningxia, Ankang, Hanzhong, and Tongchuan in Shaanxi, and Tianshui and Jiuquan in Gansu. These cities have relatively low concentrations of pollutants, especially PM2.5. 10These cities had the lowest concentrations of SO2, NO2, and CO, and their overall air quality was better than other cities. This is mainly attributed to two factors: first, these cities have smaller populations, are mostly third- or fourth-tier cities, have relatively lower levels of economic development, less industrial activity, and lower pollutant emissions; second, some of these cities are located in relatively dry areas, which is conducive to the dispersion of air pollutants, thus improving air quality.
[0046] The second cluster comprises 18 major cities in the three provinces of Shaanxi, Gansu, and Ningxia. These cities experience relatively severe air pollution, especially PM2.5. 10 The concentrations of SO2 and O3 are relatively high. A deeper analysis reveals the following main reasons: First, some cities have complex terrain, including mountains, hills, and plateaus, which causes air pollutants to linger in lower-lying areas, making it difficult for them to disperse and thus affecting air quality. Second, some cities are densely populated; for example, Xi'an has a current population of 12.9529 million and a floating population of 3.74 million, with large-scale population activity and traffic emissions exacerbating air pollution. Third, some cities are located on important transportation routes; for example, Jiayuguan and Zhangye are key cities along the ancient Silk Road and the modern Eurasian Land Bridge, with frequent population movement and logistics and transportation emissions significantly impacting air quality. Furthermore, Qingyang City is located at the junction of Shaanxi, Gansu, and Ningxia provinces, belonging to the Guanzhong Economic Belt along with Xianyang, Pingliang, and Weinan, with a strong industrial base and active economic activity; all these factors contribute to exacerbating air pollution.
[0047] Table 2
[0048] The third cluster consists of Lanzhou, the capital of Gansu Province; Guyuan, Ningxia; and Baoji, Shaanxi. The air quality in these cities is at a moderate level, neither the best nor the worst. A deeper examination of the underlying reasons reveals the following: Baoji City has a flat terrain and good air circulation, preventing pollutants from lingering for extended periods. This geographical advantage has alleviated air pollution to some extent, resulting in relatively good air quality in Baoji. However, topography is not the only factor determining air quality; economic activities and policy implementation also play crucial roles. For example, Baoji's industrialization and urban expansion in recent years may have negatively impacted air quality, requiring further research and attention.
[0049] Lanzhou, as an important industrial city in Northwest China, boasts a diverse and relatively complete economic system. However, this has also led to higher levels of PM2.5 in the air. 2.5Lanzhou had the highest average NO2 concentration, indicating generally poor air quality. Industrial emissions are one of the main factors affecting air quality in Lanzhou. Although Lanzhou has implemented a series of environmental protection measures, further efforts are needed in industrial restructuring and pollution control. Furthermore, traffic congestion in Lanzhou may also exacerbate air pollution, warranting further investigation and resolution.
[0050] Guyuan City is located at the intersection of Central Plains agricultural culture and northern nomadic culture, and is the only city in Ningxia Hui Autonomous Region not situated along the Yellow River. This unique geographical location results in relatively moderate population fluctuations and diversified economic activities in Guyuan. However, while this diversified economic activity drives urban development, it may also bring certain air pollution problems. Guyuan's air pollution level is generally moderate, which may be related to various factors such as local energy consumption structure, industrial emissions, and traffic emissions, requiring further analysis and research.
[0051] The following are recommendations for improving air quality in different types of cities: Category 1 Cities: Strengthen Environmental Protection Publicity: Popularize environmental protection knowledge through various channels to raise residents' environmental awareness and participation. Promote Renewable Energy: Formulate preferential policies to encourage residents and businesses to use renewable energy sources such as solar and wind power, reducing dependence on traditional fossil fuels such as coal and oil. Increase Urban Green Space: Develop urban greening plans to increase vegetation coverage areas such as parks and green spaces, enhancing the function of the urban ecosystem. Advocate Low-Carbon Travel: Improve public transportation systems, construct bicycle lanes and pedestrian paths, and encourage residents to choose low-carbon travel methods such as walking, cycling, or public transportation.
[0052] The second category of cities: Strict industrial emission control: Establish strict industrial emission standards, increase penalties for violating enterprises, and ensure that enterprises meet emission standards. Strengthen environmental supervision and enforcement: Establish and improve the environmental monitoring system, strengthen real-time monitoring of key areas and key enterprises, and improve the efficiency and deterrent effect of environmental law enforcement. Optimize urban transportation planning: Improve road capacity and reduce vehicle exhaust emissions through measures such as rationally planning road networks, optimizing traffic light settings, and developing intelligent transportation systems. Promote clean energy substitution: Increase the development and utilization of clean energy, reduce dependence on high-polluting fuels such as coal and oil, and improve the city's energy structure. Implement energy conservation and emission reduction measures: Encourage enterprises to adopt energy-saving equipment and technologies, improve energy efficiency, and reduce energy consumption and pollutant emissions.
[0053] The third category of cities: Strengthen air pollution monitoring and analysis: Establish a comprehensive air pollution monitoring network, strengthen real-time monitoring and data analysis of major pollutants, and provide a scientific basis for environmental management decisions. Promote green building and energy-saving technologies: Formulate green building standards and energy-saving technical specifications, encourage the adoption of energy-saving and environmentally friendly designs and materials in new buildings, and carry out energy-saving renovations on existing buildings. Encourage enterprises to implement cleaner production: Through policy guidance and financial support, encourage enterprises to adopt cleaner production technologies and processes to reduce the generation of industrial waste and pollutants at the source. Strengthen environmental planning and resource management: Formulate scientific and reasonable urban planning and industrial development planning, rationally allocate industrial zones and residential areas, optimize the use of land and water resources, and reduce the pressure of economic development on the environment.
[0054] To analyze the air quality relationships between geographically adjacent and non-adjacent cities more deeply, this embodiment applies the O-Shape of the present invention to atmospheric monitoring data from nearly 40 cities in the Jiangsu, Zhejiang, Shanghai, Shaanxi, Gansu, and Ningxia regions. This method aims to reveal the potential connections and differences in air quality between geographically proximate and non-adjacent cities. O-Shape, with its unique robust distance metric and shape extraction strategy, effectively handles the complex air quality patterns among these cities. Through O-SBD, the algorithm can identify cross-regional air quality similarity groups and accurately capture the changing trends of the air quality index among different cities. Experimental results show that the O-Shape algorithm performs excellently in processing this complex data. It successfully groups cities with similar air quality patterns, even those geographically distant. This capability is crucial for improving the reliability of air quality monitoring data and the accuracy of analysis results. The clustering results provide valuable information for environmental policymakers. By understanding the patterns of regional pollution propagation, policymakers can develop more effective air quality management and improvement strategies. For example, stricter emission control measures can be implemented in heavily polluted areas, while successful experiences can be promoted to guide environmental governance efforts in other cities for areas with better air quality.
[0055] The algorithm included in this invention innovatively introduces the maximum triangle three-segment downsampling algorithm to preprocess atmospheric sequence data, effectively preserving the basic shape of the atmospheric data. This step is crucial for processing pollutant information in atmospheric monitoring data, as it ensures that while removing redundant information, key data features important for air pollution prevention and control are not lost.
[0056] This invention proposes an O-Shape clustering method for atmospheric pollutant concentrations based on time-series shape features, targeting pollutant data from atmospheric monitoring. This method utilizes the normalized cross-correlation estimated by Theil-Sen to measure the relationship between series, constructing a novel shape-based distance O-SBD. This method is particularly suitable for processing atmospheric monitoring data because it can effectively identify and quantify the similarity between pollutant time series, even in the presence of outliers and noise. Through this method, we can more accurately perform clustering analysis on time-series data of atmospheric pollutants, thereby providing scientific data support for air pollution prevention and control.
[0057] This invention proposes a novel shape extraction strategy constrained by KL divergence for pollutant data in atmospheric monitoring. The strategy aims to ensure that the distribution of the central sequence is consistent with the atmospheric data in the cluster, thereby capturing the shape of the atmospheric sequence more accurately.
[0058] By applying this invention to atmospheric monitoring data analysis, the method combines pollutant characteristics and the air quality index to conduct in-depth analysis of air quality datasets, including multiple geographically adjacent and non-adjacent cities. Comparative analysis verifies the effectiveness of this method in actual air quality monitoring. Based on the urban characteristics revealed by clustering results, targeted air quality improvement measures are further proposed, aiming to significantly improve atmospheric environmental quality.
[0059] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for clustering atmospheric pollutant concentrations based on time-series shape features, characterized in that, Includes the following steps: The original atmospheric pollutant concentration time series was downsampled using the maximum triangle three-segment downsampling algorithm to obtain the downsampled sequence. Based on the downsampled sequence, the Theil-Sen estimation method is used to construct the normalized cross-correlation coefficient, and the O-SBD distance is obtained based on the normalized cross-correlation coefficient. Based on the O-SBD distance, a shape extraction strategy with KL divergence constraints is used to update the cluster center sequence and obtain the final center sequence. Cluster analysis of atmospheric pollutant concentration time series was performed based on the central sequence until convergence, and the clustering results were obtained.
2. The atmospheric pollutant concentration clustering method based on time series shape features according to claim 1, characterized in that, The process of obtaining the downsampled sequence includes: The maximum triangle three-segment algorithm is used to segment the original atmospheric pollutant concentration time series to obtain several data segments; Representative data points are selected from each data segment using an importance-based selection method; The downsampling sequence is obtained by combining the representative data points.
3. The atmospheric pollutant concentration clustering method based on time series shape features according to claim 1, characterized in that, The expression for constructing the normalized cross-correlation coefficient is as follows: ; In the formula, This represents the normalized cross-correlation coefficient, where X and Y represent two sequences. and This indicates that a fast Fourier transform was performed on two atmospheric pollutant concentration sequences. Indicates to Perform inverse Fourier transform Indicates to Perform Theil-Sen estimation storage.
4. The atmospheric pollutant concentration clustering method based on time series shape features according to claim 3, characterized in that, The formula for calculating the O-SBD distance is: ; In the formula, Let O-SBD be the distance between sequences X and Y.
5. The atmospheric pollutant concentration clustering method based on time series shape features according to claim 1, characterized in that, Based on the O-SBD distance, the process of updating the cluster center sequence and obtaining the final center sequence using the shape extraction strategy with KL divergence constraints includes: Each atmospheric sequence is assigned to a corresponding cluster based on the O-SBD distance; For each cluster, an optimization objective is constructed to balance the shape similarity and distribution consistency among the fused sequences, where KL divergence constraints are used to ensure that the probability distribution of the center sequence is consistent with the overall distribution of the sequences within the cluster. By solving the eigenvalue problem corresponding to the optimization objective, the eigenvector corresponding to the largest eigenvalue is used as the updated center sequence; The atmospheric sequence allocation and central sequence update steps are executed iteratively to obtain the final central sequence.
6. The atmospheric pollutant concentration clustering method based on time series shape features according to claim 5, characterized in that, The expression for calculating the central sequence is: ; In the formula, The optimal solution represents the cluster center sequence. Represents the cluster center sequence, This represents a sequence in the air pollution concentration dataset. Let M be the matrix. Let KL be the Lagrange multiplier, and KL denote the KL divergence. For sequence The probability distribution, Represents the cluster center sequence The probability distribution, This represents a cluster containing a sequence of atmospheric pollutant concentrations.
7. The atmospheric pollutant concentration clustering method based on time series shape features according to claim 6, characterized in that, The KL divergence constraint ensures consistent distributions by constructing a matrix M, the expression for which matrix M is calculated is: ; ; ; In the formula, This represents the sum of the outer products of all sequences in cluster C. It is a matrix in which all elements are 1. The number of sequences in the cluster. It is the identity matrix. for The transpose of .