A method and system for analyzing urban near-surface air pollution data

By performing multi-time scale component separation, dynamic time regularization and condensation hierarchical clustering of urban near-ground air pollution data, background atmospheric particulate concentration fluctuations are extracted, and the problem of analyzing complex pollution processes in the prior art is solved, and the accuracy and efficiency of the analysis are improved.

CN119337158BActive Publication Date: 2025-06-10GUANGDONG URBAN & RURAL PLANNING & DESIGN INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411211792.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-06-10
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively analyze urban near-ground air pollution data, especially in timing signal processing, it is difficult to decompose the composite changes caused by different emission sources and background factors, and it is impossible to localize the signals in the time and frequency domains at the same time, and it is affected by dynamic asynchronousness between sites.

Method used

A method for analyzing data of near-ground urban air pollution is proposed, including obtaining the time sequence data of atmospheric particulate concentrations of multiple monitoring stations, performing multi-time scale component separation, dynamic time regularization processing and condensation hierarchical clustering treatment to extract background atmospheric particulate concentration fluctuations.

Benefits of technology

By isolating and identifying the multi-scale components of single-site monitoring data through multiple time-scale components, dynamic asynchronousness between different monitoring sites is eliminated, background atmospheric particulate concentration fluctuations are extracted, and the analysis ability of complex pollution processes and the accuracy of data analysis is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119337158B_ABST
    Figure CN119337158B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of air pollution control, and discloses a method for analyzing urban near-surface air pollution data. The method includes obtaining time series data of atmospheric particulate matter concentration and performing multi-time scale component separation to obtain a first feature sequence of the time series data of atmospheric particulate matter concentration at different time scales; performing dynamic time warping processing on the first feature sequences between different monitoring stations to obtain a second feature sequence that eliminates the dynamic asynchrony between different monitoring stations; performing agglomerative hierarchical clustering processing on the second feature sequence to obtain the fluctuation characteristics of the background atmospheric particulate matter concentration; and analyzing the pollution process of atmospheric particulate matter on the urban near-surface according to the fluctuation characteristics of the background atmospheric particulate matter concentration. The present invention realizes a detailed analysis of the change characteristics of urban near-surface atmospheric particulate matter concentration on a minute-level time scale, provides a reliable technical basis for screening exogenous particulate matter pollution processes, and improves the efficiency and accuracy of air pollution analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air pollution control, and more specifically, to a method and system for analyzing urban near-surface air pollution data. Background Art

[0002] With China's large-scale energy conservation and emission reduction work carried out over the years, the emissions of industrial-source air pollutants have decreased significantly, and the overall urban air quality has been improved. However, the proportion of regional pollution sources has increased, and combined with the influence of adverse meteorological conditions, heavy pollution weather still occurs from time to time. In this context, the focus of air pollution control work has gradually shifted from controlling the total emissions of local pollutants to controlling the concentration of air pollutants in the affected areas. To achieve refined air pollution control and improve the urban spatial air environment, it is urgent to accurately grasp the movement and distribution of particulate matter in the urban near-surface under heavy pollution weather, distinguish exogenous and local pollution processes, and capture and characterize the influence of local pollution sources.

[0003] Currently, densely arranging a high-frequency sampling network of air pollutants near the urban surface is the main means to monitor and analyze the urban near-surface air pollution process. However, due to the simultaneous action of multiple air pollution processes with different sources and natures near the urban surface, as well as the interference of meteorological processes at different scales, conventional time-series signal processing methods are difficult to effectively analyze the data generated by such monitoring networks. Existing analysis methods, such as methods based on segmented statistical features, frequency-domain operation methods based on Fourier transform, and global time-frequency analysis methods such as empirical mode decomposition, all have their own limitations. For example, CN118070110A discloses an air pollutant monitoring method based on time-domain data analysis. It obtains the historical data of the pollutants to be monitored within a continuous period of the sampling period, uses a clustering method to divide the range of pollutant content into k state intervals, obtains the time-domain sequence D of the content of the pollutants to be monitored in the sampling, conducts a correlation analysis on the influencing factor data related to the content of the pollutants to be monitored, screens out the types of positive and negative influencing factors, classifies the sampling period according to the types of influencing factors contained, calculates the content state transition probability, obtains the content state transition probability matrix under different cycle states, and according to the cycle state and the content of the pollutants to be monitored at the current moment t, selects the corresponding state transition probability matrix to predict the content of the pollutants to be monitored and obtains the predicted value of the content of the pollutants to be monitored in the next cycle. However, the above methods either cannot effectively decompose the composite changes caused by different emission sources and background factors, or cannot localize the signal in both the time domain and the frequency domain at the same time, or are interfered by the dynamic asynchrony between stations and are difficult to make effective comparisons. These limitations lead to interference in the monitoring data analysis, which not only affects the monitoring accuracy but also increases the difficulty of extracting effective information from it to form accurate early warning indicators, and has the defects of low efficiency and low accuracy. Summary of the Invention

[0004] To overcome the defects of low efficiency and low accuracy existing in the prior art of air pollution analysis, the present invention proposes the following technical solutions:

[0005] In a first aspect, the present invention proposes a method for analyzing urban near-surface air pollution data, including:

[0006] Obtaining time-series data of atmospheric particulate matter concentration at multiple monitoring stations near the urban surface.

[0007] Performing multi-time scale component separation on the time-series data of atmospheric particulate matter concentration to obtain a first feature sequence of the time-series data of atmospheric particulate matter concentration at different time scales.

[0008] Performing dynamic time warping processing on the first feature sequences between different monitoring stations to obtain a second feature sequence that eliminates the dynamic asynchrony between different monitoring stations.

[0009] Performing agglomerative hierarchical clustering processing on the second feature sequence to obtain the fluctuation characteristics of background atmospheric particulate matter concentration.

[0010] Analyzing the pollution process of atmospheric particulate matter on the urban near-surface according to the fluctuation characteristics of background atmospheric particulate matter concentration.

[0011] As a preferred technical solution, after obtaining the time-series data of atmospheric particulate matter concentration at multiple monitoring stations near the urban surface, the method further includes:

[0012] Marking and removing the time-series data segments of atmospheric particulate matter concentration with more than M consecutive zero values, null values or unchanged values as abnormal segments. M is a positive integer

[0013] Obtaining the outliers of the time-series data of atmospheric particulate matter concentration by using the absolute median difference between the original value of the time-series data of atmospheric particulate matter concentration and the predicted value of the differential integrated moving average autoregression of the sequence, and replacing the outliers with the predicted value of the differential integrated moving average autoregression. Its expression is as follows:

[0014]

[0015] where c≈1.4826 is a constant, erfcinv() represents the inverse function of the complementary error function, spred(t) is the predicted value of s(t) under ARIMA(p, q), p is the order of the autoregressive term of the ARIMA model, q is the order of the differencing term of the ARIMA model. When res(t)>3, s(t) is considered an outlier and is replaced by spred(t).

[0016] As a preferred technical solution, multi-time scale component separation is performed on the time series data of atmospheric particulate matter concentration to obtain a first characteristic sequence of the time series data of atmospheric particulate matter concentration at different time scales, including:

[0017] The discrete wavelet decomposition method based on Meyer wavelet is used to decompose the time series data of atmospheric particulate matter concentration, and the wavelet approximation sequences of the time series data of atmospheric particulate matter concentration at different time scales are obtained. Its expression is as follows:

[0018]

[0019] where ψ(ω) is the Meyer wavelet basis function, j is the number of layers of the decomposition time scale, ω is the variable in the time domain representing the time point, v(·) = a 4 (35 - 84a + 70a 2 - 20a 3 ), a ∈ [0, 1] is the auxiliary function, a is the scale parameter in wavelet decomposition, z j is the low-frequency component of the time series data s of atmospheric particulate matter concentration, representing the wavelet approximation sequence of the time series data s of atmospheric particulate matter concentration at the time scale j, d n is the high-frequency component of the time series data s of atmospheric particulate matter concentration at the decomposition level n.

[0020] The wavelet approximation sequence is converted into a discrete first characteristic sequence, and its expression is as follows:

[0021]

[0022] where fp(·) is the discrete first characteristic sequence, z (n) (·) is the nth derivative of the first characteristic sequence z(·), t represents time, n 0 is the number of elements contained in the time series data of atmospheric particulate matter concentration, and N is a positive integer greater than 0.

[0023] As a preferred technical solution, dynamic time warping processing is performed on the first characteristic sequences between different monitoring stations to obtain a second characteristic sequence that eliminates the dynamic asynchrony between different monitoring stations, including:

[0024] The first characteristic sequences of two monitoring stations are selected respectively.

[0025] A distance matrix for storing the difference between any elements between the two first characteristic sequences is constructed, and its expression is as follows:

[0026] D(t 1 , t 2 ) = z 1 (t 1 ) - z 2 (t2 ),t 1 ∈ [1, m] ∩ N, t 2 ∈ [1, n] ∩ N

[0027] where D(t 1 , t 2 ) is the distance matrix of t 1 × t 2 , and z 1 (t 1 ) is the element of the first feature sequence z 1 of monitoring site 1 at time t 1 , z 2 (t 2 ) is the element of the first feature sequence z 2 of monitoring site 2 at time t 2 , m and n are the numbers of elements in the first feature sequence z 1 and the first feature sequence z 2 respectively, and N is a positive integer greater than 0.

[0028] Use the Floyd algorithm to calculate the shortest path from the upper left corner to the lower right corner of the distance matrix.

[0029] Determine the correspondence between the elements in the two first feature sequences according to the shortest path.

[0030] According to the correspondence, perform alignment processing on the two first feature sequences including local stretching or compression to obtain a second feature sequence that eliminates the dynamic asynchrony between the two monitoring sites.

[0031] Repeat the above steps to obtain the second feature sequences between any two monitoring sites at the same time scale.

[0032] As a preferred technical solution, perform agglomerative hierarchical clustering processing on the second feature sequence to obtain the background atmospheric particulate matter concentration fluctuation characteristics, including:

[0033] For the second feature sequence of any monitoring site, obtain the second feature sequence corresponding to any other monitoring site.

[0034] Statistical time positions and atmospheric particulate matter concentration values in the second feature sequences of other monitoring sites, and eliminate the second feature sequences whose time positions and atmospheric particulate matter concentration values do not meet the preset conditions.

[0035] Calculate the mean of the time positions and the mean of the atmospheric particulate matter concentrations of the remaining second feature sequences, construct a feature sequence using the mean of the time positions and the mean of the atmospheric particulate matter concentrations, and then perform maximum normalization on the constructed feature sequence to obtain a new second feature sequence of the current monitoring site.

[0036] Repeat the above steps to construct a new second feature sequence for each monitoring site.

[0037] Calculate the distance matrix between the new second feature sequences.

[0038] Based on the distance matrix between the new second feature sequences, use agglomerative hierarchical clustering to cluster the second feature sequences to obtain K types of clustering features. Here, K is the average length of the new second feature sequences of all monitoring sites.

[0039] Evaluate the intra-class similarity of all types of clustering features, and use the clustering features whose intra-class similarity meets the preset conditions as the final background atmospheric particulate matter concentration fluctuation features.

[0040] As a preferred technical solution, according to the background atmospheric particulate matter concentration fluctuation features, analyze the pollution process of atmospheric particulate matter on the urban near-surface, including:

[0041] According to the background atmospheric particulate matter concentration fluctuation features, divide the atmospheric particulate matter concentration fluctuation process into several pollution process segments.

[0042] Perform time series clustering processing on the pollution process segments to obtain different pollution process patterns.

[0043] Calculate the characteristic indexes of each monitoring site under the same pollution process pattern, and analyze the pollution process of atmospheric particulate matter on the urban near-surface according to the calculation results of the characteristic indexes.

[0044] As a preferred technical solution, divide the background atmospheric particulate matter concentration fluctuation features into several pollution process segments, including:

[0045] According to the peak and valley values of the atmospheric particulate matter concentration fluctuation, divide the atmospheric particulate matter concentration fluctuation into an ascending segment and a descending segment. Each individual ascending segment or descending segment represents a process of increasing or spreading of the atmospheric particulate matter concentration.

[0046] As a preferred technical solution, after dividing the background atmospheric particulate matter concentration fluctuation features into several pollution process segments, the method further includes performing minimum value normalization processing on the atmospheric particulate matter concentration in each pollution process segment.

[0047] As a preferred technical solution, perform time series clustering processing on the pollution process segments to obtain different pollution process patterns, including:

[0048] Calculate the time series difference D(s′ 1 ||s′ 2 ) between any two pollution process segments, and its expression is as follows:

[0049] D(s′ 1 ||s′ 2 ) = D DTW (s′ 1 ||s′ 2 ) × D shape (s′ 1 ||s′ 2 ) × D stretch (s′ 1 ||s′ 2 )

[0050] where s′ 1 and s′ 2 are respectively any two segments of the pollution process on the same time scale after min - max normalization, D DTW (s′ 1 ||s′ 2 ) and D shape (s′ 1 ||s′ 2 ) are respectively the warping distance and the shape distance after dynamic time warping, and D stretch (s′ 1 ||s′ 2 ) is the extension ratio of the original sequence on the time axis during the dynamic time warping process.

[0051] Construct a distance matrix between two segments of the pollution process according to the time - series difference.

[0052] Perform agglomerative hierarchical clustering on two segments of the pollution process according to the distance matrix between them to obtain clustering segments.

[0053] Repeat the above steps until the clustering segments reach the optimal number of classifications, and each clustering segment represents a pollution process pattern.

[0054] As a preferred technical solution, the method further includes calculating the maximum value of the silhouette coefficient SC of the agglomerative hierarchical clustering as the optimal number of classifications, and its expression is as follows:

[0055]

[0056] where is the mean of the silhouette values of all segments of the pollution process for agglomerative hierarchical clustering, and the silhouette value s(i) is defined as:

[0057]

[0058] where is the distance between the pollution process segment i and the clustering segment c iThe average distance of other pollution process segments, indicating the goodness or badness of assigning pollution process segment i to clustering segment c i is. is the average distance between pollution process segment i and pollution process segments in other clustering segments c i other than clustering segment c k and is the average dissimilarity between different clustering segments.

[0059] The beneficial effects of the present invention at least include:

[0060] By performing multi-time scale component separation on the time series data of atmospheric particulate matter concentration, the present invention obtains the first feature sequences at different time scales, effectively identifies the multi-scale components of single-site monitoring data, and improves the analysis ability of complex pollution processes. Dynamic time warping processing is used to eliminate the dynamic asynchrony between different monitoring sites, solves the problem that traditional methods are difficult to effectively compare multi-site data, and improves the accuracy of data analysis. Agglomerative hierarchical clustering processing is used to extract the background atmospheric particulate matter concentration fluctuation characteristics, effectively excludes the influence of irrelevant interfering factors, and improves the extraction efficiency of exogenous particulate matter pollution process characteristics. Analyzing the atmospheric particulate matter pollution process based on the extracted background fluctuation characteristics provides a more accurate method for distinguishing exogenous and local pollution processes, and improves the analysis accuracy of the urban near-surface pollution process. In summary, through a comprehensive analysis method combining time domain, frequency domain and time-frequency joint, the present invention realizes a detailed analysis of the variation characteristics of urban near-surface atmospheric particulate matter concentration on the minute time scale, provides a reliable technical basis for screening exogenous particulate matter pollution processes, and improves the efficiency and accuracy of air pollution analysis. Description of the Drawings

[0061] Figure 1 is a schematic flow chart of the method for analyzing urban near-surface air pollution data provided in Embodiment 1.

[0062] Figure 2 is a logical architecture diagram of the time series characteristics of urban near-surface PM2.5 concentration data in Embodiment 2.

[0063] Figure 3 is a schematic diagram of the results of multi-scale autocorrelation feature analysis of monitoring data in Embodiment 2.

[0064] Figure 4 is a power spectral density estimation diagram of monitoring data in Embodiment 2.

[0065] Figure 5 is a schematic diagram of the stationarity characteristics of monitoring data segments under different sample lengths and wavelet decomposition levels in Embodiment 2.

[0066] Figure 6It is the continuous wavelet decomposition result diagram of the atmospheric particulate matter monitoring data at the monitoring sites in Example 2.

[0067] Figure 7 It is the change diagram of the numerical difference and the temporal morphological difference between the monitoring data segments at each monitoring site in Example 2 with the time scale.

[0068] Figure 8 It is the logical architecture diagram for processing the monitoring data of each monitoring site in Example 3.

[0069] Figure 9 It is the result schematic diagram of the outlier marking and substitution example of the monitoring data in Example 3.

[0070] Figure 10 It is the schematic diagram of the principle of dynamic time warping in Example 3.

[0071] Figure 11 Schematic diagram of the relationship between the asynchrony index and the severity index and the change of PM2.5 concentration in Example 3. Detailed implementation manners

[0072] The following will illustrate the implementation manners of the present invention with reference to the accompanying drawings and the preferred technical solutions. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred technical solutions are only for illustrating the present invention rather than limiting the protection scope of the present invention.

[0073] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, number, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0074] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0075] Example 1

[0076] This example proposes a method for analyzing urban near-surface air pollution data, as Figure 1 shown, Figure 1The flowchart shows a method for analyzing urban near-surface air pollution data provided in this embodiment. The method includes the following steps:

[0077] S1: Obtain the time-series data of atmospheric particulate matter concentrations at multiple monitoring stations near the urban surface.

[0078] In the specific implementation process, the atmospheric particulate matter concentration is usually mainly the PM2.5 concentration.

[0079] S2: Perform multi-time scale component separation on the time-series data of atmospheric particulate matter concentrations to obtain the first feature sequences of the time-series data of atmospheric particulate matter concentrations at different time scales.

[0080] In the specific implementation process, apply discrete wavelet transform to the time-series data of atmospheric particulate matter concentrations at each monitoring station. Select appropriate wavelet basis functions and decomposition levels, and decompose the original time-series data into approximation coefficients and detail coefficients at different time scales. These coefficients constitute the first feature sequences at different time scales.

[0081] S3: Perform dynamic time warping processing on the first feature sequences between different monitoring stations to obtain the second feature sequences that eliminate the dynamic asynchrony between different monitoring stations.

[0082] In the specific implementation process, perform dynamic time warping (DTW) processing on the first feature sequences of different monitoring stations at the same time scale. The DTW algorithm eliminates the dynamic asynchrony between different stations by finding the best alignment between two time series. The sequence obtained after processing is the second feature sequence.

[0083] S4: Perform agglomerative hierarchical clustering processing on the second feature sequences to obtain the background atmospheric particulate matter concentration fluctuation characteristics.

[0084] In the specific implementation process, perform agglomerative hierarchical clustering (AHC) processing on the second feature sequences. First, calculate the distance matrix between sequences, and then use the AHC algorithm for clustering. During the clustering process, determine the optimal number of clusters through the silhouette coefficient. The clustering result represents the background atmospheric particulate matter concentration fluctuation characteristics.

[0085] S5: Analyze the pollution process of atmospheric particulate matter on the urban near-surface according to the background atmospheric particulate matter concentration fluctuation characteristics.

[0086] In the specific implementation process, based on the obtained background fluctuation characteristics, analyze the spatio-temporal relationships of each monitoring station, distinguish exogenous and local pollution processes, and describe the pollution process of atmospheric particulate matter on the urban near-surface.

[0087] It can be understood that in this embodiment, by performing multi-time scale component separation on the time series data of atmospheric particulate matter concentration, the first feature sequences at different time scales are obtained, effectively identifying the multi-scale components of single-site monitoring data and improving the analysis ability of complex pollution processes. Dynamic time warping processing is used to eliminate the dynamic asynchrony between different monitoring sites, solving the problem that traditional methods are difficult to effectively compare multi-site data and improving the accuracy of data analysis. Agglomerative hierarchical clustering processing is used to extract the background atmospheric particulate matter concentration fluctuation characteristics, effectively excluding the influence of irrelevant interference factors and improving the extraction efficiency of exogenous particulate matter pollution process characteristics. Based on the extracted background fluctuation characteristics, the atmospheric particulate matter pollution process is analyzed, providing a more accurate method for distinguishing exogenous and local pollution processes and improving the analysis accuracy of the urban near-surface pollution process. In summary, through a comprehensive analysis method combining time domain, frequency domain, and time-frequency joint analysis, this embodiment realizes a detailed analysis of the variation characteristics of urban near-surface PM2.5 concentration on a minute-level time scale, provides a reliable technical basis for screening exogenous particulate matter pollution processes, and improves the efficiency and accuracy of air pollution analysis.

[0088] Embodiment 2

[0089] Based on the urban near-surface air pollution data analysis method proposed in Embodiment 1, this embodiment makes a detailed analysis of the time series characteristics of urban near-surface PM2.5 concentration data. As Figure 2 shown, Figure 2 This is the logical architecture diagram of the time series characteristics of urban near-surface PM2.5 concentration data in this embodiment.

[0090] Regarding the time series data characteristics of a single monitoring site, this embodiment starts from the autocorrelation, local stationarity, and frequency component characteristics in the time domain, analyzes the time series data of each site, clarifies the mathematical characteristics of the single-site data sequence, and provides a basis for establishing a suitable representation method. This analysis method helps to deeply understand the internal characteristics and variation laws of PM2.5 concentration data.

[0091] As an exemplary illustration, in the autocorrelation feature analysis, the autocorrelation function is used to analyze and obtain the periodic characteristics of the PM2.5 concentration time series data. The results show that there are no significant daily and weekly periodic characteristics, but there are periodic signs in the long time series. This discovery is of great significance for understanding the long-term variation trend of PM2.5 concentration.

[0092] Autocorrelation is the degree of correlation between a time series and itself lagged by a certain continuous time period. The autocorrelation sequence of a periodic signal has the same periodic characteristics as the signal itself. Measurement uncertainties and noise sometimes make it difficult to detect the periodic oscillation behavior in the original signal, even if such periodic behavior is expected to exist; autocorrelation can help verify the existence of a period and determine its length. This method is particularly useful when dealing with complex environmental data because it can reveal potential periodic patterns that may be masked by random fluctuations.

[0093] For a time series {Xt} (t = 1, 2,...), its autocorrelation function is expressed as:

[0094] R(Δt) = E(X t ×X t +Δt)

[0095] Taking the fixed-minute monitoring data of each station as Xt and calculating Rx(Δt) (taking Δt ∈ [1, 60×24×365×1.5]), the result is as Figure 3 shown. Figure 3 This is the schematic diagram of the result of multi-scale autocorrelation feature analysis of the monitoring data in this embodiment. This calculation process covers a time range from 1 minute to 1.5 years, providing a comprehensive time-scale analysis of the PM2.5 concentration change.

[0096] The results show that the autocorrelation function of the original signal does not exhibit obvious daily or weekly periodicity. Instead, it starts from 1 and rapidly drops to about 0.4 within 4,000 minutes (about 3 days), then oscillates and drops, and stabilizes at about -0.08 at 200,000 minutes (about 140 days). This pattern indicates that the PM2.5 concentration has strong autocorrelation in the short term (within 3 days), but this correlation weakens rapidly over time. In the long time series, there are signs of a cycle of about 7 months, but the values of Rx are not significant. This finding implies that the PM2.5 concentration may be affected by seasonal factors, but this effect is not very strong.

[0097] Similar conclusions can be obtained through power spectral density estimation, as Figure 4 shown. Figure 4 This is the power spectral density estimation diagram of the monitoring data in this embodiment. Power spectral density estimation provides another representation of the signal in the frequency domain and can reveal the periodic components in the time series. The specific analysis results are as follows:

[0098] Obvious minima can be observed at 0.30 - 0.80π rad / sample, corresponding to significant inter-seasonal oscillations. This indicates that the PM2.5 concentration has obvious changes on the seasonal scale.

[0099] There are maxima near 0.57π rad / sample and 0.71π rad / sample, corresponding to relatively weak inter-annual variability. This indicates that there are also certain periodic variations in PM2.5 concentration on the annual scale, but they are not as significant as the seasonal variations.

[0100] There are maxima at 0.013π rad / sample, 0.026π rad / sample, 0.037π rad / sample, 0.048π rad / sample, etc., corresponding to weak weekly oscillations. These higher-frequency oscillations may reflect the periodic variations in PM2.5 concentration on a shorter time scale, which may be related to the weekly human activity patterns.

[0101] As an exemplary illustration, in the local stationarity feature analysis, through discrete wavelet decomposition and the ADF unit root test (Augmented Dickey-Fuller test, ADF), the stationarity of PM2.5 concentration time series with different frequency components and different lengths is analyzed to find the time scales at which different frequency components act on the temporal variation of PM2.5 concentration. This analysis method can help to understand more deeply the variation characteristics of PM2.5 concentration on different time scales.

[0102] Specifically, stationarity reflects whether the process generating the time series changes over time. For a time series {Xt} (t = 1, 2,...), when its moving mean E(Xt), variance var(Xt), and covariance cov(Xt, Xt+k) are all independent of time t, then Xt can be considered a weakly stationary time series. If this time series is generated by a certain stochastic process, then this stochastic process is a stationary stochastic process, that is, the statistical properties of the process generating the time series do not change over time. That is, if Xt is the sequence of PM2.5 concentration changing over time in a certain area within a certain time range, then it can be inferred that the factors affecting the PM2.5 concentration in this area have not changed significantly during this period. This stationarity analysis is crucial for understanding the long-term variation trend and influencing factors of PM2.5 concentration.

[0103] First, different frequency components of each station are extracted through discrete wavelet decomposition, and the frequency component at decomposition level j is set as dj; dj can be regarded as the oscillation characteristics of the original time series at the 2j scale, where the unit of the time scale is consistent with the interval of the time series, that is, 1 minute. Discrete wavelet decomposition allows the complex time series to be decomposed into components of different frequencies, thus better understanding the variation characteristics on different time scales.

[0104] Secondly, dj is segmented into data segments of n hours (n ∈ [1, 24] ∩ N) in chronological order. After dj is segmented into time segments of n hours in length, the m-th segment is denoted as segj,n,m. This segmentation method enables the analysis of the stationarity characteristics of PM2.5 concentration within different time windows.

[0105] The ADF is used to examine whether any segj,n,m is a stationary time series. The principle is that assuming Xt is autoregressive, a point on this sequence can be expressed as:

[0106]

[0107] Its characteristic equation is:

[0108]

[0109] If all the characteristic roots of this equation are inside the unit circle, that is:

[0110] λ p |<1, i ∈ [1, p] ∩ N

[0111] Then the original time series is stationary; otherwise, the series is non-stationary.

[0112] For each decomposition level j and data segment length (60 × n), calculate the proportion of non-stationary segments in all segments of segj,n,m that meet the conditions, and organize it into Figure 5 , Figure 5 which is the schematic diagram of the stationarity characteristics of the monitoring data segments under different sample lengths and wavelet decomposition levels in this embodiment.

[0113] As Figure 5 shown, 1) When j = 1 or j = 2, almost all segj,n,m are chronologically stationary regardless of the length of the data segments constructed. 2) When j ≥ 3 and j ≤ 5, most of the shorter data segments are non-stationary, but the proportion of non-stationary data segments in all segments decreases as the length of the data segments increases until it drops to nearly 0. 3) When j ≥ 6, the vast majority of the data segments are non-stationary and the proportion of non-stationary data segments hardly changes with the change in the length of the data segments. These results reveal the stationarity characteristics of PM2.5 concentration on different time scales and frequency components, providing important clues for understanding the dynamic characteristics of the pollution process.

[0114] The results show that the fluctuation law of the PM2.5 concentration at the site is relatively stable at high frequencies (such as 4 minutes). The fluctuation law at general frequencies (such as 832 minutes) will change significantly on the 116-hour scale. Short-term pollution processes or pollution events will be reflected in these frequency components. If we want to completely characterize the local change characteristics of such pollution processes in the time domain, the sampling frequency of the sensor should at least reach this level. The fluctuation law at low frequencies (such as 64 minutes) is almost significantly changing regardless of the time scale observed. Therefore, if we want to completely characterize the long-time series local change characteristics of atmospheric particulate matter, the sensor should at least reach a sampling frequency of about 1 hour. These findings have important practical significance for the design and optimization of PM2.5 monitoring systems and can guide the selection of appropriate sampling frequencies and data analysis methods.

[0115] As an exemplary illustration, continuous wavelet transform is used in the analysis of local frequency component characteristics to examine the relationship between the time-frequency characteristics of time series data and specific events. The results show that the oscillation of the PM2.5 concentration on a long time scale is related to exogenous pollution processes, while high-frequency oscillations are often related to local pollution events. This analysis method can help distinguish different types of pollution events and provide a scientific basis for formulating targeted pollution control strategies.

[0116] For the time series {Xt} (t = 1, 2,...), the result of its continuous wavelet transform can be expressed as a binary function, and there is:

[0117]

[0118] where a is the scale variable of the frequency component, and b represents the translation factor of the position in the time domain. represents the base number of the wavelet transform. The modulus and sign of the real part of ψf(a, b) express the oscillation intensity and direction of the original signal at a certain time point and a certain frequency.

[0119] Such as Figure 6 shown, Figure 6This is the continuous wavelet decomposition result graph of the atmospheric particulate matter monitoring data at the monitoring site in this embodiment. 1) For most of the time, the oscillation of the PM2.5 concentration is mainly manifested in hours and larger scales, and the intensity of high-frequency oscillation is relatively small. This phenomenon reflects that the main characteristics of the PM2.5 concentration change usually occur on a relatively long time scale. At the same time, the oscillation characteristics above several hours show obvious time-domain locality, and higher-intensity oscillations can be observed before and after the known exogenous particulate matter pollution process. This finding indicates that the impact of exogenous pollution events on the PM2.5 concentration is not only significant during the event but also causes obvious fluctuations before and after the event. 2) In addition to being related to the increase in the concentration of atmospheric particulate matter in the early morning under the diurnal periodic change of the atmospheric boundary layer, the high-frequency oscillation also has an obvious corresponding relationship with short-time events that cause the increase in the concentration of atmospheric particulate matter. This observation reveals that high-frequency oscillation may reflect two different phenomena: one is the regular change caused by the diurnal variation of the atmospheric boundary layer, and the other is the sudden change caused by local short-term events (such as traffic peaks, industrial activities, etc.).

[0120] For the comparison of time series data among multiple monitoring sites, in this embodiment, the time series data of each site is processed into approximate sequence segments on different time scales, and at each time scale, the approximate sequences of different sites within the same time segment range are compared pairwise to analyze their similarity and asynchrony. The results show that the temporal variations of the PM2.5 concentration at different sites are controlled by the same exogenous pollution process on a relatively large time scale, but the heterogeneity caused by local pollution events is relatively large. This multi-site, multi-time scale comparative analysis method provides a new perspective for comprehensively understanding the spatial distribution characteristics of urban PM2.5 pollution.

[0121] This embodiment compares based on two approximate sequence segments and The value difference and shape difference between them are respectively defined by the following formulas:

[0122]

[0123] where N is the number of elements contained in the two segments. The distribution ranges of D value and D shape under different time scales are statistically analyzed, and the results are as Figure 7 shown, Figure 7 is the graph of the variation of the numerical difference and temporal morphological difference between the monitoring data segments of each monitoring site with the time scale.

[0124] It can be understood that through the analysis of the overall characteristics of the time-series data in this embodiment, it is found that the time-series variation of the urban near-surface PM2.5 concentration has the characteristics of weak periodicity, strong localness, multi-scale nature, and dynamic asynchrony among stations. At the same time, it is also affected by some short-term and accidental events. These characteristics reflect the complexity and diversity of PM2.5 pollution, indicating that it is affected by multiple factors, including but not limited to meteorological conditions, human activities, geographical features, etc.

[0125] Weak periodicity indicates that although the PM2.5 concentration may be affected by certain periodic factors (such as daily or weekly human activity patterns), this periodicity is not significant or consistent. Strong localness means that there may be significant differences in the PM2.5 concentration at different locations, which may be caused by different local pollution sources or microenvironmental conditions. Multi-scale nature indicates that the changes in the PM2.5 concentration occur on different time scales, and there may be short-term fluctuations and long-term trends. The dynamic asynchrony among stations reflects the time lag or lead relationship of the PM2.5 concentration changes between different monitoring stations, which may be related to the transport and diffusion process of pollutants.

[0126] Through the oscillation frequencies of different time-series components, it is possible to distinguish, to a certain extent, long-term and background changes from short-term and local changes, but there is still a large overlapping area in the frequency domain between the two. This overlap means that relying solely on frequency analysis may not be able to fully distinguish different types of pollution processes, and other analysis methods need to be combined to obtain a more comprehensive understanding.

[0127] Therefore, to effectively extract the characteristics of the exogenous particulate matter pollution process near the urban surface, it is necessary to comprehensively utilize time-domain and frequency-domain information such as the multi-scale components of a single station, the similarity and asynchronous relationship among stations, and try to exclude the influence of irrelevant interfering factors. This comprehensive analysis method can help to more accurately identify and characterize exogenous pollution events, and at the same time, it can also distinguish their influence from local pollution sources.

[0128] This embodiment proposes a method that combines time-domain, frequency-domain, and time-frequency joint analysis, and elaborately analyzes the characteristics of the urban near-surface PM2.5 concentration change on a minute-level time scale, providing a technical basis for screening exogenous particulate matter pollution processes. The innovation of this method lies in that it combines multiple data analysis techniques, including but not limited to time series analysis, spectrum analysis, wavelet analysis, etc., to comprehensively capture all aspects of the PM2.5 concentration change. By analyzing on a minute-level time scale, this method can capture short-term and rapid pollution events, which is particularly important for understanding and predicting short-term fluctuations in urban air quality.

[0129] Embodiment 3

[0130] This embodiment makes improvements on the basis of the urban near-surface air pollution data analysis method proposed in the above embodiment. As Figure 8 shown, Figure 8 is the logical architecture diagram for processing the monitoring data of each monitoring station in this embodiment.

[0131] In this embodiment, after obtaining the time series data of the atmospheric particulate matter concentration at multiple monitoring stations in the urban near-surface area, the time series data segments of the atmospheric particulate matter concentration with more than 4 consecutive 0 values, null values or unchanged values are marked as abnormal segments and removed. And the absolute median difference between the original value of the time series data of the atmospheric particulate matter concentration and the differential integrated moving average autoregressive prediction value of the sequence is used to obtain the outliers of the time series data of the atmospheric particulate matter concentration, and the differential integrated moving average autoregressive prediction value is used to replace the outliers. The expression is as follows:

[0132]

[0133] where c≈1.4826 is a constant, erfcinv() represents the inverse function of the complementary error function, spred(t) is the predicted value of s(t) under ARIMA(p, q), p is the order of the autoregressive term of the ARIMA model, q is the order of the differencing term of the ARIMA model. When res(t)>3, s(t) is considered an outlier and is replaced by spred(t). As Figure 9 shown, Figure 9 is the result schematic diagram of the outlier marking and replacement example of the monitoring data in this embodiment.

[0134] In this embodiment, the multi-time scale component separation is performed on the time series data of the atmospheric particulate matter concentration to obtain the first feature sequence of the time series data of the atmospheric particulate matter concentration at different time scales.

[0135] The discrete wavelet decomposition method based on the Meyer wavelet is used to decompose the time series data of the atmospheric particulate matter concentration to obtain the wavelet approximation sequences of the time series data of the atmospheric particulate matter concentration at different time scales. Among them, the expression of the Meyer wavelet basis function ψ(ω) is as follows:

[0136]

[0137] where ψ(ω) is the Meyer wavelet basis function, j is the number of layers of the decomposition time scale, ω is the variable in the time domain representing the time point, v(·)=a 4 (35 - 84a + 70a 2 - 20a 3 ), a∈[0, 1] is an auxiliary function, and a is the scale parameter in the wavelet decomposition.

[0138] At time scale j, the time series data s of atmospheric particulate matter concentration will be decomposed into a low-frequency component (wavelet approximation), denoted as z, by the Mallat tower decomposition algorithm (Mallat Tower Decomposing Algorithm). j , and j high-frequency components (wavelet details) d n , and their expressions are as follows:

[0139]

[0140] Among them, z j is the low-frequency component of the time series data s of atmospheric particulate matter concentration, representing the wavelet approximation sequence of the time series data s of atmospheric particulate matter concentration at time scale j, and d n is the high-frequency component of the time series data s of atmospheric particulate matter concentration at decomposition level n, where j ∈ [4, 12] ∩ N.

[0141] To align the time series data of different stations and cancel the dynamic asynchronous characteristics between stations, the dynamic time warping method needs to be used to process each pair of station pairs subsequently. However, the minimum time complexity of this algorithm is O(m × n), and the time consumed by its calculation will increase quadratically with the length of the original time series. For the fine-grained, long time series used in this embodiment, this exceeds the processing capacity of a general desktop workstation. To address this problem, in this embodiment, the features of each wavelet approximation sequence are simplified and expressed by the approximate 0 points of their multi-order derivatives. That is, for the wavelet approximation sequence z j , the discrete first feature sequence fp(t) is obtained. By experimenting with the information loss of the feature point sequence compared to the original wavelet approximation sequence when n 0 takes different values, this embodiment determines that n 0 = 3. Its expression is as follows:

[0142]

[0143] Among them, fp(·) is the discrete first feature sequence, z (n) (·) is the nth derivative of the first feature sequence z(·), t represents time, and n 0 is the number of elements contained in the time series data of atmospheric particulate matter concentration, and N is a positive integer greater than 0.

[0144] In this embodiment, the dynamic time warping process is performed on the first feature sequences between different monitoring stations to obtain a second feature sequence that eliminates the dynamic asynchrony between different monitoring stations, including:

[0145] Select the first feature sequences of two monitoring stations respectively.

[0146] Construct a distance matrix for storing the difference between any elements of two first feature sequences, and its expression is as follows:

[0147] D(t 1 , t 2 ) = z 1 (t 1 ) - z 2 (t 2 ), t 1 ∈[1, m] ∩ N, t 2 ∈[1, n] ∩ N

[0148] where D(t 1 , t 2 ) is the distance matrix of t 1 ×t 2 , z 1 (t 1 ) is the element of the first feature sequence z 1 of monitoring site 1 at time t 1 , z 2 (t 2 ) is the first feature sequence z 2 of monitoring site 2 at time t 2 , m and n are the numbers of elements in the first feature sequence z 1 and the first feature sequence z 2 respectively, and N is a positive integer greater than 0.

[0149] Use the Floyd algorithm to calculate the shortest path from the upper left corner to the lower right corner of the distance matrix.

[0150] Determine the correspondence between the elements in the two first feature sequences according to the shortest path.

[0151] According to the correspondence, perform alignment processing on the two first feature sequences including local stretching or compression to obtain a second feature sequence that eliminates the dynamic asynchrony between the two monitoring sites.

[0152] Repeat the above steps to obtain the second feature sequences between any two monitoring sites at the same time scale.

[0153] As an exemplary illustration, use the Floyd algorithm to solve the shortest path from the upper left corner of the distance matrix D to the lower right corner of the matrix. If this shortest path completely coincides with the diagonal in this direction, it means that the two sequences are already completely aligned. A path offset from the diagonal means local stretching or compression done to align the two sequences. If this path passes through a point D(t, t') on the matrix, this indicates that the element s 1 (t) is aligned to the element s 2On (t′). As Figure 10 shown, Figure 10 This is the schematic diagram of dynamic time warping in this embodiment.

[0154] It should be noted that after warping, the sum of the absolute values of the point-by-point differences between the new sequences (which can be called the warping distance) represents the similarity of the original sequences after alignment. In this embodiment, it is the similarity of the wavelet approximation sequences of two stations after correcting the dynamic asynchrony between stations, and its expression is as follows:

[0155]

[0156] The path sequence in the distance matrix D represents the correspondence between the data points in the original sequence, and its expression is as follows:

[0157] P DTW (s 1 ||s 2 ) = {(t 1 , t′ 1 ), (t 2 , t′ 2 ),..., (t m , t′ m )}

[0158] In this embodiment, through the above correspondence, it can be known that for monitoring station 1, if there is a certain change feature at time t on a certain time scale, then at monitoring station 2, there is a similar change feature at time t′ on the same time scale. The wavelet approximation feature point sequences fp of each station under the same time scale are processed by the above method, and the alignment relationship between the wavelet approximation feature point sequences between any two stations under the same time scale is formed.

[0159] In this embodiment, the second feature sequence is subjected to agglomerative hierarchical clustering processing to obtain the background atmospheric particulate matter concentration fluctuation characteristics, including:

[0160] For any monitoring station m 0 of the second feature sequence obtain the second feature sequence corresponding to any other monitoring station, such as

[0161] statistically analyze the time position t m and the atmospheric particulate matter concentration value in the second feature sequences of other monitoring stations. The atmospheric particulate matter concentration value in this embodiment is the PM2.5 concentration value p m , and according to the 3σ principle, the second feature sequences whose time position and atmospheric particulate matter concentration value do not meet the conditions are eliminated.

[0162] Calculate the mean of the time positions and the mean of the atmospheric particulate matter concentrations of the remaining second feature sequences, and construct a feature sequence using the mean of the time positions and the mean of the atmospheric particulate matter concentrations. vectors of (t m , p m ), and then perform maximum normalization on the constructed feature sequence to obtain a new second feature sequence for the current monitoring site.

[0163] Repeat the above steps to construct a new second feature sequence for each monitoring site.

[0164] Calculate the distance matrix between the new second feature sequences.

[0165] According to the distance matrix between the new second feature sequences, use Agglomerative Hierarchical Clustering (AHC) to cluster the second feature sequences to obtain K types of clustering features. Where K is the average length of the new second feature sequences of all monitoring sites and is the optimal number of classifications for the agglomerative hierarchical clustering in this embodiment.

[0166] Evaluate the within-class similarity of all types of clustering features, and use the clustering features whose within-class similarity meets the preset conditions as the final background atmospheric particulate matter concentration fluctuation features. In this embodiment, the clustering center of the remaining clusters and the distribution range of each class are used as the extraction results of the background atmospheric particulate matter concentration fluctuation features at this time scale.

[0167] In this embodiment, according to the background atmospheric particulate matter concentration fluctuation features, analyze the pollution process of atmospheric particulate matter on the urban near-surface, including:

[0168] According to the background atmospheric particulate matter concentration fluctuation features, divide the atmospheric particulate matter concentration fluctuation process into several pollution process segments.

[0169] Perform time series clustering processing on the pollution process segments to obtain different pollution process patterns.

[0170] Calculate the characteristic indicators of each monitoring site under the same pollution process pattern, and analyze the pollution process of atmospheric particulate matter on the urban near-surface according to the calculation results of the characteristic indicators.

[0171] In this embodiment, dividing the background atmospheric particulate matter concentration fluctuation features into several pollution process segments includes:

[0172] According to the peak and valley values of the atmospheric particulate matter concentration fluctuation, divide the atmospheric particulate matter concentration fluctuation into an ascending segment and a descending segment, and each individual ascending segment or descending segment represents a process of increasing or spreading of the atmospheric particulate matter concentration.

[0173] In the specific implementation process, the fluctuation characteristics of the background atmospheric particulate matter concentration of the entire monitoring network under the action of external particulate matter are divided into separate rising segments and falling segments according to the peaks and valleys. One separate rising segment or falling segment represents a process of accumulation and increase or diffusion and reduction of the PM2.5 concentration at a corresponding time scale.

[0174] In this embodiment, after dividing the fluctuation characteristics of the background atmospheric particulate matter concentration into several pollution process segments, the method further includes performing minimum value normalization processing on the atmospheric particulate matter concentration in each pollution process segment, and its expression is as follows:

[0175] s′ = s - min(s)

[0176] In this embodiment, time series clustering processing is performed on the pollution process segments to obtain different pollution process patterns, including:

[0177] Calculate the time series difference D(s′ 1 ||s′ 2 ) between any two pollution process segments, and its expression is as follows:

[0178] D(s′ 1 ||s′ 2 ) = D DTW (s′ 1 ||s′ 2 ) × D shape (s′ 1 ||s′ 2 ) × D stretch (s′ 1 ||s′ 2 )

[0179] Where s′ 1 and s′ 2 are respectively any two pollution process segments at the same time scale after minimum value normalization processing, D DTW (s′ 1 ||s′ 2 ) and D shape (s′ 1 ||s′ 2 ) are respectively the warping distance and shape distance after dynamic time warping, and D stretch (s′ 1 ||s′ 2 ) is the extension ratio of the original sequence on the time axis during the dynamic time warping process.

[0180] According to the time series difference, a distance matrix between two pollution process segments is constructed.

[0181] Based on the distance matrix between two pollution process segments, hierarchical agglomerative clustering is performed on the two pollution process segments to obtain clustered segments.

[0182] Repeat the above steps until the number of clustered segments reaches the optimal classification number, and each clustered segment represents a pollution process pattern.

[0183] In this embodiment, the method further includes calculating the maximum value of the silhouette coefficient SC of the hierarchical agglomerative clustering as the optimal classification number, and its expression is as follows:

[0184]

[0185] where is the mean of the silhouette values of all pollution process segments for hierarchical agglomerative clustering, and the silhouette value s(i) is defined as:

[0186]

[0187] where is the average distance between pollution process segment i and other pollution process segments in clustered segment c i indicating the goodness or badness of assigning pollution process segment i to clustered segment c i . is the average distance between pollution process segment i and pollution process segments in other clustered segments c i except clustered segment c k , which is the average dissimilarity between different clustered segments.

[0188] Table 1 Explanation of asynchrony index

[0189]

[0190] Table 2 Explanation of severity index

[0191]

[0192] As Figure 11 shown in Table 1 and Table 2, in this embodiment, for a given pollution process pattern, the key feature points of the background segments and different site segments under this pattern are compared. This comparison allows quantifying the response differences of different sites to the same pollution event. For one peak or valley, the differences between them are described by the asynchrony index and the severity index:

[0193] The asynchrony index includes the advance and delay of the peak, the advance and delay of the start time of the rising segment, and the advance and delay of the end time of the falling segment. These indexes reflect the propagation time differences of pollutants between different sites and may be related to factors such as wind direction and terrain.

[0194] The severity indicators include the peak concentration difference and the total concentration-time product of the segment under the corresponding wavelet approximation curve. These indicators reflect the intensity differences of pollution events at different sites, which may be related to the distance of the site from the pollution source or local conditions.

[0195] The calculation of these indicators provides a basis for quantitatively comparing the pollution characteristics of different sites, which helps to more accurately describe and analyze the spatial characteristics of pollution events.

[0196] In this embodiment, the characteristic indicators of each site at different elevations, microenvironments, and relative positions are investigated. Combining with the mechanism knowledge of the relevant exogenous particulate matter pollution process, a characterization of the dynamic changes of near-surface particulate matter under this kind of atmospheric particulate matter pollution process is established. This step combines the quantitative analysis results with the actual geographical environment and pollution mechanism, aiming to construct a comprehensive pollution process model. By analyzing the response characteristics of sites under different environmental conditions, the influence of factors such as terrain and building distribution on pollutant diffusion can be better understood.

[0197] It can be understood that in this embodiment, the PM2.5 concentration time series is processed by discrete wavelet decomposition to obtain multi-time scale approximation sequences, so as to preliminarily distinguish the exogenous pollution process and local interference. The use of discrete wavelet decomposition allows the analysis of pollution data at different time scales, which helps to identify long-term trends and short-term fluctuations. Subsequently, the dynamic time warping method is used to eliminate the dynamic asynchrony between different sites, laying a foundation for comparing the response characteristics of different sites to the same exogenous pollution process. The application of dynamic time warping solves the problem of possible time delays in the data of different sites, making the comparison between sites more accurate. Then, a background time series that shows consistency between different sites is found as a characterization of the exogenous pollution process. This step helps to identify the common change patterns that may be caused by regional or large-scale pollution events. Finally, by combining the geographical locations and response characteristics of each site, the influence characteristics of exogenous particulate matter pollution on the near-surface pollution process are inferred. This comprehensive analysis method not only considers the time series data, but also incorporates spatial information and environmental factors, thus providing a more comprehensive understanding of the pollution process.

[0198] Example 4

[0199] Based on the analysis results of the urban near-surface air pollution data obtained in Example 3, this embodiment further analyzes the pollution process of atmospheric particulate matter on the urban near-surface.

[0200] In this embodiment, the changing patterns of the rising segments of the background particulate matter concentration obtained through time series clustering are compared with the actual exogenous air pollution mechanism, so as to identify the main exogenous particulate matter pollution patterns: the dominant type of northern exogenous air masses, the dominant type of southern urban pollution, the dominant type of southwestern exogenous air masses, the dominant type of sedimentation process, the dominant type of surrounding traffic pollution, and the static and stable diffusion type of surrounding traffic pollution. These patterns account for 78.0% of the total duration of the pollution process and 76.3% of the total PM2.5 concentration time integral. The three main frequency components of the temporal fluctuation of the particulate matter concentration are: sub-hourly fluctuation, hourly-level fluctuation, and sky-level fluctuation. These frequency components are closely related to factors such as meteorological conditions, input and diffusion of particulate matter during the pollution process.

[0201] In this embodiment, according to the geographical locations of different stations in the monitoring network, the stations are divided into the north-of-mountain group, the in-mountain group, the south-of-mountain group, the northeast interface group, the southwest interface group, the urban arterial road group, the lakeside arterial road group, the non-urban arterial road group, the mountaintop group, the mountainside group, the foot-of-mountain group, and the sea-level group. By analyzing the response differences of these groups during each pollution process, the spatio-temporal characteristics of the particulate matter pollution process can be inferred, and corresponding scenario patterns can be established.

[0202] This embodiment applies the extraction results of these exogenous pollution processes to the urban near-surface air pollution data analysis method to improve the ability to identify and analyze the air pollution process. By combining multi-time scale component separation of monitoring data, dynamic time warping processing, agglomerative hierarchical clustering processing, and identification of pollution process patterns, this method can more accurately characterize the pollution process of atmospheric particulate matter on the urban near-surface, providing strong technical support for air pollution control and environmental optimization design.

[0203] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms is not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0204] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0205] Any process or method description shown in the flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0206] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays, field-programmable gate arrays, etc.

[0207] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0208] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for analyzing urban near-surface air pollution data, characterized in that: include: Obtain time series data on atmospheric particulate matter concentration at multiple monitoring stations near the ground in the city; The atmospheric particulate matter concentration time series data is separated into multiple time scale components to obtain the first characteristic sequence of the atmospheric particulate matter concentration time series data at different time scales; Performing dynamic time warping on the first characteristic sequence between different monitoring sites to obtain a second characteristic sequence that eliminates the dynamic asynchrony between different monitoring sites; Performing agglomerative hierarchical clustering on the second characteristic sequence to obtain background atmospheric particulate matter concentration fluctuation characteristics; According to the fluctuation characteristics of the background atmospheric particulate matter concentration, the pollution process of atmospheric particulate matter on the urban near-surface is analyzed, including: According to the background atmospheric particulate matter concentration fluctuation characteristics, the atmospheric particulate matter concentration fluctuation process is divided into a plurality of pollution process segments; The pollution process fragments are subjected to time series clustering to obtain different pollution process patterns, including: Calculate the time series difference D(s′1||s′2) between any two pollution process segments, and its expression is as follows: D(s′1||s′2)=D DTW (s′1||s′2)×D shape (s′1||s′2)×D stretch (s′1||s′2) Among them, s′1 and s′2 are any two pollution process fragments on the same time scale after minimum value normalization, D DTW (s′1||s′2) and D shape (s′1||s′2) are the warping distance and shape distance after dynamic time warping, respectively. stretch (s′1||s′2) is the extension ratio of the original sequence on the time axis during dynamic time warping; According to the time series differences, a distance matrix between two pollution process segments is constructed; According to the distance matrix between the two pollution process fragments, the two pollution process fragments are subjected to agglomerative hierarchical clustering to obtain cluster fragments; Repeat the above steps until the cluster segments reach the optimal number of classifications, and each cluster segment represents a pollution process mode; Calculate the characteristic indicators of each monitoring station under the same pollution process mode, and analyze the pollution process of atmospheric particulate matter on the near-surface of the city based on the calculation results of the characteristic indicators.

2. The method for analyzing urban near-surface air pollution data according to claim 1, characterized in that: After obtaining the time series data of atmospheric particulate matter concentration at multiple monitoring stations near the ground in the city, the method further includes: The atmospheric particulate matter concentration time series data segments with more than M consecutive 0 values, null values ​​or unchanged values ​​are marked as abnormal segments and removed; M is a positive integer The outliers of the atmospheric particulate matter concentration time series data are obtained by using the absolute median difference between the original value of the atmospheric particulate matter concentration time series data and the differential integrated moving average autoregressive predicted value of the sequence, and the outliers are replaced by the differential integrated moving average autoregressive predicted value. The expression is as follows: Where c≈1.4826 is a constant, erfcinv() represents the inverse function of the complementary error number, spred(t) is the predicted value of s(t) under ARIMA(p,q), p is the order of the autoregressive term of the ARIMA model, q is the order of the difference term of the ARIMA model, and when res(t)>3, s(t) is considered to be an outlier and is replaced by spred(t).

3. The method for analyzing urban near-surface air pollution data according to claim 1, characterized in that: The atmospheric particulate matter concentration time series data is separated into multiple time scale components to obtain the first characteristic sequence of the atmospheric particulate matter concentration time series data at different time scales, including: The atmospheric particulate matter concentration time series data is decomposed by using a discrete wavelet decomposition method based on Meyer wavelet to obtain a wavelet approximate sequence of the atmospheric particulate matter concentration time series data on different time scales, and its expression is as follows: Where ψ(ω) is the Meyer wavelet basis function, j is the number of layers of the decomposed time scale, ω is the variable in the time domain representing the time point, and v(·) = a 4 (35-84a+70a 2 -20a 3 ), a∈[0,1] is the auxiliary function, a is the scale parameter in wavelet decomposition, z j is the low-frequency component of the atmospheric particle concentration time series data s, which represents the wavelet approximation sequence of the atmospheric particle concentration time series data s at time scale j, d n is the high-frequency component of the atmospheric particulate matter concentration time series data s at the decomposition level n; The wavelet approximation sequence is converted into a discrete first characteristic sequence, and its expression is as follows: Among them, fp(·) is the discrete first characteristic sequence, z (n) (·) is the nth-order derivative of the first characteristic sequence z(·), t represents time, n0 is the number of elements contained in the atmospheric particulate matter concentration time series data, and N is a positive integer greater than 0.

4. The method for analyzing urban near-surface air pollution data according to claim 1, characterized in that: The first characteristic sequence between different monitoring sites is subjected to dynamic time warping to obtain a second characteristic sequence that eliminates the dynamic asynchrony between different monitoring sites, including: Select the first characteristic sequences of two monitoring sites respectively; Construct a distance matrix for storing the difference between any elements of two first feature sequences. The expression is as follows: D(t1,t2)=z1(t1)-z2(t2),t1∈[1,m]∩N,t2∈[1,n]∩N Where D(t1,t2) is the distance matrix of t1×t2, z1(t1) is the element of the first characteristic sequence z1 of monitoring site 1 at time t1, z2(t2) is the element of the first characteristic sequence z2 of monitoring site 2 at time t2, m and n are the number of elements in the first characteristic sequence z1 and the first characteristic sequence z2 respectively, and N is a positive integer greater than 0; Use Floyd's algorithm to calculate the shortest path from the upper left corner to the lower right corner of the distance matrix; Determine the corresponding relationship between the elements in the two first characteristic sequences according to the shortest path; According to the corresponding relationship, the two first characteristic sequences are aligned including local stretching or compression to obtain a second characteristic sequence that eliminates the dynamic asynchronism between the two monitoring sites; Repeat the above steps to obtain the second characteristic sequence between any two monitoring sites at the same time scale.

5. The method for analyzing urban near-surface air pollution data according to claim 4, characterized in that: The second characteristic sequence is subjected to agglomerative hierarchical clustering to obtain the background atmospheric particulate matter concentration fluctuation characteristics, including: For the second characteristic sequence of any monitoring site, obtain the second characteristic sequence corresponding to any other monitoring site; Counting the time positions and atmospheric particle concentration values ​​in the second characteristic sequences of other monitoring stations, and eliminating the second characteristic sequences whose time positions and atmospheric particle concentration values ​​do not meet the preset conditions; Calculate the time position mean and the atmospheric particulate matter concentration mean of the remaining second characteristic sequence, and construct a characteristic sequence using the time position mean and the atmospheric particulate matter concentration mean, and then perform maximum value standardization on the constructed characteristic sequence to obtain a new second characteristic sequence for the current monitoring site; Repeat the above steps to construct a new second characteristic sequence for each monitoring site; Calculate the distance matrix between the new second feature sequences; According to the distance matrix between the new second characteristic sequences, the second characteristic sequences are clustered using agglomerative hierarchical clustering to obtain K types of clustering features; where K is the average length of the new second characteristic sequences of all monitoring stations; The intra-category similarity of all types of clustering features is evaluated, and the clustering features whose intra-category similarity meets the preset conditions are taken as the final background atmospheric particulate matter concentration fluctuation characteristics.

6. The method for analyzing urban near-surface air pollution data according to claim 1, characterized in that: The background atmospheric particulate matter concentration fluctuation characteristics are divided into several pollution process segments, including: According to the peak and valley values ​​of the atmospheric particle concentration fluctuation, the atmospheric particle concentration fluctuation is divided into rising and falling segments. Each individual rising or falling segment represents a process of increase or diffusion of the atmospheric particle concentration.

7. The method for analyzing urban near-surface air pollution data according to claim 1 or 6, characterized in that: After segmenting the background atmospheric particulate matter concentration fluctuation characteristics into a plurality of pollution process segments, the method further comprises performing minimum value normalization processing on the atmospheric particulate matter concentration in each pollution process segment.

8. The method for analyzing urban near-surface air pollution data according to claim 7, characterized in that: The method further includes calculating the maximum value of the silhouette coefficient SC of the agglomerative hierarchical clustering as the optimal number of classifications, and its expression is as follows: in, The profile value s(i) is the mean of the profile values ​​for agglomerative hierarchical clustering of all pollution process fragments. The profile value s(i) is defined as: in, is the pollution process segment i and cluster segment c i The average distance of other contaminated process fragments in the cluster indicates that the contaminated process fragment i is assigned to the cluster fragment c. i How good or bad is it? is the pollution process segment i and the cluster removal segment c i Other cluster fragments c k The average distance of the contaminated process fragments in , and the average dissimilarity between fragments in different clusters.

Citation Information

Patent Citations

  • Atmospheric pollutant monitoring method and system based on time domain data analysis

    CN118070110A

  • Airline planning method and device and storage medium

    CN116592894A