A VOCs multi-scale source tracing analysis method and system, and a storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-11
AI Technical Summary
然而,在实际大气环境中,排放强度和组分构成会随气象条件、污染程度以及时间变化而发生动态波动,导致解析结果往往掩盖了局部或瞬时的重要污染源信息,且人为确定因子数的过程存在较大的主观不确定性
[0007]This invention improves the representativeness of VOCs concentration data across different environmental backgrounds and rapidly changing processes by stratifying VOCs concentration data according to meteorological and pollution characteristics and employing a non-uniform resampling mechanism based on concentration gradients to construct a multi-scale sample set. Based on this, a source classification tree is constructed through hierarchical clustering, and a joint statistic integrating component spectrum similarity and time-series change distance is used to achieve deep identification of pollution sources from both spatial characteristics and temporal trends of chemical components. Statistical testing mechanisms are used to determine the confidence level of split nodes, objectively determining the optimal number of pollution sources and avoiding subjective bias and model computation uncertainty caused by manually specifying the number of sources. The obtained stable source results have high statistical confidence.
Smart Images

Figure CN122548181A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of source analysis, and in particular relates to a VOCs multi-scale source analysis method, system and storage medium. Background Technology
[0002] Volatile organic compounds (VOCs), as key precursors to the formation of tropospheric ozone (O3) and fine particulate matter (PM2.5), have a significant impact on the formation of regional atmospheric complex pollution. Accurately identifying VOC emission sources is a prerequisite for developing scientific and effective air quality control measures. Currently, receptor-based source apportionment techniques are widely used in atmospheric environmental research, among which the positive definite matrix factorization (PMF) model has become the mainstream method because it does not require prior knowledge of the source composition spectrum and the constraints conform to physical meaning. Traditional source apportionment procedures are usually based on the overall analysis of long-term observation data, assuming that the source composition spectrum remains unchanged throughout the observation period, and determining the optimal number of factors, i.e., the number of sources, through mathematical indicators such as Q-values and residual distributions combined with human experience. However, in the actual atmospheric environment, emission intensity and composition fluctuate dynamically with changes in meteorological conditions, pollution levels, and time, often causing the apportionment results to mask important local or instantaneous pollution source information, and the process of manually determining the number of factors involves significant subjective uncertainty.
[0003] Bootstrap resampling methods can assess the uncertainty of model results, but when performing cluster analysis on the massive solutions generated by resampling to identify stable sources, most existing techniques rely solely on the chemical similarity of source component spectra for matching. Especially when facing complex source classification hierarchies, there is a lack of an automated decision-making mechanism based on rigorous statistical tests to determine whether pollution sources should be further subdivided or merged into the same stable source category. This makes source apportionment results prone to over-splitting of factors or omission of characteristic sources when facing complex meteorological and variable pollution scenarios, failing to meet the urgent need for multi-scale, high-confidence source tracing of VOCs in current atmospheric environmental management. Summary of the Invention
[0004] To address the aforementioned problems, this invention first proposes a multi-scale source tracing analysis method for VOCs, comprising the following steps: Acquire VOCs concentration monitoring data and concurrent meteorological data for the area to be analyzed; The VOCs concentration monitoring data are stratified according to wind direction sector and pollution level. Within each stratum, resampling weights are assigned to samples at each time point based on the temporal gradient of VOCs species concentrations, and bootstrap resampling is performed using the weights to generate a bootstrap sample set with multi-scale characteristics. Source parsing is performed on each generated bootstrap sample to obtain multiple sets of source component spectra and source contribution time series; and hierarchical clustering is performed based on the similarity among all obtained source component spectra to construct a source classification tree; For any split node in the source classification tree, a joint similarity statistic is calculated to test the stability of the split based on the source component spectrum cosine similarity and the source contribution time series dynamic time regularization distance between the two sub-branches formed by the split. Based on the joint similarity statistics, a bootstrap test is performed to determine the confidence level of the split. When the confidence level of the split is lower than a preset threshold, the branch represented by the parent node of the split is determined to be a stable source, and the analysis results within the stable source branch are merged to output the source component spectrum and source contribution time series of the stable source.
[0005] Secondly, this invention proposes a VOCs multi-scale source tracing analysis system, comprising the following modules: The data acquisition module is used to acquire VOCs concentration monitoring data and concurrent meteorological data for the area to be analyzed. The sampling module is used to stratify the VOCs concentration monitoring data according to wind direction sector and pollution level; within each stratum, resampling weights are assigned to samples at each time point according to the time gradient of VOCs species concentration, and bootstrap resampling is performed using the weights to generate a bootstrap sample set with multi-scale features. The clustering module is used to perform source analysis on each generated bootstrap sample to obtain multiple sets of source component spectra and source contribution time series; and to perform hierarchical clustering based on the similarity among all the obtained source component spectra to construct a source classification tree; The similarity calculation module is used to calculate a joint similarity statistic for testing the stability of any split node in the source classification tree, based on the source component spectrum cosine similarity and the source contribution time series dynamic time regularization distance between the two sub-branches formed by the split. The source tracing module is used to perform a bootstrap test based on the joint similarity statistics to determine the confidence level of the split. When the confidence level of the split is lower than a preset threshold, the branch represented by the parent node of the split is determined to be a stable source, and the analysis results within the stable source branch are merged to output the source component spectrum and source contribution time series of the stable source.
[0006] Finally, the present invention proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0007] This invention improves the representativeness of VOCs concentration data across different environmental backgrounds and rapidly changing processes by stratifying VOCs concentration data according to meteorological and pollution characteristics and employing a non-uniform resampling mechanism based on concentration gradients to construct a multi-scale sample set. Based on this, a source classification tree is constructed through hierarchical clustering, and a joint statistic integrating component spectrum similarity and time-series change distance is used to achieve deep identification of pollution sources from both spatial characteristics and temporal trends of chemical components. Statistical testing mechanisms are used to determine the confidence level of split nodes, objectively determining the optimal number of pollution sources and avoiding subjective bias and model computation uncertainty caused by manually specifying the number of sources. The obtained stable source results have high statistical confidence. Attached Figure Description
[0008] Figure 1 This is a flowchart of a specific embodiment one; Figure 2 This is a schematic diagram of source component spectral hierarchical clustering. Detailed Implementation
[0009] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0010] Specific embodiment one is a VOCs multi-scale source tracing analysis method, the flowchart of which is as follows: Figure 1 This includes the following steps: S1, acquire VOCs concentration monitoring data and concurrent meteorological data for the area to be analyzed.
[0011] Online volatile organic compound (VOC) monitoring equipment is deployed at the target industrial park or urban monitoring points. This equipment uses gas chromatography-mass spectrometry (GC-MS) or gas chromatography-flame ionization (GC-FIN) detection, with a sampling time resolution of 1 hour. It continuously collects and analyzes ambient air samples throughout the year to obtain hourly concentration data for various VOC species, including alkanes, alkenes, aromatics, and halogenated hydrocarbons. Simultaneously, meteorological parameters such as wind speed, wind direction, temperature, relative humidity, and atmospheric pressure are recorded hourly using meteorological stations located at the same monitoring point or within a 500-meter radius of the monitoring point. The collected raw data is preprocessed to remove abnormal zero values and extremely high values caused by instrument malfunctions. Data below the instrument's minimum detection limit are replaced with half of the minimum detection limit value. Missing data are filled using linear interpolation with valid data from adjacent time points, resulting in time-series aligned concentration and meteorological datasets.
[0012] S2, the VOCs concentration monitoring data are stratified according to wind direction sector and pollution level; within each stratum, resampling weights are assigned to samples at each time point according to the time gradient of VOCs species concentration, and bootstrap resampling is performed using the weights to generate a bootstrap sample set with multi-scale characteristics.
[0013] The wind direction angle was divided into 16 sectors, each covering 22.5 degrees. The processed concentration monitoring data was assigned to the corresponding wind direction sector based on the wind direction at the corresponding time. Simultaneously, the sum of the concentrations of all volatile organic compounds (VOCs) at each time point was calculated, and the concentration data was divided into three levels: low pollution, medium pollution, and high pollution, using the quartile method. The original dataset was then segmented into multiple subsets combining wind direction sectors and pollution levels. In each subset, the absolute value of the rate of change of the concentration of each VOC at each time point relative to the concentration at the previous time point was calculated, and this absolute value was used as an indicator of the importance of the sample point. Normalization was applied to convert the rate of change indicators of all sample points into probability weights, giving sample points with drastic concentration changes a higher probability of being selected. The number of resampling operations was set to 100. Using the aforementioned probability weights, random sampling with replacement was performed on the original subset data to construct 100 new data matrices with the same sample size as the original subset. These matrices constitute the bootstrap sample set with multi-scale characteristics.
[0014] S3. Perform source parsing on each generated bootstrap sample to obtain multiple sets of source component spectra and source contribution time series; and perform hierarchical clustering based on the similarity among all obtained source component spectra to construct a source classification tree.
[0015] A positive definite matrix factorization model is initialized, with the number of factors set to a search range of 3 to 10, and the number of iterations set to 20. The generated 100 bootstrap sample sets are sequentially input into the positive definite matrix factorization model for computation. Under non-negative constraints, the objective function is iteratively solved using the least squares method until convergence, resulting in 100 sets of analytical results. Each set of analytical results contains a source component spectrum matrix and a source contribution time series matrix. Source component spectrum vectors are extracted from all analytical results, and the cosine similarity between any two source component spectrum vectors is calculated to construct a similarity matrix. A cohesive hierarchical clustering algorithm is used, with source component spectrum vectors as clustering objects and 1 minus the cosine similarity as the distance metric. The most similar source component spectra are gradually merged according to the principle of minimum distance priority until all source component spectra are merged into a single root node, thus establishing a tree-like hierarchical structure with this root node at the top and the individual source component spectra obtained from each of the original analyses at the bottom. Figure 2 As shown.
[0016] S4. For any split node in the source classification tree, calculate the joint similarity statistic to test the stability of the split based on the source component spectrum cosine similarity and the source contribution time series dynamic time regularization distance between the two sub-branches formed by the split.
[0017] During the traversal of the source classification tree, for each parent node with a branching structure, its directly connected left and right child nodes are identified. The mean vectors of all source component spectra and the mean vector of the source contribution time series contained in the left child node are extracted, and the corresponding mean vectors of the right child nodes are also extracted. The cosine similarity between the mean vectors of the source component spectra of the left and right child nodes is calculated using the vector dot product formula. At the same time, the distance between the mean vectors of the source contribution time series of the left and right child nodes is calculated using the dynamic time warping algorithm. By constructing a distance accumulation matrix and finding the minimum path cost, a dynamic time warping distance that can tolerate nonlinear distortion of the time axis is obtained, and this distance is mapped to the interval of 0 to 1 using the max-min normalization method. The weight coefficient is set to 0.6, and the source component spectrum cosine similarity is multiplied by this weight coefficient, plus 1 minus the normalized dynamic time warping distance multiplied by the remaining weight, to obtain the fused joint similarity statistic.
[0018] S5. Perform a bootstrap test based on the joint similarity statistics to determine the confidence level of the split. When the confidence level of the split is lower than a preset threshold, the branch represented by the parent node of the split is determined to be a stable source. The analysis results within the stable source branch are merged, and the source component spectrum and source contribution time series of the stable source are output.
[0019] A null distribution for hypothesis testing is constructed, which consists of the joint similarity statistics calculated from the datasets after random shuffling of time series order under the same clustering structure. The actual joint similarity statistics calculated in S4 are located in the null distribution, and the corresponding percentile value is calculated as the confidence score of the split node. A preset threshold of 0.85 is set. If the confidence score of a split node is less than 0.85, the split is considered to be caused by random noise or algorithm instability and is not statistically significant, so further splitting of the node is stopped. The node is identified as a stable source node, and the results of the original positive definite matrix factorization model corresponding to all leaf nodes under the node are collected. The arithmetic mean of all collected source component spectrum matrices is used to obtain the final stable source component spectrum, and the arithmetic mean of all collected source contribution time series matrices is used to obtain the stable source contribution time series. The two average matrices are output as the final analytical result of the stable source to the storage medium or display terminal.
[0020] In a preferred embodiment, acquiring VOCs concentration monitoring data and concurrent meteorological data for the area to be analyzed includes: Quality control is performed on the collected raw VOCs monitoring data, and linear interpolation is used to complete data segments with missing durations shorter than the preset duration. Construct the uncertainty matrix corresponding to the concentration matrix, following these rules: For a specific species at a specific time point, if the concentration value is less than or equal to the method detection limit for that species, the concentration value is replaced with half of the method detection limit, and the corresponding uncertainty is set to five-sixths of the method detection limit. If the concentration value is greater than the method detection limit, the uncertainty is calculated based on the method detection limit and the error fraction. The calculation method is as follows: calculate the square of the method detection limit and the square of the product of the error fraction and the concentration value, sum the two squares and take the square root. The processed concentration matrix and the corresponding uncertainty matrix are aligned with the meteorological data by time to obtain a standardized input dataset.
[0021] First, data cleaning and preprocessing steps are performed. For temporary data loss due to calibration or malfunction of monitoring equipment, the time series is automatically scanned to locate gaps of less than 3 hours. For example, if a monitoring point loses data between 10:00 AM and 12:00 PM, the program will use concentration data from 9:00 AM and 1:00 PM to calculate the slope and intercept using a linear equation, thus estimating the filler value for the missing period to ensure the continuity of the time series. Then, each element in the concentration matrix is traversed, and the method detection limit for the j-th species is set. The concentration is 0.5 micrograms per cubic meter, for the i-th time point. If the detected value is 0.2 micrograms per cubic meter, which is less than the detection limit, then the concentration value should be replaced with 0.25 micrograms per cubic meter, and the corresponding uncertainty should be set. The concentration was 0.416 micrograms per cubic meter. If the concentration is detected The concentration was 10.0 micrograms per cubic meter, which is greater than the detection limit, and an error fraction was set. If it is 0.1, then according to the formula The uncertainty of the data point is calculated to determine the measurement error range of high concentration values. The VOCs concentration matrix and its corresponding uncertainty matrix, after being cleaned and completed, are matched and aligned with meteorological parameters such as wind speed, wind direction, temperature and humidity of the same period according to a unified timestamp. The time rows that cannot be matched are removed to obtain a standardized input dataset with the number of rows representing time points and the number of columns representing species concentration and meteorological parameters.
[0022] In a preferred embodiment, assigning resampling weights to samples at each time point based on the temporal gradient of VOCs species concentrations includes: Calculate the sum of the absolute values of the first-order differences of the concentrations of all VOC species j at time point t relative to the previous time point, and use it as the temporal gradient at that time point. Dividing the time gradient at that point in time by the sum of the time gradients at all points in that stratum yields the time point within that stratum. The resampling weights are set to equal values if the sum of the time-varying gradients of all time points within the layer is zero.
[0023] To obtain the dynamic characteristics of the pollution process, the degree of concentration change at each time point relative to the previous time point is calculated, and the time gradient of time point t is defined. This is the sum of the concentration changes of all J VOC species at that time. For example, if at time t, the concentration of ethylene increased by 5 units compared to time t-1, toluene decreased by 3 units, and the concentrations of the other species remained unchanged, then the concentration at that time is... The value is the sum of the absolute values of the differences between the terms, i.e., according to the formula... Calculations show that a larger value indicates a period of drastic change in pollution emissions at that moment; After obtaining the gradient values at all time points, the non-uniform resampling weights at time point t within that layer will be calculated. Through formula The gradient values are normalized into probability values, so that time points with drastic concentration changes have a greater probability of being selected in subsequent sampling, thereby enhancing the model's ability to analyze pollution peaks and emission abrupt events. If an extreme case occurs in the calculation process where the gradient sum of all time points within the stratum is zero, that is, the data is completely stable without fluctuations, the system automatically sets the weight of all time points to an equal value of 1 / T to ensure the robustness of the sampling process.
[0024] In a preferred embodiment, the step of using the weights to perform bootstrap resampling to generate a bootstrap sample set with multi-scale features includes: For each stratum, the probability distribution calculated based on that stratum. T samples are drawn with replacement from the dataset of this stratification to construct the subset of data for the b-th resampling. ; Repeat the above process B times to obtain a sample set containing B subsets of the dataset. .
[0025] The number of resampling iterations, B, is set to 100 to ensure sufficient representativeness of the statistical results. For each defined time stratum, the probability distribution vector corresponding to that stratum is retrieved, which contains the probability of being sampled at each time point calculated based on the concentration gradient. Using the Monte Carlo simulation method, the computer performs random sampling with replacement based on probability P, sampling a complete species concentration row vector at each time point, repeating this sampling T times, thereby constructing a new resampled subset with the same number of rows as the original dataset. In this process, weight Samples from periods of high pollution variation may be in The sample may appear repeatedly during the plateau phase, while samples from the plateau phase may be overlooked. The above resampling process is repeated 100 times, with each iteration generating a subset of the dataset. Each of these 100 independent subsets retains different statistical characteristics of the original data. Finally, these 100 subsets are aggregated to obtain a large sample set containing B subsets. This sample set amplifies the data fluctuation characteristics caused by changes in the intensity of pollution source emissions through a non-uniform weighting mechanism, providing information input for the PMF model to identify weak or instantaneous emission sources.
[0026] In a preferred embodiment, the step of performing source analysis on each generated bootstrap sample to obtain multiple sets of source component spectra and source contribution time series includes: A positive definite matrix factorization (PMF) model is used for each subset of the dataset. Given a preset number of factors p, the source component spectrum matrix F and the source contribution matrix G are solved by iteratively minimizing the objective function Q.
[0027] The system employs a positive definite matrix factorization (PMF) model as its core analytical algorithm, and reads each subset of the sample set sequentially. The system initializes the preset number of factors, p, for example, setting it to resolve 6 pollution sources. Then, iterative optimization is initiated to minimize the objective function Q, which measures the residual between the model-reconstructed concentration and the actual observed concentration, and is weighted and standardized using uncertainty. The specific calculation formula is as follows: ; In the solution process, the concentration matrix is decomposed into the source contribution matrix G and the source component spectrum matrix F, where This represents the contribution concentration of the k-th pollution source at time point i. This represents the mass percentage of the j-th chemical species in the k-th pollution source. Simultaneously, the model strictly enforces non-negativity constraints, meaning it requires all... and To ensure that the analytical results conform to the non-negativity in the physical and chemical sense, after multiple iterations and convergence, 100 sets of corresponding F and G matrices will be output for 100 subsets of data as the statistical distribution range of the model results.
[0028] In a preferred embodiment, the step of constructing a source classification tree by performing hierarchical clustering based on the similarity among all obtained source component spectra includes: Collect all source component spectral vectors generated by the Bootstrap run; Calculate the source component spectral vector and Cosine distance between Construct the distance matrix; Based on the distance matrix, hierarchical clustering is performed using the Ward minimum variance method. In each merging step, the increment of the sum of squares within the cluster is minimized to construct a binary tree structure containing all Bootstrap source factors.
[0029] First, create a vector pool to collect all source component spectral vectors generated from 100 Bootstrap runs. If each run uses 6 factors, a total of 600 source component spectral vectors will be collected. Then, calculate the spectral vectors of any two source component vectors. and The cosine distance between them is used to calculate the difference in their chemical characteristics using a formula. The closer the cosine distance is to 0, the more similar the chemical fingerprints of the two are. In this way, a 600x600 distance matrix is constructed. Based on the calculated distance matrix, a bottom-up hierarchical clustering algorithm is executed using Ward's minimum variance method. In each step of clustering, all possible cluster merging schemes are traversed, and the pair of clusters that minimize the increase in the sum of squares within the merged cluster is selected for merging. For example, two industrial source factors that both show high proportions of toluene and xylene are preferentially classified into one class. Through continuous iterative merging, a binary tree structure containing all Bootstrap source factors is established. The root node of the tree represents the set of all factors, and the leaf nodes represent individual specific source factors. The splitting process of the tree represents the hierarchical relationship between different pollution source categories.
[0030] In a preferred embodiment, the calculation of the joint similarity statistic used to test the stability of the split includes: Let the two sub-branches generated by the split node be A and B, respectively. Calculate the average source component spectral vector of all source factors contained in sub-branches A and B. , and the average source contribution time series vector , ; Calculate the cosine similarity between the average component spectra, and then perform Z-score normalization on the average source contribution time series to obtain the results. and Calculate the dynamic time-warped distance and convert it into time series similarity. The arithmetic mean of the cosine similarity and the time series similarity is used as the joint similarity statistic of the distance between the centers of the two sub-branches.
[0031] For any split node in the source classification tree, assuming it divides the dataset of its parent node into two sub-branches A and B, calculate the arithmetic mean of the component spectral vectors of all source factors in sub-branch A. and the average value of the corresponding source contribution time series Similarly, calculate sub-branch B. and , using these as the central representatives of their respective branches; calculate the cosine similarity between the two average component spectra; To assess the consistency of the temporal variation trend, the average source contribution time series is first Z-score standardized to eliminate the influence of dimensions, resulting in the series... and Then, the dynamic time warping algorithm is applied to calculate the distance DTW between the two. , ), and according to the formula This is converted into time series similarity, an indicator that can tolerate small shifts on the timeline; through the formula... Calculate the joint similarity statistic.
[0032] In a preferred embodiment, performing a bootstrap test based on the joint similarity statistic to determine the confidence level of the split includes: For a specific split node in the source classification tree, extract the set of source factors mapped to sub-branch A. and the set of source factors mapped to subbranch B ; Perform a paired bootstrap test: randomize from Extract a factor from and from Extract a factor from Calculate the joint similarity between the two, and repeat the sampling N times, where N≥B; In N sampling iterations, the joint similarity is less than a preset discrimination threshold. The number of times; The confidence level of the split is the ratio of the number of times the joint similarity is less than the preset discrimination threshold to the total number of extractions. If the confidence level of the split is less than the preset threshold, the two sub-branches generated by the split are determined to be statistically insignificant, the split of the node is stopped and it is reverted to merge into a single stable source.
[0033] To verify whether a specific split node in the source classification tree is statistically significant, the set of source factors mapped to the left sub-branch A of that node is extracted. The source factor set of the right sub-branch B And execute the paired bootstrap test procedure, setting the number of samplings N to 1000, each time randomly selecting from... Extract a factor from At the same time from Extract a factor from Calculate the joint similarity between this pair of factors. ; Among 1000 samplings, those satisfying the joint similarity... Less than the preset discrimination threshold For example, the number of times 0.8 According to the formula Calculate the confidence level of the split. If the calculated confidence level Conf is less than a preset threshold... For example, if the result is 80%, it indicates that the two subclasses generated by the split are not statistically distinguishable enough. The split will be judged as unstable, and further subdivision at that node will be stopped. Instead, it will be backed up and merged into a single stable source category, thereby avoiding the model from producing false source resolution results.
[0034] To facilitate understanding of the present invention, the solution is described below with reference to specific data. Assume that VOCs data were collected in an industrial park. At 10:00 AM on October 1st, the concentration of benzene was detected to be 3.0%. However, toluene data was missing, so linear interpolation was used to complete the toluene data, resulting in a value of 1.5. Given that the method detection limit (MDL) for benzene is 0.5 and the error fraction is set to 0.1, the intermediate variable for the uncertainty of benzene is... The processed concentration data was aligned with the concurrent northwesterly wind meteorological data to construct a standardized input matrix.
[0035] After stratifying the data, the resampling weights are calculated. Assuming that within the northwest wind stratum, the total VOCs concentration at time t increases sharply compared to time t-1, the sum of the absolute values of the first-order differences for all species at that time is calculated to be 200. If the sum of gradients at all time points within this stratum is 10000, then the intermediate variable for the resampling weights at a specific time t is... Utilizing the weight distribution, perform 100 Bootstrap resampling cycles with replacement to generate a sample set containing 100 subsets. This ensures that pollution peaks with high variability are prioritized and preserved in the sample.
[0036] The PMF model was run on 100 Bootstrap subsets, with a factor count of 5, generating 500 source component spectral vectors. The cosine distance between the vectors was calculated. For example, "Source A" from the first run is very similar to "Source B" from the 50th run, with a cosine angle of 0.98. Therefore, the intermediate distance between them is d = 1 - 0.98 = 0.02. Based on this distance matrix, Ward's minimum variance method was used for clustering, progressively merging the 500 results to obtain a multi-level binary tree structure.
[0037] When traversing the classification tree, it is necessary to examine whether sub-branch A (suspected motor vehicle source) and sub-branch B (suspected combustion source) splitting from a certain node are truly different. The average component spectrum cosine similarity between the two is calculated to be 0.6; simultaneously, the source contribution time series of both are Z-score normalized, the DTW distance is calculated and converted to a time series similarity of 0.4. The joint similarity statistic used for the test is calculated as (0.6+0.4) / 2=0.5. The value is low, indicating that the two branches have significant differences in chemical composition and emission time, and may indeed be two independent sources.
[0038] Furthermore, a confidence test is performed on the split node. Pairs are randomly selected 1000 times from the factor sets of sub-branches A and B, and a discrimination threshold is set. Statistical analysis revealed that in 1000 extractions, the calculated joint similarity was less than 0.55 in 950 cases, indicating significant differences. Therefore, the confidence level for this split is 950 / 1000 = 0.95. Since 0.95 is higher than a preset threshold (e.g., 0.9), the split is deemed valid, and A and B are retained as two independent stable sources. Conversely, if the confidence level is extremely low, they are merged, and the statistically validated stable source component spectra and contribution time series are output.
[0039] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A multi-scale source tracing analysis method for VOCs, characterized in that, Includes the following steps: Acquire VOCs concentration monitoring data and concurrent meteorological data for the area to be analyzed; The VOCs concentration monitoring data were stratified according to wind direction, sector, and pollution level. Within each stratum, resampling weights are assigned to samples at each time point based on the temporal gradient of VOC species concentrations, and bootstrap resampling is performed using these weights to generate a bootstrap sample set with multi-scale features. Source parsing is performed on each generated bootstrap sample to obtain multiple sets of source component spectra and source contribution time series; and hierarchical clustering is performed based on the similarity among all obtained source component spectra to construct a source classification tree; For any split node in the source classification tree, a joint similarity statistic is calculated to test the stability of the split based on the source component spectrum cosine similarity and the source contribution time series dynamic time regularization distance between the two sub-branches formed by the split. Based on the joint similarity statistics, a bootstrap test is performed to determine the confidence level of the split. When the confidence level of the split is lower than a preset threshold, the branch represented by the parent node of the split is determined to be a stable source, and the analysis results within the stable source branch are merged to output the source component spectrum and source contribution time series of the stable source.
2. The method according to claim 1, characterized in that, The acquisition of VOCs concentration monitoring data and concurrent meteorological data for the area to be analyzed includes: Quality control is performed on the collected raw VOCs monitoring data, and linear interpolation is used to complete data segments with missing durations shorter than the preset duration. Construct the uncertainty matrix corresponding to the concentration matrix, following these rules: For a specific species concentration at a specific time point, if the concentration value is less than or equal to the method detection limit for that species, the concentration value is replaced with half of the method detection limit, and the corresponding uncertainty is set to five-sixths of the method detection limit. If the concentration value is greater than the method detection limit, the uncertainty is calculated based on the method detection limit and the error fraction. The calculation method is as follows: calculate the square of the method detection limit and the square of the product of the error fraction and the concentration value, sum the two squares, and then take the square root. The processed concentration matrix and the corresponding uncertainty matrix are aligned with the meteorological data by time to obtain a standardized input dataset.
3. The method according to claim 1, characterized in that, The method of assigning resampling weights to samples at each time point based on the temporal gradient of VOC species concentrations includes: Calculate the sum of the absolute values of the first-order differences of the concentrations of all VOC species j at time point t relative to the previous time point, and use it as the temporal gradient at that time point. Divide the time gradient of the time point by the sum of the time gradients of all time points within the stratum to obtain the resampling weight of time point t within the stratum; if the sum of the time gradients of all time points within the stratum is zero, then set the resampling weight of all time points to an equal value.
4. The method according to claim 1, characterized in that, The step of using the weights to perform bootstrap resampling to generate a bootstrap sample set with multi-scale features includes: For each stratum, the probability distribution calculated based on that stratum. T samples are drawn with replacement from the dataset of this stratification to construct the subset of data for the b-th resampling. ; Repeat the above procedure B times to obtain a sample set comprising B sub-data sets .
5. The method of claim 1, wherein, The process involves performing source analysis on each generated bootstrap sample to obtain multiple sets of source component spectra and source contribution time series, including: Using the positive matrix factorization (PMF) model, for each sub-data set The source composition spectral matrix F and the source contribution matrix G are solved by iteratively minimizing the objective function Q at a pre-set number of factors p.
6. The method of claim 1, wherein, The step of constructing a source classification tree by performing hierarchical clustering based on the similarity among all obtained source component spectra includes: Collect all source component spectral vectors generated by the Bootstrap run; Calculate the source component spectral vector and Cosine distance between Construct the distance matrix; Based on the distance matrix, hierarchical clustering is performed using the Ward minimum variance method. In each merging step, the increment of the sum of squares within the cluster is minimized to construct a binary tree structure containing all Bootstrap source factors.
7. The method according to claim 1, characterized in that, The calculation of the joint similarity statistic used to test the stability of the split includes: Let the two sub-branches generated by the split node be A and B, respectively. Calculate the average source component spectral vector of all source factors contained in sub-branches A and B. , and the average source contribution time series vector , ; The cosine similarity between the average ingredient profiles is calculated, and the average source contribution time series is Z-score standardized to obtain and The dynamic time warping distance is calculated, and the dynamic time warping distance is converted into a time series similarity. The arithmetic mean of the cosine similarity and the time series similarity is used as the joint similarity statistic of the distance between the centers of the two sub-branches.
8. The method of claim 1, wherein, The step of performing a bootstrap test based on the joint similarity statistic to determine the confidence level of the split includes: For a particular split node in the source taxonomy tree, extract the set of source factors that map to child branch A and the set of source factors that map to child branch B ; Performing the paired Bootstrap test: randomly draw one factor from and one factor from and compute the joint similarity between the two, repeat drawing N times, N ≥ B and one factor from and compute the joint similarity between the two, repeat drawing N times, N ≥ B the number of times that the joint similarity is less than the preset discrimination threshold in N times of extraction of times The confidence level of the split is the ratio of the number of times the joint similarity is less than the preset discrimination threshold to the total number of extractions. If the confidence level of the split is less than the preset threshold, the two sub-branches generated by the split are determined to be statistically insignificant, the split of the node is stopped and it is reverted to merge into a single stable source.
9. A VOCs multi-scale source analysis system, characterized in that, Includes the following modules: The data acquisition module is used to acquire VOCs concentration monitoring data and concurrent meteorological data for the area to be analyzed. The sampling module is used to stratify the VOCs concentration monitoring data according to wind direction, sector, and pollution level. Within each stratum, resampling weights are assigned to samples at each time point based on the temporal gradient of VOC species concentrations, and bootstrap resampling is performed using these weights to generate a bootstrap sample set with multi-scale features. The clustering module is used to perform source analysis on each generated bootstrap sample to obtain multiple sets of source component spectra and source contribution time series. Based on the similarity among all the obtained source component spectra, hierarchical clustering is performed to construct a source classification tree; The similarity calculation module is used to calculate a joint similarity statistic for testing the stability of any split node in the source classification tree, based on the source component spectrum cosine similarity and the source contribution time series dynamic time regularization distance between the two sub-branches formed by the split. The source tracing module is used to perform a bootstrap test based on the joint similarity statistics to determine the confidence level of the split. When the confidence level of the split is lower than a preset threshold, the branch represented by the parent node of the split is determined to be a stable source, and the analysis results within the stable source branch are merged to output the source component spectrum and source contribution time series of the stable source.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, implements the method as described in any one of claims 1-8.