Incremental clustering method and system of atmospheric aerosol particle mass spectrum data set, terminal and storage medium
The cluster analysis of the atmospheric aerosol particle material spectrum data set was solved through the FASC algorithm, which improved the efficiency and accuracy of cluster analysis, and generated cluster sets that satisfies the constraints of similarity between clusters and similarity within clusters.
Patent Information
- Application Number
- CN202510839139.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the prior art, cluster analysis of atmospheric aerosol particle material spectrum data sets has problems of intercluster confusion, resulting in low accuracy and low efficiency of cluster analysis results.
A flexible incremental clustering algorithm (FASC) combining similarity and density is adopted to generate cluster sets that meet the similarity between clusters and similarity constraints between clusters, reduce inter-cluster confusion and improve cluster analysis efficiency and accuracy through target cluster selection strategy, cluster allocation processing, merging processing and denoising processing.
It significantly reduces intercluster confusion, improves the cluster analysis efficiency and accuracy of atmospheric aerosol particle matter spectrum data set under high-dimensional big data conditions, and generates intuitive and quantitative clustering results.
Smart Images

Figure CN120354154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to an incremental clustering method, system, terminal and computer-readable storage medium for an atmospheric aerosol particle mass spectrometry data set. Background Art
[0002] Clustering is an unsupervised learning technique that can automatically identify the internal structure of data through metrics such as distance or density, and divide data points into clusters without prior labels, minimizing the differences within clusters and maximizing the differences between clusters. Clustering analysis technology can reveal the internal structure patterns of data sets and has a wide range of applications in fields such as computer vision (such as image segmentation), bioinformatics (such as gene expression analysis), and complex network analysis (such as community detection).
[0003] Performing clustering analysis on a large data set of atmospheric aerosol particle mass spectrometry is a currently focused direction in atmospheric research and protection. However, in traditional techniques, there are problems of confusion between clusters in the clustering analysis method for the atmospheric aerosol particle mass spectrometry data set, resulting in low accuracy and low analysis efficiency of the clustering analysis results for the atmospheric aerosol particle mass spectrometry data set.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide an incremental clustering method, system, terminal and computer-readable storage medium for an atmospheric aerosol particle mass spectrometry data set, aiming to solve the problem that in the existing technology, there is a situation of confusion between clusters in the clustering analysis method for the atmospheric aerosol particle mass spectrometry data set, resulting in low accuracy and low analysis efficiency of the clustering analysis results for the atmospheric aerosol particle mass spectrometry data set.
[0006] To achieve the above purpose, the present invention provides an incremental clustering method for an atmospheric aerosol particle mass spectrometry data set, and the incremental clustering method for the atmospheric aerosol particle mass spectrometry data set includes the following steps: Obtain an atmospheric aerosol particle mass spectrometry data set, and perform clustering initialization processing on the atmospheric aerosol particle mass spectrometry data set to obtain a plurality of clustering initialization parameters; Determine a target cluster selection strategy, and use the target cluster selection strategy to perform a first similarity calculation and cluster assignment processing on the plurality of clustering initialization parameters to obtain a plurality of cluster center matrices; Perform merging processing and denoising processing on the plurality of cluster center matrices to obtain an initial clustering result, and perform a second similarity calculation on the initial clustering result to obtain an iterative similarity; Perform clustering convergence determination and iterative loop clustering processing on the initial clustering result according to the iterative similarity, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset.
[0007] Optionally, in the incremental clustering method for the atmospheric aerosol particle mass spectrometry dataset, the clustering initialization parameters include target clusters, the cluster center of each target cluster, the number of data points in each target cluster, and the cluster number. The method for obtaining the atmospheric aerosol particle mass spectrometry dataset and performing clustering initialization processing on the atmospheric aerosol particle mass spectrometry dataset to obtain multiple clustering initialization parameters specifically includes: Obtain the atmospheric aerosol particle mass spectrometry dataset and convert the atmospheric aerosol particle mass spectrometry dataset into a target data matrix. Obtain the vector of each row in the target data matrix, and perform normalization processing based on the L2 norm on the vector to obtain a normalized data vector. Determine multiple random sample vectors in the normalized data vector, where the random sample vector is a vector randomly extracted from the normalized data vector, and set each random sample vector as the cluster center to obtain multiple cluster centers and multiple target clusters. Perform initialization processing on each target cluster to obtain the number of data points in each target cluster and the cluster number.
[0008] Optionally, in the incremental clustering method for the atmospheric aerosol particle mass spectrometry dataset, the target cluster selection strategy includes a similarity priority strategy and a density priority strategy. Determine the target cluster selection strategy, and perform the first similarity calculation and cluster assignment processing on multiple clustering initialization parameters by using the target cluster selection strategy to obtain multiple cluster center matrices, specifically including: Use any random number generation algorithm to construct an integer array according to the number of data points in each target cluster and the cluster number, and sequentially extract each data point in the integer array. If the target cluster selection strategy is the similarity priority strategy, calculate the similarity for each data point, and perform cluster assignment processing according to the similarity priority strategy to obtain multiple cluster center matrices. If the target cluster selection strategy is the density priority strategy, calculate the similarity for each data point, and perform cluster assignment processing according to the density priority strategy to obtain multiple cluster center matrices.
[0009] Optionally, in the incremental clustering method for the atmospheric aerosol particle mass spectrometry data set, the step of calculating the similarity for each data point and performing cluster assignment processing according to the similarity priority strategy to obtain multiple cluster center matrices specifically includes: Calculate the similarity between each data point and multiple target clusters, and determine the first assigned target cluster corresponding to each data point when the similarity is the largest, where the first assigned target cluster is the target cluster with the largest similarity to the data point among the multiple target clusters; Add each data point to the corresponding first assigned target cluster to obtain multiple cluster center matrices.
[0010] Optionally, in the incremental clustering method for the atmospheric aerosol particle mass spectrometry data set, the step of calculating the similarity for each data point and performing cluster assignment processing according to the density priority strategy to obtain multiple cluster center matrices specifically includes: Calculate the similarity between each data point and multiple target clusters; Determine the preset intra-cluster similarity, and extract multiple second assigned target clusters whose similarity is greater than or equal to the preset intra-cluster similarity; Calculate the cluster density of each second assigned target cluster, and obtain the third assigned target cluster with the largest cluster density among the second assigned target clusters, where the third assigned target cluster is the target cluster with the largest cluster density among the multiple second assigned target clusters; Add each data point to the corresponding third assigned target cluster to obtain multiple cluster center matrices.
[0011] Optionally, in the incremental clustering method for the atmospheric aerosol particle mass spectrometry data set, the step of performing merging processing and denoising processing on the multiple cluster center matrices to obtain an initial clustering result, and performing a second similarity calculation on the initial clustering result to obtain an iterative similarity specifically includes: Calculate the inter-cluster similarity between multiple cluster center matrices, and extract the cluster center matrices whose inter-cluster similarity is greater than the preset inter-cluster similarity to obtain a cluster index set; Perform merging processing on the cluster index set to obtain a redundancy-removed cluster set; Delete the clusters with 0 data vectors in the redundancy-removed cluster set, and perform cluster sorting processing and cluster numbering processing to obtain an initial clustering result; Perform iterative inter-similarity calculation on the initial clustering result to obtain an iterative similarity, where the iterative inter-similarity calculation includes iterative inter-similarity calculation based on cluster distribution and iterative inter-similarity calculation based on cluster numbering.
[0012] Optionally, for the incremental clustering method of the atmospheric aerosol particle mass spectrometry data set, the method of performing clustering convergence determination and iterative loop clustering processing on the initial clustering result according to the iterative similarity, and outputting the clustering result of the atmospheric aerosol particle mass spectrometry data set specifically includes: Determine a preset convergence value, and determine whether the iterative similarity is less than the preset convergence value; If so, perform iterative loop clustering processing on the initial clustering result, where the input of each iteration is the clustering result output by the previous iteration, until the iterative similarity is greater than or equal to the preset convergence value or the number of iterations is greater than the preset iteration number threshold, and output the clustering result of the atmospheric aerosol particle mass spectrometry data set.
[0013] In addition, to achieve the above object, the present invention also provides an incremental clustering system for an atmospheric aerosol particle mass spectrometry data set, where the incremental clustering system for the atmospheric aerosol particle mass spectrometry data set includes: A clustering initialization module, configured to obtain an atmospheric aerosol particle mass spectrometry data set, and perform clustering initialization processing on the atmospheric aerosol particle mass spectrometry data set to obtain a plurality of clustering initialization parameters; A cluster selection strategy determination module, configured to determine a target cluster selection strategy, and perform a first similarity calculation and cluster assignment processing on the plurality of clustering initialization parameters by using the target cluster selection strategy to obtain a plurality of cluster center matrices; An inter-iteration similarity calculation module, configured to perform merging processing and denoising processing on the plurality of cluster center matrices to obtain an initial clustering result, and perform a second similarity calculation on the initial clustering result to obtain an iterative similarity; A clustering result output module, configured to perform clustering convergence determination and iterative loop clustering processing on the initial clustering result according to the iterative similarity, and output the clustering result of the atmospheric aerosol particle mass spectrometry data set.
[0014] In the present invention, an atmospheric aerosol particle mass spectrum data set is obtained, and the atmospheric aerosol particle mass spectrum data set is subjected to clustering initialization processing to obtain a plurality of clustering initialization parameters; a target cluster selection strategy is determined, and the target cluster selection strategy is used to perform a first similarity calculation and cluster assignment processing on the plurality of clustering initialization parameters to obtain a plurality of cluster center matrices; the plurality of cluster center matrices are subjected to merging processing and denoising processing to obtain an initial clustering result, and a second similarity calculation is performed on the initial clustering result to obtain an iterative similarity; according to the iterative similarity, clustering convergence determination and iterative loop clustering processing are performed on the initial clustering result, and the clustering result of the atmospheric aerosol particle mass spectrum data set is output. By adopting the target cluster selection strategy to perform similarity calculation and cluster assignment processing on the clustering initialization parameters corresponding to the atmospheric aerosol particle mass spectrum data set, the cluster to which each data point belongs can be flexibly dynamically selected by two strategies, including the traditional similarity priority strategy and the specially designed density priority strategy, and a cluster set that meets the cluster - to - cluster similarity and intra - cluster similarity constraint conditions can be generated. Further, by performing merging processing and denoising processing on the cluster center matrix and performing iterative loop clustering processing according to the inter - iterative similarity, the inter - cluster confusion can be significantly reduced, and the clustering analysis efficiency of the atmospheric aerosol particle mass spectrum data set and the accuracy of the clustering result output under the condition of high - dimensional big data can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flowchart of a preferred embodiment of the incremental clustering method for the atmospheric aerosol particle mass spectrum data set of the present invention; Figure 2 is a schematic diagram of the inter - cluster relationship between the similarity priority strategy and the density priority strategy of a preferred embodiment of the incremental clustering method for the atmospheric aerosol particle mass spectrum data set of the present invention; Figure 3 is a schematic diagram of the proportion of data contained in each cluster to the total data volume and its blending situation of a preferred embodiment of the incremental clustering method for the atmospheric aerosol particle mass spectrum data set of the present invention; Figure 4 is a schematic diagram of the average mass spectrum of the main clusters of a preferred embodiment of the incremental clustering method for the atmospheric aerosol particle mass spectrum data set of the present invention; Figure 5 is a structural diagram of a preferred embodiment of the incremental clustering system for the atmospheric aerosol particle mass spectrum data set of the present invention; Figure 6 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the following further elaborates on the present invention with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] Clustering is an unsupervised learning technique that automatically identifies the intrinsic structure of data through metrics such as distance or density, divides data points into clusters without prior labels, and minimizes the differences within clusters and maximizes the differences between clusters. This technique can reveal the intrinsic structure patterns of data and has wide applications in fields such as computer vision (e.g., image segmentation), bioinformatics (e.g., gene expression analysis), and complex network analysis (e.g., community detection). However, modern data science faces two core challenges: the exponential growth of data dimensions and the massive scale of sample sizes, which pose requirements for clustering algorithms in terms of computational efficiency, parameter adaptability, noise robustness, and big data processing capabilities.
[0018] Various clustering algorithms achieve data partitioning through different similarity metrics and optimization strategies. The traditional clustering method system mainly includes the following types: prototype-based K-means, density-based DBSCAN (Density-Based Spatial Clustering of Applications with Noise), hierarchical clustering, spectral clustering, and clustering based on the Adaptive Resonance Theory (ART).
[0019] The K-means algorithm uses the number of clusters preset by the user as a parameter. First, it randomly initializes the centroid positions of the clusters, and through iterative minimization of the sum of the squared Euclidean distances between the data points within the clusters and the centroids, it optimizes the centroid positions and outputs the clusters and their corresponding centroid vectors. The overall time complexity of this algorithm is , with relatively high computational efficiency but unable to handle non-convex clusters. Its improved version K-means++ enhances the algorithm's stability through a probabilistic centroid initialization strategy, and the time complexity remains at order of magnitude; another improved version K-medoids algorithm (i.e., k-center clustering algorithm) uses actual data points as clustering centers, which enhances the robustness to outliers, but the computational complexity increases to . However, the clustering results of the K-means algorithm are significantly affected by the number of clusters, the selection of initial centroids, and noise, and due to the failure of distance metrics in high dimensions, this type of algorithm is not suitable for clustering large-scale high-dimensional data.
[0020] DBSCAN divides clusters based on density connectivity into core points, border points, and noise points. It does not require presetting the number of clusters and has strong anti-noise ability, but it needs to pre-specify the neighborhood radius and the minimum number of points . The algorithm expands clusters by connecting the neighborhoods of core points. The neighborhood contains at least data points. Border points belong to the neighborhood of a certain core point but do not meet the core point conditions themselves. Those not covered by any cluster are noise points. DBSCAN needs to calculate the distances between all pairs of data points, and the complexity is . After using a spatial index such as the R-tree (a multi-dimensional data structure for spatial data indexing) to optimize neighborhood queries, the complexity can be reduced to . Variations of DBSCAN include HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) and OPTICS (Ordering Points To Identify the Clustering Structure), etc. The clustering results need to be empirically and manually divided. And due to the failure of distance metrics in high-dimensional data, the DBSCAN algorithm is not applicable to high-dimensional big data
[0021] Hierarchical clustering is divided into two types: divisive and agglomerative. Divisive hierarchical clustering recursively divides data from top to bottom based on the maximum distance criterion, and the computational complexity is extremely high, being . Agglomerative hierarchical clustering merges the nearest neighbor clusters from bottom to top and finally generates a dendrogram. In the initial stage, the distances between all pairs of data points are calculated. Subsequently, the closest clusters are iteratively merged and the distance matrix is updated until all points are grouped into one cluster. The complexity of agglomerative clustering is , and it can be reduced to after using a priority queue to optimize distance updates. The space complexity of storing the distance matrix is , it has a large memory overhead. Although hierarchical clustering can provide an intuitive cluster hierarchy, its high computational and storage costs limit its application in large-scale high-dimensional data. As a variant of it, the BIRCH algorithm (Balanced Iterative Reducing and Clustering using Hierarchies) compresses data by constructing a clustering feature tree to meet the needs of big data. However, when the data dimension exceeds 20, the statistical information of the CF tree becomes invalid, resulting in a significant decline in clustering quality; in addition, the clustering results of BIRCH are greatly affected by the data insertion order. Points in the same cluster may be assigned to different subtrees due to different insertion orders, and the final clustering results may deviate greatly from the true distribution; moreover, the construction of the CF tree (Clustering Feature Tree) requires adjusting the branching factor, leaf node capacity, and radius threshold, and the parameter selection is complex, which is likely to cause the tree structure to be unbalanced or the cluster division to be unreasonable. Therefore, in practice, the use of the BIRCH algorithm is highly restricted in big data.
[0022] Spectral Clustering needs to select a similarity matrix construction method (such as Gaussian kernel) and the number of clusters, and outputs clusters based on graph partitioning. The algorithm first constructs a similarity matrix, calculates the normalized Laplacian matrix and performs eigen decomposition, selects the first eigenvectors to form a low-dimensional embedding space, and finally applies K-means clustering to obtain the results. Spectral Clustering can identify complex non-convex clusters, but its time complexity is high. The complexity of constructing the similarity matrix is , the complexity of eigen decomposition is , and the total complexity is , and a large amount of space is required for calculation. The space complexity of storing the similarity matrix is , so it is difficult to scale to big data.
[0023] Clustering based on ART theory is a type of incremental learning, which can adaptively generate clusters without specifying the number of clusters. The principle is as follows: The new data point is compared with the existing clusters. If the similarity (such as cosine similarity) exceeds the preset threshold, it is assigned to the cluster and the cluster center is updated; otherwise, a new cluster is created. The time complexity of this method is Although ART can cluster a large amount of high-dimensional data after fixing the number of clusters, the selection of the learning rate requires empirical attempts to make the clustering process converge. In addition, once the clusters are generated in this method, they will not disappear, resulting in a large number of fragmented and redundant clusters with a lot of overlap between them. Coupled with the iteration based on the learning rate, the cluster centers do not coincide with the cluster averages, and subsequent empirical merging by humans is still required, which is difficult to achieve when dealing with a large amount of data. Therefore, when applying clustering based on the ART theory, the results are not intuitive and difficult to quantify, and it is difficult to be effectively applied to the clustering of big data.
[0024] To solve the above problems, the present invention proposes a flexible incremental clustering algorithm combining similarity and density (FASC, Flexible Adaptive Similarity Clustering), which can give a set of clusters that meet the constraints of inter-cluster similarity and intra-cluster similarity. The clustering results are intuitive and quantitative, the input parameters are simple, and it is easy to deploy and use. The cluster to which each data point belongs can be flexibly dynamically selected by two strategies, including the traditional similarity-first strategy and a specially designed density-first strategy. Among them, the latter can significantly reduce the problem of inter-cluster confusion in principle. The number of clusters is optimized by an innovative cluster assignment mechanism, similarity-based dynamic cluster merging, and cluster tidying methods, and is adaptively generated and annihilated online to the optimal, while automatically identifying noise data. The FASC algorithm supports real-time data stream processing and uses an incremental update method when updating clusters to avoid global recomputation. Therefore, in terms of efficiency, the time and space complexities of the FASC algorithm are significantly better than traditional clustering methods, and it is particularly suitable for the mining of high-dimensional big data.
[0025] The incremental clustering method for the mass spectrometry dataset of atmospheric aerosol particles according to a preferred embodiment of the present invention, as Figure 1 shown, the incremental clustering method for the mass spectrometry dataset of atmospheric aerosol particles includes the following steps: Step S10: Obtain the mass spectrometry dataset of atmospheric aerosol particles, and perform clustering initialization processing on the mass spectrometry dataset of atmospheric aerosol particles to obtain a plurality of clustering initialization parameters. The clustering initialization parameters include target clusters, the cluster centers of each target cluster, the number of data points in each target cluster, and cluster numbers.
[0026] The FASC algorithm set by the present invention can cluster any dataset without any prior information (the FASC clustering algorithm in the present invention is applicable to any dataset, but in this embodiment, it is mainly applied to the mass spectrometry dataset of atmospheric aerosol particles) into clusters that satisfy the inter-cluster similarity less than a threshold (i.e., the preset inter-cluster similarity in the present invention) and the intra-cluster similarity greater than a threshold Under the condition of [[ID=]], the data points within each cluster have a high similarity, while the similarity between different clusters is low, and these similarities are quantitative and controllable.
[0027] For the convenience of description, in the present invention, the symbol for the logical operation "equal" is "==", and the assignment symbol is "="; the operation to obtain the size of a set X is X.size, the operation to obtain the length of a data structure Y in the form of an array is Y.length, the center of a certain cluster is vector, and the cluster count is count (the same meaning will be followed in the following text. For example, represents the cluster count, represents the center of a certain cluster, and will not be elaborated one by one in the following text).
[0028] The input parameters are as follows: The present invention is introduced with the mass spectrometry dataset of atmospheric aerosol particles. The data matrix of the mass spectrometry dataset of atmospheric aerosol particles is , where is the number of samples, is the number of feature dimensions. Each row of [[ID=]] is a data vector, and each column is a feature dimension. The method to convert a dataset containing samples into a data matrix is as follows: The value in [[ID=]] is the value of the th dimension of the th sample. If there is no value or the value is 0, then . The following parameters are user-defined according to the application scenario and requirements: The inter-cluster similarity threshold (i.e., the preset inter-cluster similarity in the present invention): ; The intra-cluster similarity threshold (i.e., the preset intra-cluster similarity in the present invention): ; The initial maximum number of clusters: , if there is no limit, let ; The cyclic maximum number of clusters: , if there is no limit, let ; The maximum number of iterative loops: , if there is no limit, let ; The similarity convergence threshold between iterations: , if absolute convergence is required, let ; The similarity algorithm: , the input values and are vectors or matrices with the same feature dimensions, and the output is a similarity value or matrix; The cluster selection strategy is: , where is density priority, is similarity priority; The method for calculating the inter-cluster similarity is: , where is the cluster distribution, is the cluster number.
[0029] Specifically, obtain the mass spectrometry data set of atmospheric aerosol particles and convert the mass spectrometry data set of atmospheric aerosol particles into a target data matrix; obtain the vector of each row in the target data matrix, and perform L2-norm based normalization processing on the vector to obtain a normalized data vector.
[0030] Determine multiple random sample vectors in the normalized data vector, where the random sample vector is a vector randomly extracted from the normalized data vector, and set each random sample vector as the cluster center to obtain multiple cluster centers and multiple target clusters; perform initialization processing on each target cluster to obtain the number of data points in each target cluster and the cluster number.
[0031] The clustering process of the FASC algorithm in the present invention consists of three stages: Stage 1: Clustering initialization. Let , where is the number of iterations; perform L2-norm based normalization on each row of (for a vector, the L2 norm is its modulus length, and the L2 normalization here is to divide the vector represented by each row in by its modulus length so that its modulus length is 1). Take a random number between 1 and , and select the random sample vector with the row index of in , and let it be the center of the first cluster: ; and let the cluster count ( ) be 1: , where is the vector of the first cluster center, is the vector corresponding to this cluster center, is the count corresponding to this cluster center, and so on. Subsequently, initialize the cluster count array ( ): Let , which stores the number of data points assigned to each cluster; initialize the array ( ) that stores the cluster number to which each data point belongs: .
[0032] Step S20, determine the target cluster selection strategy, and perform the first similarity calculation and cluster assignment processing on multiple clustering initialization parameters by using the target cluster selection strategy to obtain multiple cluster center matrices. The target cluster selection strategy includes a similarity priority strategy and a density priority strategy.
[0033] Specifically, an arbitrary random number generation algorithm is used to construct an integer array according to the number of data points in each of the target clusters and the cluster number, and each data point in the integer array is sequentially extracted.
[0034] Phase 2: Perform similarity calculation, cluster selection, and data allocation. Update the iteration count. , record the cluster count array of the previous loop , and the number array of the previous loop . Use an arbitrary random number generation algorithm (such as a pseudo-random number generation algorithm based on modulo operation) to obtain an integer array with a length of and a range from 1 to , and sequentially select from , is a randomly selected array from , and sequentially and without repetition, select the data vectors with row indices . Such random selection of data vectors can avoid the optimization process falling into a local optimum. If , then calculate the similarity between and all cluster centers at this time: , , is the transpose of the matrix formed by arranging all cluster center vectors row by row. Among them, the similarity algorithm can be selected according to the application scenario and requirements. Here, the cosine similarity is taken as an example, and so on. , where is 's L2 norm. If , directly pre-calculate the similarity matrix of all data points and all cluster centers , that is , is the cluster center of the th cluster. At this time . Specifically, when using the cosine similarity metric, , this matrix operation process can significantly reduce the computational complexity in engineering implementation. Subsequently, perform cluster allocation and update. The cluster to which
[0035] belongs is selected according to one of the following two strategies. For general data, the density-first strategy is preferably used; if it is known in advance that the data can be clearly divided into spherical clusters with obvious boundaries, or if there is no need to avoid the problem of cluster mixing, the similarity-first strategy can be used.If the target cluster selection strategy is the similarity - priority strategy, calculate the similarity between each data point and multiple target clusters, and determine the first assigned target cluster corresponding to each data point when the similarity is the largest, where the first assigned target cluster is the target cluster with the largest similarity to the data point among multiple target clusters; add each data point to the corresponding first assigned target cluster to obtain multiple cluster - center matrices.
[0036] A. Similarity - priority strategy: Select ( is the number of clusters) the cluster with the highest similarity to , where , (indicating the input), and the corresponding similarity is (indicating the output).
[0037] If the target cluster selection strategy is the density - priority strategy, calculate the similarity between each data point and multiple target clusters; determine the preset intra - cluster similarity, and extract multiple second - assigned target clusters whose similarity is greater than or equal to the preset intra - cluster similarity; calculate the cluster density of each second - assigned target cluster, and obtain the third - assigned target cluster with the largest cluster density among the second - assigned target clusters, where the third - assigned target cluster is the target cluster with the largest cluster density among multiple second - assigned target clusters; add each data point to the corresponding third - assigned target cluster to obtain multiple cluster - center matrices.
[0038] B. Density - priority strategy: Take the index that satisfies , is the data point whose similarity is greater than or equal to the preset intra - cluster similarity, is to calculate the similarity between the data point and each cluster, and include the index in the set . If , then select the cluster with the largest density in , if the densities of multiple clusters are the same, preferentially select the one with the highest similarity, that is, , , and the corresponding similarity is . If , let , .
[0039] As shown in Figure 2 ( Figure 2 different colors in represent different clusters), according to theoretical analysis, the cluster mixing will occur when the distance between cluster centers is less than the inter - cluster similarity in the similarity - priority selection strategy, while the inter - cluster mixing situation of the density - priority strategy will only occur whenFigure 2 In the extreme case, therefore, the density-first strategy significantly reduces the mixing frequency of clusters in theory.
[0040] The pseudo-code for cluster selection is as follows:
[0041] Subsequently, the cluster assignment process is carried out: If (i.e., the similarity is less than the preset intra-cluster similarity), when the current number of clusters does not exceed the preset threshold ( when , indicating that the current number of clusters is less than the maximum number of clusters in the loop ; when , indicating that the current number of clusters is less than or equal to the initial maximum number of clusters ), the following steps are taken: , indicating that the length of the current cluster is incremented by 1; , indicating that the updated cluster center of the current cluster is assigned as ; Generate a new cluster, indicating that the updated cluster count of the current cluster is assigned 1. Specifically: First, the total number of clusters is incremented by 1, then a new cluster numbered with the current number of clusters is created and its count is set to 1. Assign to this cluster ( ). At and , or at and , outlier processing is triggered, i.e., the point is regarded as an outlier: Set the cluster number of this point to -1 ( ). Conversely, if , at the following operations are carried out for incremental cluster update: , where is the cluster count of cluster ; , where is the cluster center of cluster ; .
[0042] Subsequently, the current vector is assigned a cluster number ( ). At , the cluster number is directly assigned without cluster update. This stage of operation continues until all the numbers in are selected once.
[0043] Step S30: Perform merging processing and denoising processing on the multiple cluster center matrices to obtain an initial clustering result, and perform a second similarity calculation on the initial clustering result to obtain an iterative similarity.
[0044] Specifically, calculate the inter-cluster similarity between the multiple cluster center matrices, and extract the cluster center matrices with the inter-cluster similarity greater than a preset inter-cluster similarity to obtain a cluster index set; perform merging processing on the cluster index set to obtain a non-redundant cluster set.
[0045] Phase 3: Dynamic merging and sorting of redundant clusters: The first part of this phase, namely dynamic merging of redundant clusters, is only performed when. For the cluster center matrices generated in Phase 2 Performing dynamic merging can reduce redundant clusters and accelerate the convergence speed.
[0046] The method is as follows: Start from the first cluster (let ), enter the dynamic cluster merging and execute the following loop: Calculate the inter-cluster similarity between each cluster and other clusters (here, the cosine similarity is used as an example) , is the cluster center of the th cluster. According to the logical operation Obtain the set of cluster indices to be merged , are the clusters to be merged. If , it means there are no redundant clusters to merge, then let , return for calculation, and repeat the dynamic cluster merging loop. Conversely, if , it proves that there are redundant clusters to be merged and updated. The cluster count expression of ; Merge by weighted average into a new cluster and replace , and the replaced expression is: .
[0047] Correspondingly, update to match the merged cluster, and the expression is: , that is, until the size of the cluster index set is less than 2, indicating that there are no redundant clusters to merge.
[0048] Subsequently, let , reset the index, and delete the redundant clusters. The expression is: .
[0049] The dynamic cluster merging loop continues until no clusters can be merged ( Stop when the number of clusters has reached the threshold ( ).
[0050] The pseudo - code of the redundant cluster dynamic merging algorithm is as follows:
[0051] Delete the clusters with 0 data vectors in the redundant - cluster set, and perform cluster sorting and cluster numbering to obtain the initial clustering result.
[0052] The second part of this stage is cluster arrangement, which needs to be carried out after any iteration loop. The main task is to delete the clusters that occupy 0 data vectors during the iteration and calculate the inter - iteration similarity.
[0053] The steps of cluster arrangement after iteration are: , is the count of the cluster with cluster number , is the size of the set of row indices of the data vectors; ; Take an integer without repetition , , and obtain the set of row indices of the data vectors with the assigned cluster number . If , then delete , otherwise update .
[0054] After that, use any sorting algorithm (including quick sort, merge sort, and heap sort, etc.) to sort the clusters in descending order according to the value and update the cluster numbers in to match the sorted . There are many simple ways to implement the update. An intuitive implementation is to permute the cluster numbers in corresponding to the transposition operations performed on during sorting. The pseudo - code of cluster arrangement is as follows:
[0055] Calculate the inter - iteration similarity for the initial clustering result to obtain the iteration similarity. Among them, the inter - iteration similarity calculation includes the inter - iteration similarity calculation based on cluster distribution and the inter - iteration similarity calculation based on cluster numbers.
[0056] The inter - iteration similarity can be calculated in the following two ways:
[0057] For calculating the inter - iteration similarity, the following two methods can be used: 1. Iterative similarity calculation based on cluster distribution ( ) First, obtain the cluster count array , and the expression is: ; Subsequently, calculate the iterative similarity between and , , and the expression of is: Using this method can quickly achieve convergence under a relatively high iterative similarity threshold, and it is more applicable to the case where the distribution of cluster counts is relatively uniform.
[0058] 2. Iterative similarity calculation based on cluster number ( ), and the expression is: ; ; Calculate the repetition ratio of at each index in the current iteration and the previous iteration, and its value is . This method is for strict convergence and is applicable in any case.
[0059] The pseudo-code for iterative similarity calculation is as follows:
[0060] Step S40. Perform clustering convergence determination and iterative loop clustering processing on the initial clustering result according to the iterative similarity, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset.
[0061] Specifically, determine a preset convergence value, and judge whether the iterative similarity is less than the preset convergence value; if so, perform iterative loop clustering processing on the initial clustering result, where the input of each iteration is the clustering result output by the previous iteration, until the iterative similarity is greater than or equal to the preset convergence value or the number of iterations is greater than the preset iteration number threshold, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset.
[0062] After the end of stage 3, the clustering enters the next iterative loop and starts running from stage 1 again to continuously update the cluster partition until convergence ( ) or exceeds the iteration number threshold ( ). When performing incremental learning, for streaming input data, you can set and perform continuous clustering on the input data. The main pseudo-code of the FASC algorithm is as follows:
[0063]
[0064]
[0065]
[0066] Complexity Analysis of the FASC Algorithm: The time complexity of FASC mainly consists of the following parts. 1. In data preprocessing, the time complexity of the L2 normalization operation is . 2. Cluster assignment and update: In each iteration, the similarity between each sample and all cluster centers needs to be calculated, and the time complexity is ( is the number of dynamic clusters, that is, ). 3. Dynamic cluster merging: By screening and merging cluster pairs through a similarity threshold, the complexity is . 4. The comprehensive time complexity is . FASC reduces redundant clusters through dynamic merging. In big data , it can be regarded as a constant, and the total algorithm time complexity can be approximated as , which is better than traditional algorithms. In terms of space complexity, FASC only needs to store the cluster center matrix (with a complexity of ) and the inter-cluster similarity matrix (with a complexity of ), and the total space complexity is . Compared with traditional methods, such as the complexity of spectral clustering being and the complexity of K-means being , FASC occupies less memory space during calculation and is suitable for high-dimensional big data scenarios.
[0067] For an application example of the FASC algorithm set in the present invention: Taking a big data set of atmospheric aerosol particle mass spectra as an example, this data set contains a total of 10 million 600-dimensional data: each data is a single-particle laser desorption ionization mass spectrum of a real environmental atmospheric aerosol particle, and each dimension is a specific mass-to-charge ratio, ranging from -300 to +300, with integer precision. Implement the above FASC algorithm using the matlab language and calculate using a single CPU core. The preset parameters are as follows: , ; ; ; ; ; Since the mass spectrometry data set is a sparse matrix, cosine similarity is selected as the similarity metric. After three iterative loops, the algorithm converges, taking 12,867 seconds. 50 clusters that meet the preset conditions are obtained; there are 1.33 million outliers, accounting for 13.3% of the total. Due to the complex real atmospheric environmental conditions and the large heterogeneity of particulate matter components, the number of outliers is in line with expectations.
[0068] As Figure 3 shown, it is the ratio of the data contained in each cluster to the total data volume. There is no mixing pair, indicating that there is no mixing between clusters, which proves that FASC effectively solves the problem of inter-cluster mixing of traditional algorithms. In addition, according to the previous theoretical description of the FASC algorithm, it is easy to know that the cluster center coincides with the cluster average in the clustering result of FASC. Therefore, the cluster center can be used as the representative of the data contained in the cluster. The central vectors of the 5 main clusters are shown in the form of mass spectrometry as follows (as Figure 4 shown). Clusters 1 to 5 are common particulate matter categories in the atmospheric environment, which is in line with the previous conventional observation results. These results prove the correctness of FASC clustering.
[0069] The key points of innovation of the present invention mainly include: 1. The specially designed clustering process can give quantitative and intuitive clustering results under simple custom parameters. In the cluster assignment mechanism of data, an online dynamic generation and annihilation mechanism of clusters is proposed, which can obtain the optimal number of clusters without pre-specifying the number of clusters. In the selection of the cluster to which the data belongs and the merging of clusters, a density-priority cluster selection strategy and a method of dynamically merging and sorting clusters are innovatively proposed, taking into account both cluster density and similarity, and significantly reducing inter-cluster mixing. In the generation of clusters, the algorithm supports the specification of inter-cluster similarity and intra-cluster similarity, and the clustering result is controllable and quantitative, and the interpretability and usability are better than traditional methods.
[0070] 2. Deeply optimized constraint conditions, operation steps and convergence determination methods make this algorithm particularly suitable for high-dimensional big data. Restricting the initial and cyclic number of clusters and pre-computing the similarity within the loop significantly reduces the time complexity of the algorithm. The efficient clustering process only needs to store the information of the clusters and the data belonging, with low space complexity, which is significantly better than traditional algorithms. The flexible selection of the inter-iteration similarity calculation method enables the present invention to converge quickly under different data conditions.
[0071] In addition, in the dynamic adjustment and optimization of the number of clusters, the present invention innovatively uses the method of online generation and merging and sorting, which is superior to the traditional incremental learning algorithm. Among them, after the clusters are sorted, a sorting process of counting by cluster is performed. In actual engineering implementation, sorting can be omitted and the calculation process of similarity between iterations can be modified accordingly. In addition, the merging method of redundant clusters can be implemented in other ways, and the present invention only shows the most intuitive one. The determination of the convergence condition can have other methods, such as the number of clusters, the number of data points contained in the cluster, or the proportion of the total number of data points, etc.
[0072] In addition, the present invention is set at When using the method of pre-computing similarity to accelerate the calculation, a possible variant is that at When, the similarity pre-computation between the data and the clusters is not performed, and the same dynamic cluster generation method of detecting data points one by one as at When is still used. The computational complexity of this method is relatively high, but it may converge faster.
[0073] Furthermore, as Figure 5 Shown, based on the above incremental clustering method for the atmospheric aerosol particle mass spectrometry dataset, the present invention also correspondingly provides an incremental clustering system for the atmospheric aerosol particle mass spectrometry dataset, wherein, the incremental clustering system for the atmospheric aerosol particle mass spectrometry dataset includes: A clustering initialization module 51, configured to obtain the atmospheric aerosol particle mass spectrometry dataset, and perform clustering initialization processing on the atmospheric aerosol particle mass spectrometry dataset to obtain a plurality of clustering initialization parameters; A cluster selection strategy determination module 52, configured to determine a target cluster selection strategy, and perform a first similarity calculation and cluster assignment processing on the plurality of clustering initialization parameters by using the target cluster selection strategy to obtain a plurality of cluster center matrices; An inter-iteration similarity calculation module 53, configured to perform merging processing and denoising processing on the plurality of cluster center matrices to obtain an initial clustering result, and perform a second similarity calculation on the initial clustering result to obtain an iterative similarity; A clustering result output module 54, configured to perform clustering convergence determination and iterative loop clustering processing on the initial clustering result according to the iterative similarity, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset.
[0074] Furthermore, as Figure 6 Shown, based on the above incremental clustering method and system for the atmospheric aerosol particle mass spectrometry dataset, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0075] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In some other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, an incremental clustering program 40 for an atmospheric aerosol particle mass spectrometry dataset is stored on the memory 20, and this incremental clustering program 40 for the atmospheric aerosol particle mass spectrometry dataset can be executed by the processor 10, thereby implementing the incremental clustering method for the atmospheric aerosol particle mass spectrometry dataset in this application.
[0076] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing the incremental clustering method for the atmospheric aerosol particle mass spectrometry dataset, etc.
[0077] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display information on the terminal and to display a visual user interface.
[0078] In one embodiment, when the processor 10 executes the incremental clustering program 40 for the atmospheric aerosol particle mass spectrometry dataset in the memory 20, the steps of the incremental clustering method for the atmospheric aerosol particle mass spectrometry dataset are implemented.
[0079] In summary, the present invention provides an incremental clustering method, system and terminal for an atmospheric aerosol particle mass spectrometry data set. The method includes: obtaining an atmospheric aerosol particle mass spectrometry data set, and performing clustering initialization processing on the atmospheric aerosol particle mass spectrometry data set to obtain a plurality of clustering initialization parameters; determining a target cluster selection strategy, and using the target cluster selection strategy to perform a first similarity calculation and cluster assignment processing on the plurality of clustering initialization parameters to obtain a plurality of cluster center matrices; performing a merging process and a denoising process on the plurality of cluster center matrices to obtain an initial clustering result, and performing a second similarity calculation on the initial clustering result to obtain an iterative similarity; performing a clustering convergence determination and an iterative loop clustering process on the initial clustering result according to the iterative similarity, and outputting a clustering result of the atmospheric aerosol particle mass spectrometry data set. By using the target cluster selection strategy to perform similarity calculation and cluster assignment processing on the clustering initialization parameters corresponding to the atmospheric aerosol particle mass spectrometry data set, the present invention enables the cluster to which each data point belongs to be dynamically selected by two strategies flexibly, including the traditional similarity priority strategy and the specially designed density priority strategy, and can generate a cluster set that satisfies the cluster - to - cluster similarity and within - cluster similarity constraint conditions. Further, by performing a merging process and a denoising process on the cluster center matrices, and performing an iterative loop clustering process according to the inter - iterative similarity, the inter - cluster confusion can be significantly reduced, and the clustering analysis efficiency of the atmospheric aerosol particle mass spectrometry data set and the accuracy of the clustering result output under the condition of high - dimensional big data can be effectively improved.
[0080] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non - exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal including that element.
[0081] Certainly, those of ordinary skill in the art can understand that all or part of the processes of implementing the above - mentioned method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer - readable storage medium readable by a computer. When the program is executed, it can include the processes of the above - mentioned method embodiments. The computer - readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0082] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of the present invention.
Claims
1. An incremental clustering method for an atmospheric aerosol particle mass spectrometry data set, characterized in that, The incremental clustering method for the atmospheric aerosol particle mass spectrometry dataset includes: Obtain the atmospheric aerosol particle mass spectrometry dataset, and perform clustering initialization processing on the atmospheric aerosol particle mass spectrometry dataset to obtain multiple clustering initialization parameters; Determine the target cluster selection strategy, and use the target cluster selection strategy to perform the first similarity calculation and cluster assignment processing on multiple clustering initialization parameters to obtain multiple cluster center matrices; Perform merging processing and denoising processing on multiple cluster center matrices to obtain an initial clustering result, and perform the second similarity calculation on the initial clustering result to obtain an iterative similarity; Perform clustering convergence determination and iterative loop clustering processing on the initial clustering result according to the iterative similarity, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset.
2. The incremental clustering method for the atmospheric aerosol particle mass spectrometry data set according to claim 1, wherein The clustering initialization parameters include target clusters, the cluster center of each target cluster, the number of data points in each target cluster, and the cluster number; The obtaining the atmospheric aerosol particle mass spectrometry dataset and performing clustering initialization processing on the atmospheric aerosol particle mass spectrometry dataset to obtain multiple clustering initialization parameters specifically includes: Obtain the atmospheric aerosol particle mass spectrometry dataset, and convert the atmospheric aerosol particle mass spectrometry dataset into a target data matrix; Obtain the vector of each row in the target data matrix, and perform normalization processing based on the L2 norm on the vector to obtain a normalized data vector; Determine multiple random sample vectors in the normalized data vector, where the random sample vector is a vector randomly extracted from the normalized data vector, and set each random sample vector as the cluster center to obtain multiple cluster centers and multiple target clusters; Perform initialization processing on each target cluster to obtain the number of data points in each target cluster and the cluster number.
3. The incremental clustering method for the mass spectrometry data set of atmospheric aerosol particles according to claim 2, wherein The target cluster selection strategy includes a similarity priority strategy and a density priority strategy; The determining the target cluster selection strategy and using the target cluster selection strategy to perform the first similarity calculation and cluster assignment processing on multiple clustering initialization parameters to obtain multiple cluster center matrices specifically includes: Use any random number generation algorithm to construct an integer array according to the number of data points in each target cluster and the cluster number, and sequentially extract each data point in the integer array; If the target cluster selection strategy is the similarity priority strategy, calculate the similarity for each data point, and perform cluster assignment processing according to the similarity priority strategy to obtain multiple cluster center matrices; If the target cluster selection strategy is the density priority strategy, calculate the similarity for each data point, and perform cluster assignment processing according to the density priority strategy to obtain multiple cluster center matrices.
4. The incremental clustering method for the atmospheric aerosol particle mass spectrometry data set according to claim 3, characterized in that The calculating the similarity for each data point and performing cluster assignment processing according to the similarity priority strategy to obtain multiple cluster center matrices specifically includes: Calculate the similarity between each data point and multiple target clusters, and determine the first assigned target cluster corresponding to each data point when the similarity is the largest, where the first assigned target cluster is the target cluster with the largest similarity to the data point among the multiple target clusters; Add each data point to the corresponding first assigned target cluster to obtain multiple cluster center matrices.
5. The incremental clustering method for the atmospheric aerosol particle mass spectrometry data set according to claim 3, wherein The similarity calculation for each data point and the cluster assignment process according to the density priority strategy to obtain multiple cluster center matrices specifically include: Calculate the similarity between each data point and multiple target clusters; Determine the preset intra-cluster similarity, and extract multiple second assigned target clusters whose similarity is greater than or equal to the preset intra-cluster similarity; Calculate the cluster density of each second assigned target cluster, and obtain the third assigned target cluster with the largest cluster density among the second assigned target clusters, where the third assigned target cluster is the target cluster with the largest cluster density among the multiple second assigned target clusters; Add each data point to the corresponding third assigned target cluster to obtain multiple cluster center matrices.
6. The incremental clustering method for the atmospheric aerosol particle mass spectrometry data set according to claim 1, characterized in that, The merging process and denoising process for multiple cluster center matrices to obtain the initial clustering result, and the second similarity calculation for the initial clustering result to obtain the iterative similarity specifically include: Calculate the inter-cluster similarity between multiple cluster center matrices, and extract the cluster center matrices whose inter-cluster similarity is greater than the preset inter-cluster similarity to obtain a cluster index set; Perform a merging process on the cluster index set to obtain a redundant cluster-free set; Delete the clusters with zero data vectors in the redundant cluster-free set, and perform cluster sorting and cluster numbering to obtain the initial clustering result; Perform iterative inter-similarity calculation on the initial clustering result to obtain the iterative similarity, where the iterative inter-similarity calculation includes iterative inter-similarity calculation based on cluster distribution and iterative inter-similarity calculation based on cluster numbering.
7. The incremental clustering method for the mass spectrometry data set of atmospheric aerosol particles according to claim 1, wherein The clustering convergence determination and iterative loop clustering process for the initial clustering result according to the iterative similarity, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset specifically include: Determine the preset convergence value, and judge whether the iterative similarity is less than the preset convergence value; If so, perform iterative loop clustering on the initial clustering result, where the input of each iteration is the clustering result output by the previous iteration, until the iterative similarity is greater than or equal to the preset convergence value or the number of iterations is greater than the preset iteration number threshold, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset.
8. An incremental clustering system for an atmospheric aerosol particle mass spectrometry data set, characterized in that, The incremental clustering system for the atmospheric aerosol particle mass spectrometry dataset includes: A clustering initialization module for obtaining the atmospheric aerosol particle mass spectrometry dataset and performing clustering initialization processing on the atmospheric aerosol particle mass spectrometry dataset to obtain multiple clustering initialization parameters; A cluster selection strategy determination module for determining the target cluster selection strategy and performing the first similarity calculation and cluster assignment process on multiple clustering initialization parameters using the target cluster selection strategy to obtain multiple cluster center matrices; The inter-iteration similarity calculation module is used to perform merging processing and denoising processing on multiple said cluster center matrices to obtain an initial clustering result, and perform a second similarity calculation on the initial clustering result to obtain an iterative similarity; The clustering result output module is used to perform clustering convergence determination and iterative loop clustering processing on the initial clustering result according to the iterative similarity, and output the clustering result of the atmospheric aerosol particle mass spectrometry dataset.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an incremental clustering program of the atmospheric aerosol particle mass spectrometry dataset stored on the memory and executable on the processor. When the incremental clustering program of the atmospheric aerosol particle mass spectrometry dataset is executed by the processor, the steps of the incremental clustering method of the atmospheric aerosol particle mass spectrometry dataset according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an incremental clustering program of the atmospheric aerosol particle mass spectrometry dataset. When the incremental clustering program of the atmospheric aerosol particle mass spectrometry dataset is executed by a processor, the steps of the incremental clustering method of the atmospheric aerosol particle mass spectrometry dataset according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Method for performing particle source analysis by using single-particle aerosol mass spectrometer
CN108680473A
PART2A algorithm-based particulate matter clustering method, apparatus and device, and storage medium
CN115840899A
Atmospheric particulate source analysis method and system based on single-particle aerosol mass spectrometry
CN117216659A
Aerosol component detection method based on machine learning
CN117542432A
Gravity model-based image clustering method and system
CN119445174A
Cited By
Aero-engine group health evaluation method based on multi-working-condition dynamic clustering
CN120822059A
Model updating training method and device, electronic equipment and storage medium
CN121561495A
Mass spectrum data clustering method and system, terminal and storage medium
CN121786520A
Mass spectrometry data clustering method, system, terminal and storage medium
CN121786520B