Deep clustering ensemble method and apparatus based on cluster confidence, device and medium

By employing a deep clustering ensemble method based on cluster confidence and utilizing variational autoencoders and local weighting strategies, the instability of deep clustering methods in large-scale, high-dimensional unstructured data is addressed, resulting in more stable and robust clustering results.

CN116150638BActive Publication Date: 2025-11-18NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310068211.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-12
Publication Date
2025-11-18
Estimated Expiration
2043-01-12

AI Technical Summary

Technical Problem

Existing deep clustering methods are unstable and have poor robustness when dealing with large-scale, high-dimensional unstructured data, and the low-dimensional embeddings that rely on the pre-training stage are not suitable for clustering tasks.

Method used

A deep clustering ensemble method based on cluster confidence is adopted. Through variational autoencoder network pre-training, Student t-distribution and KL divergence loss training, cluster confidence scores are calculated and ranked. Clustering segmentation is performed by combining local weighting strategy and Tcut graph cutting algorithm to obtain robust ensemble clustering results.

Benefits of technology

It achieves more stable and robust clustering results, improves clustering performance, and is suitable for large-scale, high-dimensional unstructured data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150638B_ABST
    Figure CN116150638B_ABST
Patent Text Reader

Abstract

The application relates to a deep clustering ensemble method and device based on cluster confidence, equipment and medium. The method generates an input sample set by preprocessing and cleaning original data, pretrains a variational auto-encoding network and calculates cluster confidence of initial low-dimensional embedding; cluster confidence of final low-dimensional embedding generated by cluster loss training and calculation of the variational auto-encoding network based on student t distribution and KL divergence loss; cluster confidence score of the final low-dimensional embedding of the variational auto-encoding network is calculated according to the cluster confidence of the initial low-dimensional embedding and the final low-dimensional embedding, and the cluster confidence score is sorted; the target base cluster corresponding to the final low-dimensional embedding of the first set number of final low-dimensional embedding with high cluster confidence score is selected in each base cluster, the cluster reliability in each target base cluster is calculated by using a local weighting strategy, a local weighted bipartite graph is constructed and a Tcut graph cutting algorithm is used for segmentation to obtain a final integrated clustering result. The robustness and clustering performance are better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data mining and artificial intelligence technology, and in particular to a deep clustering integration method, apparatus, device and medium based on cluster confidence. Background Technology

[0002] With the advent of the 5G era, big data applications have developed rapidly. The data generated by these applications are often characterized by large scale, unstructured nature, and high dimensionality. Extracting simple yet effective information from this complex data is a highly challenging task. Data clustering analysis is a classic unsupervised machine learning method that can effectively reveal and mine the potential knowledge patterns in data. Its purpose is to group data based on similarity, density, intervals, or specific statistical distribution measures in the data space. Traditional clustering methods, such as K-means and Gaussian mixture clustering, have achieved good clustering performance in many fields. However, when faced with large-scale, high-dimensional unstructured data, traditional clustering methods are not ideal, and may even fail. This is because, on the one hand, this data often exhibits a relatively sparse distribution, making it difficult to segment; on the other hand, most traditional clustering methods can only utilize the shallow features of the data and cannot uncover the interdependencies of complex data features in the latent space.

[0003] In recent years, the emergence and development of deep clustering methods have provided a solution to this problem. Deep clustering combines the unsupervised nature of deep representation learning and clustering methods, using deep representation networks to map the original data to a low-dimensional embedding representation space, and then jointly adjusting the network parameters with the clustering algorithm to achieve clustering. However, existing deep clustering methods almost all suffer from the following limitations: the clustering results are highly dependent on the low-dimensional embeddings generated during the pre-training stage. The low-dimensional embeddings generated by the deep representation network during the pre-training stage are not necessarily suitable for the clustering task. Furthermore, the deep representation network can exhibit different mappings due to different network parameter initializations, resulting in unstable clustering results and poor robustness. Summary of the Invention

[0004] To address the problems existing in the above-mentioned traditional methods, this invention proposes a deep clustering ensemble method based on cluster confidence, a deep clustering ensemble device based on cluster confidence, a computer device, and a computer-readable storage medium, which can obtain clustering results with better robustness and clustering performance.

[0005] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:

[0006] On the one hand, a deep clustering ensemble method based on cluster confidence is provided, including the following steps:

[0007] The raw data to be processed is preprocessed and cleaned to generate the input sample set;

[0008] The variational autoencoder network is pre-trained using the input sample set, and the cluster confidence of the initial low-dimensional embeddings generated by the pre-training is calculated.

[0009] The variational autoencoder network is trained using clustering loss based on the student t-distribution and KL divergence loss, and the cluster confidence of the final low-dimensional embedding generated after training is calculated.

[0010] Based on the cluster confidence of the initial low-dimensional embedding and the cluster confidence of the final low-dimensional embedding, calculate and sort the cluster confidence scores of the final low-dimensional embedding of the variational autoencoder network.

[0011] Among the base clusters generated by the final low-dimensional embedding, select the target base clusters corresponding to the pre-determined number of final low-dimensional embeddings with high cluster confidence scores;

[0012] The cluster reliability in each target basis cluster is calculated using a local weighting strategy. A local weighted bipartite graph is constructed and segmented using the Tcut graph cutting algorithm to obtain the final ensemble clustering result of the input sample set.

[0013] In one embodiment, the process of pre-training the variational autoencoder network using an input sample set includes:

[0014] The hidden layer variables for the posterior probability of each input sample are assumed to follow a normal distribution.

[0015] In pre-training, KL divergence is used to measure the distribution of hidden layer variables and the standard normal distribution of posterior probabilities to determine the non-clustering loss;

[0016] Pre-training of variational autoencoders using the input sample set based on non-clustering loss.

[0017] In one embodiment, the process of calculating the cluster confidence of the initial low-dimensional embeddings generated during pre-training includes:

[0018] Calculate the confidence score of each cluster of each variational autoencoder in the variational autoencoder network;

[0019] The cluster confidence of the initial low-dimensional embedding of each variational autoencoder is calculated based on the confidence of each cluster of each variational autoencoder.

[0020] In one embodiment, the process of training the variational autoencoder network with clustering loss based on the Student t-distribution and KL divergence loss includes:

[0021] The Student t-distribution is used to measure the similarity between the hidden layer variables and the cluster centers of the variational autoencoder network;

[0022] The clustering loss is trained by using KL divergence loss as the clustering loss between the constructed auxiliary distribution and the soft cluster assignment to be iteratively optimized.

[0023] In one embodiment, the cluster reliability in each target base cluster is calculated and measured using the following formula:

[0024]

[0025] Among them, ECE(C i ) represents cluster C i H is an integrated cluster reliability metric. Π (C i ) represents the cluster C in the integrated Π i The uncertainty is denoted by M, which represents the number of base clusters.

[0026] In one embodiment, point v in the locally weighted bipartite graph i and point v j The edge weights between them are calculated as follows:

[0027]

[0028] Among them, ECE(v i ) represents point v i Integrated cluster reliability metric, ECE(v j ) represents point v j Integrated cluster reliability metric Represents the sample set, Indicates a cluster.

[0029] In one embodiment, the pre-training loss function of the variational autoencoder network is:

[0030]

[0031] Where x represents the input sample, Let σ represent the reconstructed sample, σ represent the standard deviation of the input sample during training, and μ represent the mean of the input sample during training.

[0032] On the other hand, a deep clustering ensemble based on cluster confidence is provided, comprising:

[0033] The preprocessing module is used to preprocess and clean the raw data to be processed, and generate the input sample set;

[0034] The pre-training module is used to pre-train the variational autoencoder network using the input sample set and to calculate the cluster confidence of the initial low-dimensional embeddings generated by the pre-training.

[0035] The clustering training module is used to train the variational autoencoder network with clustering loss based on the student t-distribution and KL divergence loss, and to calculate the cluster confidence of the final low-dimensional embedding generated after the clustering loss training.

[0036] The scoring module is used to calculate and sort the cluster confidence scores of the final low-dimensional embeddings of the variational autoencoder network based on the cluster confidence scores of the initial low-dimensional embeddings and the final low-dimensional embeddings.

[0037] The cluster selection module is used to select the target base clusters with high cluster confidence scores from the base clusters generated by the final low-dimensional embeddings, based on a predetermined number of the final low-dimensional embeddings.

[0038] The ensemble clustering module is used to calculate the cluster reliability of each target base cluster using a local weighting strategy, construct a local weighted bipartite graph, and use the Tcut graph cutting algorithm to segment the local weighted bipartite graph to obtain the final ensemble clustering result of the input sample set.

[0039] On the other hand, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described deep clustering ensemble method based on cluster confidence.

[0040] Furthermore, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the aforementioned deep clustering ensemble method based on cluster confidence.

[0041] One of the above technical solutions has the following advantages and beneficial effects:

[0042] The aforementioned deep clustering ensemble method, apparatus, device, and medium based on cluster confidence involves preprocessing and cleaning the raw data to be processed to generate an input sample set. This input sample set is then used to pre-train a variational autoencoder network (VAN) and calculate the cluster confidence of the initial low-dimensional embeddings generated during pre-training. Next, clustering loss training is performed based on the Student's t-distribution and KL divergence loss. The cluster confidence of the final low-dimensional embeddings generated after clustering loss training is calculated. The cluster confidence scores of the final low-dimensional embeddings of the VVA are then calculated and sorted. A predetermined number of target base clusters corresponding to the top few final low-dimensional embeddings with high cluster confidence scores are selected as inputs for subsequent clustering ensembles. Finally, a local weighting strategy is used to calculate the cluster reliability in each target base cluster. A locally weighted bipartite graph is constructed, and the Tcut graph cutting algorithm is used to segment the locally weighted bipartite graph to obtain the final ensemble clustering result.

[0043] Compared with traditional methods, the above scheme combines cluster confidence assessment and cluster ensemble method, maps the original unlabeled sample data to low-dimensional (deep) embedding, constructs cluster confidence assessment to evaluate the quality of low-dimensional deep embedding and introduces cluster ensemble method, and achieves clustering results with better robustness and clustering performance. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating a deep clustering ensemble method based on cluster confidence in one embodiment;

[0046] Figure 2 This is a schematic diagram of the network structure of a variational autoencoder network and a cluster confidence evaluation process in one embodiment.

[0047] Figure 3 This is a schematic diagram of the training model structure of a variational autoencoder network in one embodiment;

[0048] Figure 4 This is a schematic diagram of the module structure of a deep clustering integration device based on cluster confidence in one embodiment. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application.

[0050] It should be noted that, in this document, the reference to "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The presentation of this phrase in various locations throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments.

[0051] Those skilled in the art will understand that the embodiments described herein can be combined with other embodiments. The term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items, and all possible combinations thereof.

[0052] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0053] In one embodiment, such as Figure 1 As shown, a deep clustering ensemble method based on cluster confidence includes the following processing steps S12 to S22:

[0054] S12, preprocess and clean the raw data to be processed to generate the input sample set.

[0055] It is understandable that after collecting raw data samples such as gene data, image data, and DNA sequence data to be processed, necessary preprocessing and data cleaning can be performed to obtain input samples suitable for subsequent training and clustering. The set of these input samples is called the input sample set. Specifically, computing devices can be used to collect the sample and category information contained in the dataset to be processed. Then, these collected unlabeled raw data of various types are preprocessed and cleaned to generate input samples X = {X1, X2, ..., X...} n}

[0056] S14, use the input sample set to pre-train the variational autoencoder network and calculate the cluster confidence of the initial low-dimensional embeddings generated by the pre-training.

[0057] It is understandable that the network structure of variational autoencoders and cluster confidence evaluation processes can be as follows: Figure 2 As shown, the aim is to train low-dimensional (deep) embeddings suitable for clustering tasks: The original data X is mapped to initial low-dimensional deep embeddings using a variational autoencoder network (VAE) pre-trained. To obtain suitable low-dimensional embeddings for clustering tasks, cluster confidence is constructed to evaluate S low-dimensional embeddings. Higher-scoring low-dimensional embeddings indicate better suitability for clustering tasks. S represents the total number of variational autoencoders (VAEs) included in the variational autoencoder network.

[0058] S16. The variational autoencoder network is trained with clustering loss based on the student t-distribution and KL divergence loss, and the cluster confidence of the final low-dimensional embedding generated after training with clustering loss is calculated.

[0059] It is understandable that the low-dimensional embeddings generated after pre-training the variational autoencoder network are called initial low-dimensional embeddings. After pre-training, the cluster confidence of the initial low-dimensional embeddings is calculated. The low-dimensional embeddings generated after further training the variational autoencoder network with clustering loss based on the Student's t-distribution and KL divergence loss are called final low-dimensional embeddings. At this time, the cluster confidence of these final low-dimensional embeddings is also calculated so as to accurately calculate the cluster confidence scores of the low-dimensional embeddings of the S variational autoencoders.

[0060] S18. Based on the cluster confidence of the initial low-dimensional embedding and the cluster confidence of the final low-dimensional embedding, calculate and sort the cluster confidence scores of the final low-dimensional embedding of the variational autoencoder network.

[0061] S20, among the base clusters generated by the final low-dimensional embedding, select the target base clusters corresponding to the predetermined number of final low-dimensional embeddings with high cluster confidence scores.

[0062] It is understandable that after obtaining the cluster confidence scores of the initial and final low-dimensional embeddings, the sum of the cluster confidence scores of the initial and final low-dimensional embeddings of the s-th variational autoencoder is calculated as the cluster confidence score of the final low-dimensional embedding of that s-th variational autoencoder, where s = 1, 2, ..., S. The cluster confidence scores of the final low-dimensional embeddings of all variational autoencoder networks are sorted in descending order. The top T final low-dimensional embeddings with high scores (high-quality embeddings) will be used to generate subsequent base clustering results. The specific value of the number T can be determined based on the overall size of the base clustering, ensuring that multiple high-quality and differentiated base clusters can be integrated.

[0063] S22. The cluster reliability in each target base cluster is calculated using a local weighting strategy. A local weighted bipartite graph is constructed and the Tcut graph cutting algorithm is used to segment the local weighted bipartite graph to obtain the final ensemble clustering result of the input sample set.

[0064] It is understandable that after generating multiple base clustering results through the aforementioned steps, the base clusters generated by the top T low-dimensional embeddings with high scores are selected. The reliability of the clusters in the base clusters is evaluated using a local weighting strategy. A local weighted bipartite graph is constructed, and then the Tcut graph cutting algorithm is used to segment the local weighted bipartite graph to obtain the final ensemble clustering result, thereby achieving the purpose of clustering ensemble processing of the original data.

[0065] The aforementioned deep clustering ensemble method based on cluster confidence preprocesses and cleans the original data to generate an input sample set. This input sample set is then used to pre-train a variational autoencoder network (VAN) and calculate the cluster confidence of the initial low-dimensional embeddings generated during pre-training. Next, clustering loss training is performed based on the Student's t-distribution and KL divergence loss. The cluster confidence of the final low-dimensional embeddings generated after clustering loss training is calculated. The cluster confidence scores of the final low-dimensional embeddings of the VVA are then calculated and sorted. A predetermined number of target base clusters with high cluster confidence scores are selected as inputs for subsequent clustering ensembles. Finally, a local weighting strategy is used to calculate the cluster reliability in each target base cluster. A locally weighted bipartite graph is constructed, and the Tcut graph cutting algorithm is used to segment the locally weighted bipartite graph to obtain the final ensemble clustering result.

[0066] Compared with traditional methods, the above scheme combines cluster confidence assessment and cluster ensemble method, maps the original unlabeled sample data to low-dimensional (deep) embedding, constructs cluster confidence assessment to evaluate the quality of low-dimensional deep embedding and introduces cluster ensemble method, and achieves clustering results with better robustness and clustering performance.

[0067] In one embodiment, the process of pre-training the variational autoencoder network using the input sample set in step S14 above may specifically include the following steps:

[0068] The hidden layer variables for the posterior probability of each input sample are assumed to follow a normal distribution.

[0069] In pre-training, KL divergence is used to measure the distribution of hidden layer variables and the standard normal distribution of posterior probabilities to determine the non-clustering loss;

[0070] Pre-training of variational autoencoders using the input sample set based on non-clustering loss.

[0071] It is understandable that the training model structure of variational autoencoders can be as follows: Figure 3 As shown. Variational autoencoders differ from general autoencoders; they are probabilistic autoencoders. The encoding process involves encoding the original data into a distribution in the hidden space. It is assumed that the hidden layer variable z, representing the posterior probability of each input sample, follows a normal distribution p(z|x), so that the hidden layer variable z collected by p(z) corresponds only to a specific input sample. During training, each input sample x... i Both targets need to be trained, namely the mean μ. i and logarithmic variance During the decoding process, the input samples are resampled based on the posterior probability. This process may be affected by noise. To ensure the model's generative capability, variational autoencoders allow all posterior distributions p(z|x) to approximate the standard normal distribution, i.e., p(z|x) → N(0,1). Here, KL divergence is used to measure the two, and the calculation formula is as follows:

[0072]

[0073] Therefore, the pre-training loss function (non-clustering loss) for the variational autoencoder network is further as follows:

[0074]

[0075] Where x represents the input sample, Let σ represent the reconstructed sample, μ represent the standard deviation of the input sample during training, and μ represent the mean of the input sample during training. Pre-training is considered complete after reaching the set number of iterations.

[0076] In one embodiment, further, regarding step S14 above, the process of calculating the cluster confidence of the initial low-dimensional embeddings generated by pre-training may specifically include the following processing:

[0077] Calculate the confidence score of each cluster of each variational autoencoder in the variational autoencoder network;

[0078] The cluster confidence of the initial low-dimensional embedding of each variational autoencoder is calculated based on the confidence of each cluster of each variational autoencoder.

[0079] It is understandable that since low-dimensional embeddings generated by pre-training variational autoencoders are not always suitable for clustering tasks, a cluster confidence metric can be constructed to evaluate the quality of low-dimensional embeddings. For any variational autoencoder, the cluster confidence cc of the j-th cluster is... j The calculation is constructed as shown in the following formula:

[0080]

[0081] in, This represents the highest soft clustering probability that sample i belongs to cluster j. This indicates that sample i belongs to the second highest soft cluster in cluster j.

[0082] Based on this, the total cluster confidence of the s-th variational autoencoder network is calculated as follows:

[0083]

[0084] Where δ>0 is the adjustment shrinkage coefficient.

[0085] In one embodiment, further regarding step S16 above, the process of training the variational autoencoder network using clustering loss based on the Student t-distribution and KL divergence loss may specifically include the following processing:

[0086] The Student t-distribution is used to measure the similarity between the hidden layer variables and the cluster centers of the variational autoencoder network;

[0087] The clustering loss is trained by using KL divergence loss as the clustering loss between the constructed auxiliary distribution and the soft cluster assignment to be iteratively optimized.

[0088] It is understandable that the training of a variational autoencoder network for deep clustering involves both non-clustering loss and clustering loss training. The non-clustering loss training of this model was introduced above, while the clustering loss training uses the Student's t-distribution to measure the depth representation z. i With cluster center c j similarity q ij q ij Let represent the probability that sample i belongs to cluster j, and it is calculated as follows:

[0089]

[0090] Here, α represents the degrees of freedom. Since cross-validation is not performed on the validation set in an unsupervised environment, learning this parameter is redundant, so we set α = 1. j This represents the j-th cluster center obtained during k-means initialization. Then, an auxiliary distribution P is constructed to help iteratively optimize the soft cluster assignment Q. Here, the KL divergence loss is used as the clustering loss between the soft cluster assignment Q and the auxiliary distribution P, calculated as follows:

[0091]

[0092] The auxiliary distribution P is calculated as follows:

[0093]

[0094]

[0095] In one embodiment, the cluster reliability in each target base cluster is further calculated and measured using the following formula:

[0096]

[0097] Among them, ECE(C i ) represents cluster C i H is an integrated cluster reliability metric. Π (C i ) represents the cluster C in the integrated Πi The uncertainty is denoted by M, which represents the number of base clusters.

[0098] Specifically, to define the reliability of ensemble clusters, we first construct the uncertainty of clusters in the base clusters. Here, we use the concept of entropy. For a given ensemble Π, the base cluster π... m ∈Πcluster C i Uncertainty is defined as:

[0099]

[0100]

[0101] Where, n m It is a base cluster π m The number of clusters, It is a base cluster π m The j-th cluster, |C i | represents cluster C i The number of samples in the middle. Then, the cluster C in the entire ensemble Π. i Uncertainty is defined as:

[0102]

[0103] Thus, cluster C is constructed. i The integrated cluster reliability metric calculation is the aforementioned ECE(C) i As shown in the figure, it is easy to see that the smaller the uncertainty of the cluster, the larger the value of ECE, and ECE(C i )∈(0,1).

[0104] In one embodiment, further regarding the integration consistency function, unlike the globally weighted integration strategy, a locally weighted bipartite graph is constructed here. in represents the point set, Represents the edge set, and the point v in the locally weighted bipartite graph. i and point v j The edge weights between them are calculated as follows:

[0105]

[0106] Among them, ECE(v i ) represents point v i Integrated cluster reliability metric, ECE(v j ) represents point v j Integrated cluster reliability metric Represents the sample set, This represents a cluster. It can be understood that after constructing a locally weighted bipartite graph, the Tcut graph cutting algorithm is applied to segment the bipartite graph to obtain the final ensemble clustering result.

[0107] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order requirement for the execution of these steps; they can be executed in other orders. Figure 1 At least some of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0108] In one embodiment, such as Figure 4 As shown, a deep clustering ensemble device 100 based on cluster confidence is also provided, including a preprocessing module 11, a pretraining module 13, a clustering training module 15, a score calculation module 17, a cluster selection module 19, and an ensemble clustering module 21. The preprocessing module 11 preprocesses and cleans the raw data to be processed, generating an input sample set. The pretraining module 13 uses the input sample set to pretrain a variational autoencoder network and calculates the cluster confidence of the initial low-dimensional embeddings generated during pretraining. The clustering training module 15 trains the variational autoencoder network using clustering loss based on the Student's t-distribution and KL divergence loss, and calculates the cluster confidence of the final low-dimensional embeddings generated after clustering loss training. The score calculation module 17 calculates and sorts the cluster confidence scores of the final low-dimensional embeddings of the variational autoencoder network based on the cluster confidence scores of the initial and final low-dimensional embeddings. Cluster selection module 19 is used to select a predetermined number of target base clusters with high cluster confidence scores from the base clusters generated by the final low-dimensional embeddings. Ensemble clustering module 21 is used to calculate the cluster reliability in each target base cluster using a local weighting strategy, construct a locally weighted bipartite graph, and use the Tcut graph cutting algorithm to segment the locally weighted bipartite graph to obtain the final ensemble clustering result of the input sample set.

[0109] The aforementioned deep clustering ensemble device 100 based on cluster confidence, through the collaboration of various modules, preprocesses and cleans the raw data to be processed to generate an input sample set. Then, it uses the input sample set to pre-train a variational autoencoder network and calculates the cluster confidence of the initial low-dimensional embeddings generated by the pre-training. Next, it trains the clustering loss based on the Student's t-distribution and KL divergence loss, calculates the cluster confidence of the final low-dimensional embeddings generated after the clustering loss training, calculates the cluster confidence score of the final low-dimensional embeddings of the variational autoencoder network, sorts them, selects the target base clusters corresponding to the top set number of final low-dimensional embeddings with high cluster confidence scores as the input for subsequent clustering ensemble, and finally uses a local weighting strategy to calculate the cluster reliability in each target base cluster, constructs a local weighted bipartite graph, and uses the Tcut graph cutting algorithm to segment the local weighted bipartite graph to obtain the final ensemble clustering result.

[0110] Compared with traditional techniques, the above scheme combines cluster confidence assessment and cluster ensemble methods, maps the original unlabeled sample data to low-dimensional (deep) embeddings, constructs cluster confidence assessment to evaluate the quality of low-dimensional deep embeddings and introduces cluster ensemble methods, achieving clustering results with better robustness and clustering performance.

[0111] In one embodiment, the pre-training module 13, during the pre-training of the variational autoencoder network using the input sample set, can specifically be used to set the distribution of the hidden layer variables of the posterior probability of each input sample to follow a normal distribution; in the pre-training process, KL divergence is used to measure the distribution of the hidden layer variables of the posterior probability and the standard normal distribution to determine the non-clustering loss; and the variational autoencoder network is pre-trained using the input sample set based on the non-clustering loss.

[0112] In one embodiment, the pre-training module 13, in the process of calculating the cluster confidence of the initial low-dimensional embedding generated by pre-training, can also be used to calculate the confidence of each cluster of each variational autoencoder in the variational autoencoder network; and calculate the cluster confidence of the initial low-dimensional embedding of each variational autoencoder based on the confidence of each cluster of each variational autoencoder.

[0113] In one embodiment, during the training of clustering loss for the variational autoencoder network based on the Student t-distribution and KL divergence loss, the clustering training module 15 can specifically be used to measure the similarity between the hidden layer variables of the variational autoencoder network and the cluster centers using the Student t-distribution; and to train the clustering loss by using the KL divergence loss as the clustering loss between the constructed auxiliary distribution and the soft cluster assignment to be iteratively optimized.

[0114] In one embodiment, the cluster reliability in each target basis cluster is calculated and measured using the following formula:

[0115]

[0116] Among them, ECE(C i ) represents cluster C i H is an integrated cluster reliability metric. Π (C i ) represents the cluster C in the integrated Π i The uncertainty is that M represents the number of base clusters.

[0117] In one embodiment, point v in a locally weighted bipartite graph i and point v j The edge weights between them are calculated as follows:

[0118]

[0119] Among them, ECE(v i ) represents point v i Integrated cluster reliability metric, ECE(v j ) represents point v j Integrated cluster reliability metric Represents the sample set, Indicates a cluster.

[0120] In one embodiment, the pre-training loss function of the variational autoencoder network is:

[0121]

[0122] Where x represents the input sample, Let σ represent the reconstructed sample, σ represent the standard deviation of the input sample during training, and μ represent the mean of the input sample during training.

[0123] For specific limitations regarding the cluster confidence-based deep clustering ensemble device 100, please refer to the corresponding limitations of the cluster confidence-based deep clustering ensemble method mentioned above, which will not be repeated here. Each module in the aforementioned cluster confidence-based deep clustering ensemble device 100 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of a device with data processing capabilities, or stored in software within the memory of the aforementioned device, so that the processor can call and execute the operations corresponding to each module. The aforementioned device can be, but is not limited to, various types of data processing devices already existing in the art.

[0124] In one embodiment, a computer device is also provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following processing steps: preprocessing and cleaning the raw data to be processed to generate an input sample set; pre-training a variational autoencoder network using the input sample set and calculating the cluster confidence of the initial low-dimensional embeddings generated by the pre-training; training the variational autoencoder network with clustering loss based on the Student's t-distribution and KL divergence loss, and calculating the cluster confidence of the final low-dimensional embeddings generated after the clustering loss training; calculating and sorting the cluster confidence scores of the final low-dimensional embeddings of the variational autoencoder network according to the cluster confidence scores of the initial low-dimensional embeddings and the final low-dimensional embeddings; selecting the target base clusters corresponding to the top set number of final low-dimensional embeddings with high cluster confidence scores from each base cluster generated by the final low-dimensional embeddings; calculating the cluster reliability in each target base cluster using a local weighting strategy, constructing a local weighted bipartite graph and segmenting the local weighted bipartite graph using the Tcut graph cutting algorithm to obtain the final ensemble clustering result of the input sample set.

[0125] It is understood that, in addition to the memory and processor mentioned above, the computer equipment described above also includes other hardware and software components not listed in this specification. The specific components can be determined according to the model of the computer equipment in different application scenarios, and will not be listed and described in detail in this specification.

[0126] In one embodiment, when the processor executes the computer program, it can also implement the steps or sub-steps added in the various embodiments of the deep clustering ensemble method based on cluster confidence.

[0127] In one embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, it performs the following processing steps: preprocessing and cleaning the raw data to be processed to generate an input sample set; pre-training a variational autoencoder network using the input sample set and calculating the cluster confidence of the initial low-dimensional embeddings generated by the pre-training; training the variational autoencoder network with clustering loss based on the Student's t-distribution and KL divergence loss, and calculating the cluster confidence of the final low-dimensional embeddings generated after the clustering loss training; calculating and sorting the cluster confidence scores of the final low-dimensional embeddings of the variational autoencoder network according to the cluster confidence scores of the initial low-dimensional embeddings and the final low-dimensional embeddings; selecting the target base clusters corresponding to the top set number of final low-dimensional embeddings with high cluster confidence scores from each base cluster generated by the final low-dimensional embeddings; calculating the cluster reliability in each target base cluster using a local weighting strategy, constructing a local weighted bipartite graph and segmenting the local weighted bipartite graph using the Tcut graph cutting algorithm to obtain the final ensemble clustering result of the input sample set.

[0128] In one embodiment, when the computer program is executed by a processor, it can also implement the steps or sub-steps added to the various embodiments of the deep clustering ensemble method based on cluster confidence described above.

[0129] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus DRAM (RDRAM), and interface DRAM (DRDRAM), etc.

[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and all of these fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A deep clustering ensemble method based on cluster confidence, characterized in that, Including the following steps: The raw data to be processed is preprocessed and cleaned to generate the input sample set; The variational autoencoder network is pre-trained using the input sample set, and the cluster confidence of the initial low-dimensional embeddings generated by the pre-training is calculated. The variational autoencoder network is trained using clustering loss based on the student t-distribution and KL divergence loss, and the cluster confidence of the final low-dimensional embedding generated after training with clustering loss is calculated. Based on the cluster confidence of the initial low-dimensional embedding and the cluster confidence of the final low-dimensional embedding, the cluster confidence score of the variational autoencoder network is calculated and sorted. Among the base clusters generated by the final low-dimensional embedding, the target base clusters corresponding to the first set number of final low-dimensional embeddings with high cluster confidence scores are selected. The cluster reliability in each target base cluster is calculated using a local weighting strategy. A local weighted bipartite graph is constructed and the Tcut graph cutting algorithm is used to segment the local weighted bipartite graph to obtain the final ensemble clustering result of the input sample set. The process of pre-training the variational autoencoder network using the input sample set includes: The hidden layer variables for the posterior probability of each input sample are assumed to follow a normal distribution. In pre-training, KL divergence is used to measure the distribution of hidden layer variables and the standard normal distribution of posterior probabilities to determine the non-clustering loss; The variational autoencoder network is pre-trained using the input sample set based on the non-clustering loss. The process of calculating the cluster confidence of the initial low-dimensional embeddings generated during pre-training includes: Calculate the confidence score of each cluster of each variational autoencoder in the variational autoencoder network; The cluster confidence of the initial low-dimensional embedding of each variational autoencoder is calculated based on the confidence of each cluster of each variational autoencoder.

2. The deep clustering ensemble method based on cluster confidence according to claim 1, characterized in that, The process of training the variational autoencoder network with clustering loss based on the Student t-distribution and KL divergence loss includes: The Student t-distribution is used to measure the similarity between the hidden layer variables of the variational autoencoder network and the cluster centers; The clustering loss is trained by using KL divergence loss as the clustering loss between the constructed auxiliary distribution and the soft cluster assignment to be iteratively optimized.

3. The deep clustering ensemble method based on cluster confidence according to claim 2, characterized in that, The cluster reliability in each of the target base clusters is calculated and measured using the following formula: Among them, ECE(C i ) represents cluster C i H is an integrated cluster reliability metric. Π (C i ) represents cluster C in the integrated Π i The uncertainty is denoted by M, which represents the number of base clusters.

4. The deep clustering ensemble method based on cluster confidence according to claim 2, characterized in that, Point v in the locally weighted bipartite graph i and point v j The edge weights between them are calculated as follows: Among them, ECE(v i ) represents point v i Integrated cluster reliability metric, ECE(v j ) represents point v j Integrated cluster reliability metric Represents the sample set, Indicates a cluster.

5. The deep clustering ensemble method based on cluster confidence according to claim 2, characterized in that, The pre-training loss function of the variational autoencoder network is: Where x represents the input sample, Let σ represent the reconstructed sample, σ represent the standard deviation of the input sample during training, and μ represent the mean of the input sample during training.

6. A deep clustering ensemble device based on cluster confidence, characterized in that, include: The preprocessing module is used to preprocess and clean the raw data to be processed, and generate the input sample set; The pre-training module is used to pre-train the variational autoencoder network using the input sample set and to calculate the cluster confidence of the initial low-dimensional embeddings generated by the pre-training. The clustering training module is used to train the variational autoencoder network with clustering loss based on the student t-distribution and KL divergence loss, and to calculate the cluster confidence of the final low-dimensional embedding generated after the clustering loss training. The scoring module is used to calculate and sort the cluster confidence scores of the final low-dimensional embeddings of the variational autoencoder network based on the cluster confidence scores of the initial low-dimensional embeddings and the final low-dimensional embeddings. The clustering selection module is used to select, from the base clusters generated by the final low-dimensional embedding, the target base clusters corresponding to the first set number of final low-dimensional embeddings with high cluster confidence scores; An ensemble clustering module is used to calculate the cluster reliability of each target base cluster using a local weighting strategy, construct a local weighted bipartite graph and segment the local weighted bipartite graph using the Tcut graph cutting algorithm to obtain the final ensemble clustering result of the input sample set; The process of pre-training the variational autoencoder network using the input sample set includes: The hidden layer variables for the posterior probability of each input sample are assumed to follow a normal distribution. In pre-training, KL divergence is used to measure the distribution of hidden layer variables and the standard normal distribution of posterior probabilities to determine the non-clustering loss; The variational autoencoder network is pre-trained using the input sample set based on the non-clustering loss. The process of calculating the cluster confidence of the initial low-dimensional embeddings generated during pre-training includes: Calculate the confidence score of each cluster of each variational autoencoder in the variational autoencoder network; The cluster confidence of the initial low-dimensional embedding of each variational autoencoder is calculated based on the confidence of each cluster of each variational autoencoder.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the deep clustering ensemble method based on cluster confidence as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the deep clustering ensemble method based on cluster confidence as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Feature clustering method, database updating method, electronic equipment and storage medium

    CN112560731A

  • Product allocation strategy and system based on deep clustering, storage medium and equipment

    CN115310554A