Three-Layer Weighted Clustering Ensemble Method Based on Tensor Decomposition
Through the three-layer weighted clustering integration method based on tensor decomposition, the problem of insufficient clustering in the existing technology is solved. By constructing a three-layer weighted co-coordinated matrix and tensor decomposition, higher clustering accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202311526928.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-15
AI Technical Summary
In the prior art, the determination of weighted values and the quality of clustering integration are affected by cluster members, resulting in insufficient clustering.
The three-layer weighted clustering integration method based on tensor decomposition is adopted to generate cluster members through a preset clustering algorithm, build a hypergraph adjacency matrix, and perform weight analysis on points, clusters and partitioned three layers, convert them into three-layer weighted co-coordinated matrix, and overlap them into three-dimensional tensors to capture the diversity of data from a global perspective.
The accuracy and robustness of clustering are improved, thereby improving the quality of clustering results.
Smart Images

Figure CN117540236B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of clustering technology, and particularly to a three-layer weighted clustering ensemble method and system based on tensor decomposition. Background Art
[0002] Cluster analysis is one of the research hotspots in the field of machine learning, and is widely used in fields such as data compression, information retrieval, image segmentation, and text clustering. At the same time, it has also received increasing attention in many fields such as biology, geology, geography, and anomaly data detection. It belongs to the category of unsupervised machine learning. Usually, it does not rely on prior knowledge of the data set, but on the similarity measure between data points, samples, or objects. It automatically divides the data set into different groups or clusters, aiming to make the similarity between points within the same cluster as high as possible, while the similarity between points in different clusters as low as possible.
[0003] Clustering ensemble introduces the idea of ensemble learning into the field of cluster analysis, and it has initiated the research on clustering ensemble. This process usually includes two main steps. First, perform multiple clustering operations on the data set to generate multiple different clustering results. This step is called clustering member generation. Then, combine all clustering members into a set, called the clustering collective, and merge them to produce the final clustering result. This step is called clustering ensemble, also known as consensus function design. The advantage of clustering ensemble is that it can obtain a better final result by combining multiple different clustering results. Compared with traditional clustering methods, it has many unique advantages, including the integrated reuse of clustering results, processing data sets with categorical attributes, detection and processing of noise and outliers, and applicability to parallel processing and distributed data sources. Therefore, clustering ensemble has rapidly developed into an important extension of traditional clustering methods, attracting many scholars at home and abroad. In order to improve the accuracy of clustering ensemble, researchers have proposed many methods, such as clustering member selection and weighted clustering ensemble.
[0004] The invention patent with the application number CN202310891062.5 discloses an unsupervised vehicle re-identification method based on adaptive clustering and weighted hard samples. Among them, the method includes: First, calculate the most suitable radius value using the current clustering parameters to improve the robustness of the clustering pseudo-labels to vehicle sample noise; Second, the memory module records all vehicle sample feature vectors, and uses the distance as the basis for weighting the difficulty of vehicle samples to improve the problem that the model pays insufficient attention to hard vehicle samples; Finally, use the weighted hard vehicle samples to train the vehicle re-identification model in combination with the contrast learning method. The above invention can be widely applied to the intelligent video surveillance system in intelligent transportation and intelligent security.
[0005] However, the determination of the weighted values of the above-mentioned existing technologies and the quality of clustering integration are greatly affected by the clustering members. When the weighted values are inaccurate and the clustering integration is inappropriate, the clustering is not precise enough.
[0006] In view of this, there is an urgent need for a three-layer weighted clustering integration method and system based on tensor decomposition to at least solve the above deficiencies. Summary of the Invention
[0007] One of the objectives of the present invention is to provide a three-layer weighted clustering integration method based on tensor decomposition, which introduces a preset clustering algorithm to generate clustering members, constructs a hypergraph adjacency matrix based on the clustering members, and then analyzes and assigns corresponding weights to the three layers of points, clusters, and partitions respectively. The three-layer weighted hypergraph adjacency matrix is converted into a three-layer weighted co-covariance matrix, and the coherent link matrix and the three-layer weighted co-covariance matrix are stacked into a three-dimensional tensor to capture the diversity of data from a global perspective and further improve the accuracy of clustering.
[0008] The three-layer weighted clustering integration method based on tensor decomposition provided by the embodiments of the present invention includes:
[0009] Step 1: Generate clustering members based on a preset clustering algorithm;
[0010] Step 2: Construct a coherent link matrix and a hypergraph adjacency matrix based on the clustering members;
[0011] Step 3: Perform a three-layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members;
[0012] Step 4: Determine a weighted hypergraph adjacency matrix according to the hypergraph adjacency matrix and the target weights, and convert the weighted hypergraph adjacency matrix into a weighted co-covariance matrix;
[0013] Step 5: Stack the coherent link matrix and the weighted co-covariance matrix into a tensor to obtain a refined co-covariance matrix;
[0014] Step 6: Perform a hierarchical clustering algorithm based on the refined co-covariance matrix to obtain a clustering result.
[0015] Preferably, Step 1: Generate clustering members based on a preset clustering algorithm, including:
[0016] Obtain data sample points;
[0017] Obtain the number of members of the clustering members;
[0018] Obtain the number of clusters of each clustering member;
[0019] Run a preset clustering algorithm for random generation according to the number of members, the number of clusters, and the data sample points to obtain clustering members. The preset clustering algorithm includes one or more of the k-means clustering algorithm, the hierarchical clustering algorithm, and the spectral clustering algorithm.
[0020] Preferably, step 3: Perform a three-layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members, including:
[0021] Obtain the first weight of the data sample points in the clustering members. The calculation formula of the first weight is as follows:
[0022]
[0023] where w i is the first weight of the i-th data sample point, a ij is the element value of the covariance matrix of the i-th data sample point, i represents the i-th row of the matrix, j represents the j-th column of the matrix, and n is the number of data sample points;
[0024] Obtain the second weight of the clusters in the clustering members. The calculation formula of the second weight is as follows:
[0025]
[0026]
[0027]
[0028] where ECI(C i ) represents the second weight of the i-th cluster, C i represents the i-th cluster, represents the j-th cluster in the m-th clustering member, H m (C i ) represents the information entropy of the i-th cluster in the clustering member π m , n m is the total number of clusters in the m-th clustering member, is a function to measure and the uncertainty between C i , H Π (C i ) represents the information entropy of the cluster C i in the clustering collective Π relative to the whole Π, M is the integration scale, and θ is a hyperparameter;
[0029] Obtain the third weight of the clustering members. The calculation formula of the third weight is as follows:
[0030]
[0031] where, is the s-th cluster in the m-th cluster member, k m is the total number of clusters in the m-th cluster member, NMI(π m , π q ) represents the similarity between the m-th cluster member and the q-th cluster member;
[0032] Take the first weight, the second weight, and the third weight together as the target weight.
[0033] Preferably, step 4: According to the hypergraph adjacency matrix and the target weight, determine the weighted hypergraph adjacency matrix, and convert the weighted hypergraph adjacency matrix into a weighted co-covariance matrix, including:
[0034] Obtain the integration scale;
[0035] According to the integration scale and the weighted hypergraph adjacency matrix, determine the weighted co-covariance matrix. The conversion model of the weighted co-covariance matrix is as follows:
[0036]
[0037] where WCA is the weighted co-covariance matrix, H is the weighted hypergraph adjacency matrix, T is the matrix transpose operator, and M is the integration scale.
[0038] The three-layer weighted clustering ensemble method based on tensor decomposition provided by the embodiments of the present invention further includes:
[0039] Before determining the weighted co-covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix, conduct a necessity analysis for constructing the weighted co-covariance matrix. If necessary, perform the corresponding construction;
[0040] Conduct a necessity analysis for constructing the weighted co-covariance matrix, including:
[0041] Calculate the first complexity required to determine the weighted co-covariance matrix based on the weighted hypergraph adjacency matrix;
[0042] Calculate the second complexity of running a preset target algorithm based on the weighted hypergraph adjacency matrix;
[0043] If the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is greater than the second complexity, the necessity analysis result for constructing the weighted co-covariance matrix is unnecessary, and determine the clustering result based on the weighted hypergraph adjacency matrix and the preset target algorithm;
[0044] If the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is less than the second complexity, the necessity analysis result for constructing the weighted co-covariance matrix is necessary;
[0045] If the complexity difference is less than the preset complexity threshold, the necessity analysis result for constructing the weighted co-covariance matrix is necessary.
[0046] Preferably, step 5: Stack the coherent link matrix and the weighted covariance matrix into a tensor to obtain a refined covariance matrix, including:
[0047] Obtain the first dimension of the coherent link matrix and the second dimension of the weighted covariance matrix;
[0048] Create a target tensor according to the first dimension and the second dimension;
[0049] Copy the coherent link matrix to the first channel of the target tensor, and at the same time, copy the weighted covariance matrix to the second channel of the target tensor;
[0050] After the copying is completed, use the corresponding target tensor as the refined covariance matrix.
[0051] The three-layer weighted clustering integration method based on tensor decomposition provided by the embodiments of the present invention further includes:
[0052] Step 7: Perform validity analysis on the clustering results, obtain the analysis results, and evaluate the quality of the clustering results according to the analysis results;
[0053] Performing validity analysis on the clustering results and obtaining the analysis results includes:
[0054] Determine the index type of the validity analysis index of the clustering results, and the index type includes one or more of F value, NMI value, ARI, NMI, CH, and Dunn;
[0055] Obtain the first algorithm type set of the clustering algorithm, and at the same time, obtain the first data feature set of the data sample points in the clustering members;
[0056] Based on the historical evaluation records, determine the preset second algorithm type set and the second data feature set corresponding to the index type;
[0057] If the second algorithm type set contains the first algorithm type set and the second data feature set contains the first data feature set, use the validity analysis index corresponding to the index type as the target analysis index;
[0058] Determine the analysis results according to the target analysis index;
[0059] Evaluating the quality of the clustering results according to the analysis results includes:
[0060] According to the analysis results, compare the clustering effects of the control clustering results of different weighting types, and obtain the comparison results. The weighting types include: point weighting, cluster weighting, partition weighting, point-cluster weighting, point-partition weighting, cluster-partition weighting, and point-cluster-partition three-layer weighting;
[0061] Determine the quality of the clustering results according to the comparison results.
[0062] Preferably, obtaining data sample points includes:
[0063] Obtaining the sample point acquisition requirements input manually, and extracting the first sample point requirement features of the sample point acquisition requirements;
[0064] Accessing multiple first target databases;
[0065] Each time of access, obtaining the database content of the first target database being accessed, and extracting the first content features of the database content, where the content features include: data type and data generation time;
[0066] Matching the first sample point requirement features with the first content features, if the match is met, using the corresponding first sample point requirement features as the second sample point requirement features, and using the corresponding first content features as the second content features;
[0067] Calculating the ratio of the number of features of the second sample point requirement features to the number of features of the first sample point requirement features, and associating it with the corresponding first target database;
[0068] Based on a preset content feature - access cost value library, determining the access cost value of each second content feature, accumulating the access cost values to obtain the sum of access cost values, and associating it with the corresponding first target database;
[0069] If the ratio of the number of features associated with the first target database is greater than or equal to a preset ratio threshold, and the sum of the access cost values associated with the first target database is less than or equal to a preset cost value threshold, then using the corresponding first target database as the second target database;
[0070] Sending a data sample point acquisition request to the second target database to obtain data sample points.
[0071] The three - layer weighted clustering integration system based on tensor decomposition provided by the embodiments of the present invention includes:
[0072] A clustering member generation subsystem, configured to generate clustering members based on a preset clustering algorithm;
[0073] A matrix construction subsystem, configured to construct a coherent link matrix and a hypergraph adjacency matrix based on the clustering members;
[0074] A target weight acquisition subsystem, configured to perform three - layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members;
[0075] A weight fusion subsystem, configured to determine a weighted hypergraph adjacency matrix according to the hypergraph adjacency matrix and the target weights, and convert the weighted hypergraph adjacency matrix into a weighted co - covariance matrix;
[0076] The refined co - covariance matrix acquisition subsystem is used to stack the coherent link matrix and the weighted co - covariance matrix into a tensor to obtain the refined co - covariance matrix;
[0077] The clustering result determination subsystem is used to perform a hierarchical clustering algorithm based on the refined co - covariance matrix to obtain the clustering result.
[0078] Preferably, the weight fusion subsystem includes:
[0079] The scale acquisition module is used to acquire the integration scale;
[0080] The weighted co - covariance matrix determination module is used to determine the weighted co - covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix. The conversion model of the weighted co - covariance matrix is as follows:
[0081]
[0082] where WCA is the weighted co - covariance matrix, H is the weighted hypergraph adjacency matrix, T is the matrix transpose operator, and M is the integration scale;
[0083] The three - layer weighted clustering integration system based on tensor decomposition further includes:
[0084] The weighted co - covariance matrix construction necessity analysis module is used to perform a necessity analysis of constructing the weighted co - covariance matrix before determining the weighted co - covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix. If necessary, the corresponding construction is carried out;
[0085] The weighted co - covariance matrix construction necessity analysis module includes:
[0086] The first complexity calculation sub - module is used to calculate the first complexity required to determine the weighted co - covariance matrix based on the weighted hypergraph adjacency matrix;
[0087] The second complexity calculation sub - module is used to calculate the second complexity of running a preset target algorithm based on the weighted hypergraph adjacency matrix;
[0088] The first analysis sub - module is used to determine that the necessity analysis result of constructing the weighted co - covariance matrix is unnecessary if the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is greater than the second complexity, and determine the clustering result based on the weighted hypergraph adjacency matrix and the preset target algorithm;
[0089] The second analysis sub - module is used to determine that the necessity analysis result of constructing the weighted co - covariance matrix is necessary if the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is less than the second complexity;
[0090] The third analysis sub-module is used to determine that the necessity analysis result of constructing the weighted co-covariance matrix is necessary if the complexity difference is less than a preset complexity threshold.
[0091] The beneficial effects of the present invention are as follows:
[0092] The present invention introduces a preset clustering algorithm to generate clustering members, constructs a hypergraph adjacency matrix based on the clustering members, then analyzes and assigns corresponding weights to the three layers of points, clusters, and partitions respectively, converts the three-layer weighted hypergraph adjacency matrix into a three-layer weighted co-covariance matrix, stacks the coherence link matrix and the three-layer weighted co-covariance matrix into a three-dimensional tensor, captures the diversity of data from a global perspective, and further improves the accuracy of clustering.
[0093] Other features and advantages of the present invention will be described in the following specification, and part of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0094] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0095] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the accompanying drawings:
[0096] Figure 1 is a schematic diagram of the three-layer weighted clustering integration method based on tensor decomposition in an embodiment of the present invention;
[0097] Figure 2 is a schematic diagram of the three-layer weighted clustering integration system based on tensor decomposition in an embodiment of the present invention. Detailed Embodiments
[0098] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described here are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0099] The embodiment of the present invention provides a three-layer weighted clustering integration method based on tensor decomposition, as Figure 1 shown, including:
[0100] Step 1: Generate clustering members based on a preset clustering algorithm; where the clustering members are a set of clusters composed of data sample points randomly generated based on the preset clustering algorithm;
[0101] Step 2: Based on the clustering members, construct a coherence link matrix and a hypergraph adjacency matrix; among them, the coherence link matrix is a symmetric matrix, and its elements represent the similarity or correlation between different clustering members; the hypergraph adjacency matrix is: The hypergraph adjacency matrix is a binary matrix used to describe the connection relationship between clustering member nodes. The rows and columns of the matrix correspond to the clustering member nodes respectively, and the elements of the matrix represent the connection situation between the clustering member nodes. For example, if there is a connection relationship between clustering nodes, the value of the corresponding element is 1, otherwise it is 0;
[0102] Step 3: Perform a three-layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members; among them, the target weights are: the importance or weights of different attributes or features of the clustering members determined by the three-layer weighted clustering ensemble algorithm based on points, clusters, and partitions during the clustering process;
[0103] Step 4: According to the hypergraph adjacency matrix and the target weights, determine the weighted hypergraph adjacency matrix, and convert the weighted hypergraph adjacency matrix into a weighted co-covariance matrix; among them, the weighted hypergraph adjacency matrix is: obtained by assigning corresponding weights to the connection relationships in the hypergraph adjacency matrix;
[0104] Step 5: Stack the coherence link matrix and the weighted co-covariance matrix into a tensor to obtain a refined co-covariance matrix; among them, when stacking the coherence link matrix and the weighted co-covariance matrix into a tensor, use the low-rank property of the tensor to transfer the reliable information of the coherence link matrix to the three-layer weighted co-covariance matrix to obtain an enhanced co-covariance matrix, that is, the refined co-covariance matrix. The refined co-covariance matrix is: a matrix that contains the coherence link information between clustering members and the weighted co-covariance relationship between attributes;
[0105] Step 6: Based on the refined co-covariance matrix, execute the hierarchical clustering algorithm to obtain the clustering result. Among them, the hierarchical clustering algorithm is: the average linkage algorithm.
[0106] The working principle and beneficial effects of the above technical solution are:
[0107] This application introduces a preset clustering algorithm to generate clustering members, constructs a hypergraph adjacency matrix based on the clustering members, and then analyzes and assigns corresponding weights to the three layers of points, clusters, and partitions respectively. The three-layer weighted hypergraph adjacency matrix is converted into a three-layer weighted co-covariance matrix, and the coherence link matrix and the three-layer weighted co-covariance matrix are stacked into a three-dimensional tensor to capture the diversity of data from a global perspective and further improve the accuracy of clustering.
[0108] In one embodiment, Step 1: Based on a preset clustering algorithm, generate clustering members, including:
[0109] Obtain data sample points; among them, the data sample points are: the clustering data in the clustering dataset;
[0110] Obtain the number of members of the clustering members; wherein, the number of members is set manually;
[0111] Obtain the number of clusters of each clustering member; wherein, the number of clusters is: the number of subsets obtained by dividing the data sample points corresponding to each clustering member;
[0112] According to the number of members, the number of clusters, and the data sample points, run a preset clustering algorithm for random generation to obtain clustering members. The preset clustering algorithm includes: one or more of the k-means clustering algorithm, the hierarchical clustering algorithm, and the spectral clustering algorithm. Among them, the clustering members are: a set of clusters composed of data sample points randomly generated based on the preset clustering algorithm.
[0113] The working principle and beneficial effects of the above technical solution are:
[0114] This application obtains data sample points, the number of clusters of clustering members, and the number of members of clustering members, and randomly generates clustering members according to the introduced preset clustering algorithm, improving the randomness of the generation of clustering members and the credibility of the clustering results.
[0115] In one embodiment, step 3: perform a three-layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members, including:
[0116] Obtain the first weight of the data sample points in the clustering members. The calculation formula of the first weight is as follows:
[0117]
[0118] where, w′ i is the first weight of the i-th data sample point, a ij is the element value of the covariance matrix of the i-th data sample point, i represents the i-th row of the matrix, j represents the j-th column of the matrix, and n is the number of data sample points;
[0119] Obtain the second weight of the clusters in the clustering members. The calculation formula of the second weight is as follows:
[0120]
[0121]
[0122]
[0123] where, ECI(C i ) represents the second weight of the i-th cluster, C i represents the i-th cluster, represents the j-th cluster in the m-th clustering member, H m (C i ) represents the clustering member πm The information entropy of the i-th cluster in, n m is the total number of clusters in the m-th clustering member, To measure and C i The function of the uncertainty between, H Π (C i ) represents the information entropy of the cluster C in the clustering collective Π relative to the overall Π, M is the integration scale, and θ is the hyperparameter; i m
[0124] Obtain the third weight value of the clustering member, and the calculation formula of the third weight value is as follows:
[0125]
[0126] Among them, is the s-th cluster in the m-th clustering member, k m is the total number of clusters in the m-th clustering member, NMI(π m , π q ) represents the similarity between the m-th clustering member and the q-th clustering member;
[0127] Take the first weight value, the second weight value, and the third weight value as the target weight value together.
[0128] The working principle and beneficial effects of the above technical solution are:
[0129] This application analyzes and assigns corresponding weight values to the point, cluster, and division layers respectively based on the uncertainty of data sample points, the uncertainty of clusters, and the similarity of clustering members, and the determination of the target weight value is more appropriate.
[0130] In one embodiment, step 4: According to the hypergraph adjacency matrix and the target weight value, determine the weighted hypergraph adjacency matrix, and convert the weighted hypergraph adjacency matrix into a weighted co-covariance matrix, including:
[0131] Obtain the integration scale;
[0132] According to the integration scale and the weighted hypergraph adjacency matrix, determine the weighted co-covariance matrix, and the conversion model of the weighted co-covariance matrix is as follows:
[0133]
[0134] Among them, WCA is the weighted co-covariance matrix, H is the weighted hypergraph adjacency matrix, T is the matrix transpose operator, and M is the integration scale.
[0135] The working principle and beneficial effects of the above technical solution are:
[0136] This application is based on a hypergraph adjacency matrix and a target weight. Additionally, an integration scale is introduced. According to the integration scale and the weighted hypergraph adjacency matrix, a weighted co-covariance matrix is determined, and the weighted hypergraph adjacency matrix is converted into a three-layer weighted co-covariance matrix, which focuses on the diversity of clustering members from a global perspective, improving the accuracy and robustness of clustering analysis.
[0137] The embodiment of the present invention provides a three-layer weighted clustering integration method based on tensor decomposition, further including:
[0138] Before determining the weighted co-covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix, an analysis of the necessity of constructing the weighted co-covariance matrix is performed. If necessary, the corresponding construction is carried out;
[0139] Performing an analysis of the necessity of constructing the weighted co-covariance matrix includes:
[0140] Calculating the first complexity required to determine the weighted co-covariance matrix based on the weighted hypergraph adjacency matrix; wherein, the first complexity is: the amount of computation required for clustering operations based on the weighted co-covariance matrix;
[0141] Calculating the second complexity of running a preset target algorithm based on the weighted hypergraph adjacency matrix; wherein, the second complexity is: the amount of computation required for directly using MLRAA, DLRSE, and K-means for clustering operations;
[0142] If the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is greater than the second complexity, the result of the analysis of the necessity of constructing the weighted co-covariance matrix is unnecessary, and the clustering result is determined based on the weighted hypergraph adjacency matrix and the preset target algorithm; wherein, the preset complexity threshold is set manually in advance;
[0143] If the complexity difference between the first complexity and the second complexity is greater than or equal to the preset complexity threshold and the first complexity is less than the second complexity, the result of the analysis of the necessity of constructing the weighted co-covariance matrix is necessary;
[0144] If the complexity difference is less than the preset complexity threshold, the result of the analysis of the necessity of constructing the weighted co-covariance matrix is necessary.
[0145] The working principle and beneficial effects of the above technical solution are:
[0146] Constructing a similarity matrix can obtain better clustering results, but for a large amount of data, the complexity is relatively high. Therefore, this application performs an analysis of the necessity of constructing the weighted co-covariance matrix, calculates the first complexity required to determine the weighted co-covariance matrix based on the weighted hypergraph adjacency matrix and the second complexity of running a preset target algorithm based on the weighted hypergraph adjacency matrix respectively, and determines different clustering situations according to the complexity difference between the first complexity and the second complexity and the size comparison result, improving the rationality of clustering.
[0147] In one embodiment, step 5: Stack the coherent link matrix and the weighted co-covariance matrix into a tensor to obtain a refined co-covariance matrix, including:
[0148] Obtain the first dimension of the coherent link matrix and the second dimension of the weighted co-covariance matrix; wherein, the first dimension is, for example: (N, N); the first dimension is, for example: (N, N);
[0149] Create a target tensor according to the first dimension and the second dimension; wherein, the target tensor is, for example: (N, N, 2);
[0150] Copy the coherent link matrix to the first channel of the target tensor, and at the same time, copy the weighted co-covariance matrix to the second channel of the target tensor; wherein, when copying, the Python language and a tensor library (such as TensorFlow or PyTorch) can be used to operate to copy the matrix to the specified channel of the tensor.
[0151] After the copying is completed, use the corresponding target tensor as the refined co-covariance matrix.
[0152] The working principle and beneficial effects of the above technical solution are as follows:
[0153] This application utilizes the low rank property of the tensor to propagate the information of the coherent link matrix to the weighted co-covariance matrix. The coherent link matrix and the weighted co-covariance matrix respectively capture different association and correlation information. By stacking them into a tensor, these two types of information can be comprehensively utilized, thereby providing a more comprehensive and rich refined co-covariance matrix. The refined co-covariance matrix provides a fine-grained association analysis. The coherent link matrix describes the connection relationship between nodes, while the weighted co-covariance matrix measures the correlation strength between nodes. By combining them into a tensor, the connection and correlation between nodes can be considered simultaneously, thus better revealing the potential association between nodes, which helps to deeply understand the network structure, node association and potential complex relationships.
[0154] The embodiment of the present invention provides a three-layer weighted clustering ensemble method based on tensor decomposition, further including:
[0155] Step 7: Perform validity analysis on the clustering result, obtain the analysis result, and evaluate the quality of the clustering result according to the analysis result;
[0156] Performing validity analysis on the clustering result and obtaining the analysis result includes:
[0157] Determine the index type of the validity analysis index of the clustering result, and the index type includes one or more of F value, NMI value, ARI, NMI, CH, and Dunn;
[0158] Obtain the first set of algorithm types of the clustering algorithm. At the same time, obtain the first set of data features of the data sample points in the clustering members; wherein, the first set of algorithm types is the set of clustering algorithms included in the preset clustering algorithm; the first set of data features is the data features of the data sample points, and the data features are data type, data length, etc.;
[0159] Based on the historical evaluation records, determine the second set of algorithm types and the second set of data features corresponding to the preset index type; wherein, the historical evaluation records are the process records of evaluating the clustering results in history; wherein, the second set of algorithm types includes the set of clustering algorithms suitable for evaluating the index corresponding to the index type; the second set of data features is the set of data features of the data suitable for evaluating the index corresponding to the index type;
[0160] If the second set of algorithm types contains the first set of algorithm types and the second set of data features contains the first set of data features, use the effectiveness analysis index corresponding to the index type as the target analysis index;
[0161] According to the target analysis index, determine the analysis result; wherein, when determining the analysis result, it can be determined according to the preset analysis template corresponding to the target analysis index. The analysis template restricts the analysis of the clustering result according to the target analysis index and does not perform other contents;
[0162] According to the analysis result, evaluate the quality of the clustering result, including:
[0163] According to the analysis result, compare the clustering effects of the control clustering results of different weighting types. The weighting types include point weighting, cluster weighting, partition weighting, point-cluster weighting, point-partition weighting, cluster-partition weighting, and three-layer point-cluster-partition weighting; wherein, the control clustering result is the clustering result of the clustering methods of different weighting types;
[0164] According to the comparison result, determine the quality of the clustering result. Among them, when determining the quality of the clustering result according to the comparison result, compare the clustering situations of different control clustering results for the same clustering target, and determine the quality of the clustering result according to the ranking of the clustering situation of the clustering result in the clustering situations of the control clustering results. The higher the ranking, the higher the quality of the corresponding clustering result. For example, the clustering result with a ranking of 1 has the highest quality.
[0165] The working principle and beneficial effects of the above technical solution are:
[0166] After determining the clustering results in this application, an effectiveness analysis index of the clustering results is introduced to analyze the effectiveness of the clustering results. When determining the analysis index applicable to this solution in the effectiveness analysis index, according to the first algorithm type set, the first data feature set, the second algorithm type set, and the second data feature set, the target analysis index is determined, which improves the accuracy of the target analysis index. Different weighted types of control clustering results are introduced to determine the comparison results for evaluating the quality of the clustering results, and the evaluation process is more appropriate.
[0167] In one embodiment, obtaining data sample points includes:
[0168] Obtaining the sample point acquisition requirements input manually and extracting the first sample point requirement features of the sample point acquisition requirements; wherein, the first sample point requirement features are: the characteristic representation of the sample point acquisition requirements, such as: what data types and within what time range, etc.;
[0169] Accessing multiple first target databases; wherein, the first target databases are: data sample point storage databases that can be accessed based on big data;
[0170] Each time of access, obtaining the database content of the first target database being accessed and extracting the first content features of the database content, and the content features include: data types and data generation times;
[0171] Matching the first sample point requirement features and the first content features. If the match is met, taking the corresponding first sample point requirement features as the second sample point requirement features and taking the corresponding first content features as the second content features;
[0172] Calculating the ratio of the number of features between the second sample point requirement features and the first sample point requirement features and associating it with the corresponding first target database; wherein, the ratio of the number of features is: the result obtained by dividing the number of the second sample point requirement features by the number of the first sample point requirement features;
[0173] Based on a preset content feature - access cost value library, determining the access cost value of each second content feature, accumulating the access cost values to obtain the sum of the access cost values, and associating it with the corresponding first target database; wherein, the preset content feature - access cost value library is set manually in advance and includes the corresponding relationships between multiple preset content features and access cost values;
[0174] If the ratio of the number of features associated with the first target database is greater than or equal to a preset ratio threshold, and the sum of the access cost values associated with the first target database is less than or equal to a preset cost threshold, then taking the corresponding first target database as the second target database; wherein, the ratio threshold and the cost threshold are set manually in advance;
[0175] Send a data sample point acquisition request to the second target database to obtain data sample points.
[0176] The working principle and beneficial effects of the above technical solution are as follows:
[0177] When obtaining data sample points for clustering integration experiments, in order to improve the richness of data sample points, the database of the public platform can be accessed for acquisition. However, there are access costs for the database of the public platform, and not all of them meet the access requirements. Therefore, this application extracts the first sample point requirement characteristics of the manually input sample point acquisition requirements. At the same time, multiple first target databases are accessed, and the database content of the first target database being accessed is obtained to determine the first content characteristics. The first sample point requirement characteristics and the first content characteristics are matched to determine the matching second sample point requirement characteristics and second content characteristics. Determine the sum of the feature number ratio and access cost value associated with the first target database. When the feature number ratio is greater than or equal to the ratio threshold and the sum of the access cost values is less than or equal to the preset cost value threshold, a data sample point acquisition request is sent to the corresponding second target database, improving the acquisition efficiency of data sample points.
[0178] In one embodiment, it further includes:
[0179] Step 8: If the quality of the clustering result meets the preset quality standard, obtain the patient data of the target patient, and use the patient data as a clustering member for three-layer weighted clustering integration based on tensor decomposition to obtain patient classification. Among them, the preset quality standard is set manually in advance, and the target patient is: the patient of the hospital using the three-layer weighted clustering integration platform based on tensor decomposition; the patient data is: the pathological data of the target patient; the patient classification is: the result of classifying target patients with similar symptoms into one category.
[0180] The working principle and beneficial effects of the above technical solution are as follows:
[0181] This application inputs the patient data into the three-layer weighted clustering integration platform based on tensor decomposition to determine the patient classification. The three-layer weighted clustering integration platform based on tensor decomposition is a medical platform constructed according to the three-layer weighted clustering integration method based on tensor decomposition whose clustering result quality meets the quality standard. Therefore, the suitability of patient classification can be improved, and the medical management efficiency can be enhanced.
[0182] The embodiment of the present invention provides a three-layer weighted clustering integration system based on tensor decomposition, as Figure 2 shown, including:
[0183] A clustering member generation subsystem 1, configured to generate clustering members based on a preset clustering algorithm;
[0184] A matrix construction subsystem 2, configured to construct a coherent link matrix and a hypergraph adjacency matrix based on the clustering members;
[0185] The target weight acquisition subsystem 3 is used to perform a three-layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members;
[0186] The weight fusion subsystem 4 is used to determine the weighted hypergraph adjacency matrix according to the hypergraph adjacency matrix and the target weights, and convert the weighted hypergraph adjacency matrix into a weighted co-covariance matrix;
[0187] The refined co-covariance matrix acquisition subsystem 5 is used to stack the coherence link matrix and the weighted co-covariance matrix into a tensor to obtain the refined co-covariance matrix;
[0188] The clustering result determination subsystem 6 is used to perform a hierarchical clustering algorithm based on the refined co-covariance matrix to obtain the clustering result.
[0189] In one embodiment, the weight fusion subsystem includes:
[0190] The scale acquisition module is used to acquire the integration scale;
[0191] The weighted co-covariance matrix determination module is used to determine the weighted co-covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix. The conversion model of the weighted co-covariance matrix is as follows:
[0192]
[0193] where WCA is the weighted co-covariance matrix, H is the weighted hypergraph adjacency matrix, T is the matrix transpose operator, and M is the integration scale;
[0194] The three-layer weighted clustering integration system based on tensor decomposition further includes:
[0195] The weighted co-covariance matrix construction necessity analysis module is used to perform a necessity analysis on the construction of the weighted co-covariance matrix before determining the weighted co-covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix. If necessary, the corresponding construction is performed;
[0196] The weighted co-covariance matrix construction necessity analysis module includes:
[0197] The first complexity calculation sub-module is used to calculate the first complexity required to determine the weighted co-covariance matrix based on the weighted hypergraph adjacency matrix;
[0198] The second complexity calculation sub-module is used to calculate the second complexity of running a preset target algorithm based on the weighted hypergraph adjacency matrix;
[0199] The first analysis sub-module is used to determine that the necessity analysis result of constructing the weighted co-covariance matrix is unnecessary if the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is greater than the second complexity, and determine the clustering result based on the weighted hypergraph adjacency matrix and a preset target algorithm;
[0200] The second analysis sub-module is used to determine that the necessity analysis result of constructing the weighted co-covariance matrix is necessary if the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is less than the second complexity;
[0201] The third analysis sub-module is used to determine that the necessity analysis result of constructing the weighted co-covariance matrix is necessary if the complexity difference is less than a preset complexity threshold.
[0202] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. Three - layer weighted clustering ensemble method based on tensor decomposition, characterized in that, it includes: Step 1: Generate clustering members based on a preset clustering algorithm; Step 2: Construct a coherence link matrix and a hypergraph adjacency matrix based on the clustering members; Step 3: Conduct a three - layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members; Step 4: Determine a weighted hypergraph adjacency matrix according to the hypergraph adjacency matrix and the target weights, and convert the weighted hypergraph adjacency matrix into a weighted co - covariance matrix; Step 5: Stack the coherence link matrix and the weighted co - covariance matrix into a tensor to obtain a refined co - covariance matrix; Step 6: Based on the refined co - covariance matrix, execute a hierarchical clustering algorithm to obtain a clustering result; Step 7: Conduct an effectiveness analysis on the clustering result to obtain an analysis result, and evaluate the quality of the clustering result according to the analysis result; Step 8: If the quality of the clustering result meets the preset quality standard, obtain the patient data of the target patient, use the patient data as clustering members for three - layer weighted clustering ensemble based on tensor decomposition to obtain patient classification; the patient data is: the pathological data of the target patient.
2. The three - layer weighted clustering ensemble method based on tensor decomposition according to claim 1, characterized in that, Step 1: Generate clustering members based on a preset clustering algorithm, including: Obtain data sample points; Obtain the number of members of the clustering members; Obtain the number of clusters of each clustering member; According to the number of members, the number of clusters and the data sample points, run a preset clustering algorithm for random generation to obtain clustering members, and the preset clustering algorithm includes: one or more of the k - means clustering algorithm, the hierarchical clustering algorithm and the spectral clustering algorithm.
3. The three - layer weighted clustering ensemble method based on tensor decomposition according to claim 1, characterized in that, Step 3: Conduct a three - layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members, including: Obtain the first weight of the data sample points in the clustering members, and the calculation formula of the first weight is as follows: where, w′ i is the first weight of the i-th data sample point, a ij is the element value of the covariance matrix of the i-th data sample point, i represents the i-th row of the matrix, j represents the j-th column of the matrix, and n is the number of data sample points; Obtain the second weight of the clusters in the clustering members, and the calculation formula of the second weight is as follows: Among them, ECI(C i ) represents the second weight of the i-th cluster, C i represents the i-th cluster, represents the j-th cluster among the m-th clustering members, H m (C i ) represents the information entropy of the i-th cluster in the clustering member π m , n m is the total number of clusters in the m-th clustering member, is a function for measuring and the uncertainty between C i , H Π (C i ) represents the information entropy of the cluster C i in the clustering collective Π relative to the overall Π, M is the integration scale, and θ is a hyperparameter; Obtain the third weight of the clustering members, and the calculation formula of the third weight is as follows: Among them, is the s-th cluster in the m-th clustering member, and k m is the total number of clusters in the m-th clustering member, and NMI(π m , π q ) represents the similarity between the m-th clustering member and the q-th clustering member; Take the first weight, the second weight and the third weight together as the target weights.
4. The three - layer weighted clustering ensemble method based on tensor decomposition according to claim 1, characterized in that, Step 4: Determine a weighted hypergraph adjacency matrix according to the hypergraph adjacency matrix and the target weights, and convert the weighted hypergraph adjacency matrix into a weighted co - covariance matrix, including: Obtain the integration scale; Determine the weighted co - covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix, and the conversion model of the weighted co - covariance matrix is as follows: where WCA is the weighted co - covariance matrix, H is the weighted hypergraph adjacency matrix, T is the matrix transpose operator, and M is the integration scale.
5. The three - layer weighted clustering ensemble method based on tensor decomposition according to claim 4, characterized in that, it further includes: Before determining the weighted co - covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix, conduct a necessity analysis for constructing the weighted co - covariance matrix. If necessary, conduct the corresponding construction; Conducting a necessity analysis for constructing the weighted co - covariance matrix includes: Calculate the first complexity required to determine the weighted co-covariance matrix based on the weighted hypergraph adjacency matrix; Calculate the second complexity of running a preset target algorithm based on the weighted hypergraph adjacency matrix; If the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is greater than the second complexity, the analysis result of the necessity of constructing the weighted co-covariance matrix is unnecessary, and based on the weighted hypergraph adjacency matrix and the preset target algorithm, determine the clustering result; If the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is less than the second complexity, the analysis result of the necessity of constructing the weighted co-covariance matrix is necessary; If the complexity difference is less than the preset complexity threshold, the analysis result of the necessity of constructing the weighted co-covariance matrix is necessary.
6. The three-layer weighted clustering integration method based on tensor decomposition according to claim 1, characterized in that Step 5: Stack the coherent link matrix and the weighted co-covariance matrix into a tensor to obtain a refined co-covariance matrix, including: Obtain the first dimension of the coherent link matrix and the second dimension of the weighted co-covariance matrix; Create a target tensor according to the first dimension and the second dimension; Copy the coherent link matrix to the first channel of the target tensor, and at the same time, copy the weighted co-covariance matrix to the second channel of the target tensor; After the copying is completed, use the corresponding target tensor as the refined co-covariance matrix.
7. The three-layer weighted clustering integration method based on tensor decomposition according to claim 1, characterized in that Evaluate the quality of the clustering result according to the analysis result, including: According to the analysis result, compare the clustering effects of the control clustering results of different weighted types, and obtain a comparison result. The weighted types include: point weighting, cluster weighting, partition weighting, point-cluster weighting, point-partition weighting, cluster-partition weighting, and point-cluster-partition three-layer weighting; Determine the quality of the clustering result according to the comparison result.
8. The three-layer weighted clustering integration method based on tensor decomposition according to claim 2, characterized in that Obtain data sample points, including: Obtain the sample point acquisition requirements input manually and extract the first sample point requirement feature of the sample point acquisition requirements; Access multiple first target databases; Each time of access, obtain the database content of the first target database being accessed and extract the first content feature of the database content. The content features include: data type and data generation time; Match the first sample point requirement feature and the first content feature. If the match is met, use the corresponding first sample point requirement feature as the second sample point requirement feature and the corresponding first content feature as the second content feature; Calculate the ratio of the number of features of the second sample point requirement feature to the first sample point requirement feature and associate it with the corresponding first target database; Based on a preset content feature - access cost value library, determine the access cost value of each second content feature, accumulate the access cost values, obtain the sum of the access cost values, and associate it with the corresponding first target database; If the ratio of the number of features associated with the first target database is greater than or equal to a preset ratio threshold, and the sum of the access cost values associated with the first target database is less than or equal to a preset cost value threshold, then the corresponding first target database is taken as the second target database; Send a data sample point acquisition request to the second target database to obtain data sample points.
9. A three-layer weighted clustering ensemble system based on tensor decomposition, characterized in that, comprising: A clustering member generation subsystem for generating clustering members based on a preset clustering algorithm; A matrix construction subsystem for constructing a coherence link matrix and a hypergraph adjacency matrix based on the clustering members; A target weight acquisition subsystem for performing a three-layer weighted weight analysis on the clustering members to obtain the target weights of the clustering members; A weight fusion subsystem for determining a weighted hypergraph adjacency matrix according to the hypergraph adjacency matrix and the target weights, and converting the weighted hypergraph adjacency matrix into a weighted co-covariance matrix; A refined co-covariance matrix acquisition subsystem for stacking the coherence link matrix and the weighted co-covariance matrix into a tensor to obtain a refined co-covariance matrix; A clustering result determination subsystem for performing a hierarchical clustering algorithm based on the refined co-covariance matrix to obtain a clustering result; The three-layer weighted clustering ensemble system based on tensor decomposition also performs the following operations: Perform an effectiveness analysis on the clustering result to obtain an analysis result, and evaluate the quality of the clustering result according to the analysis result; If the quality of the clustering result meets a preset quality standard, obtain the patient data of the target patient, and use the patient data as clustering members for three-layer weighted clustering integration based on tensor decomposition to obtain a patient classification; the patient data is: the pathological data of the target patient.
10. The three-layer weighted clustering ensemble system based on tensor decomposition according to claim 9, characterized in that, The weight fusion subsystem includes: A scale acquisition module for acquiring the integration scale; A weighted co-covariance matrix determination module for determining a weighted co-covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix, and the conversion model of the weighted co-covariance matrix is as follows: where WCA is the weighted co-covariance matrix, H is the weighted hypergraph adjacency matrix, T is the matrix transpose operator, and M is the integration scale; The three-layer weighted clustering ensemble system based on tensor decomposition further includes: A weighted co-covariance matrix construction necessity analysis module for performing a weighted co-covariance matrix construction necessity analysis before determining the weighted co-covariance matrix according to the integration scale and the weighted hypergraph adjacency matrix, and if necessary, performing corresponding construction; The weighted co-covariance matrix construction necessity analysis module includes: A first complexity calculation sub-module for calculating the first complexity required to determine the weighted co-covariance matrix based on the weighted hypergraph adjacency matrix; A second complexity calculation sub-module for calculating the second complexity of running a preset target algorithm based on the weighted hypergraph adjacency matrix; A first analysis sub-module for determining that the weighted co-covariance matrix construction necessity analysis result is unnecessary if the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is greater than the second complexity, and determining the clustering result based on the weighted hypergraph adjacency matrix and the preset target algorithm; The second analysis sub-module is used to determine that the necessity analysis result of weighted co-covariance matrix construction is necessary if the complexity difference between the first complexity and the second complexity is greater than or equal to a preset complexity threshold and the first complexity is less than the second complexity; The third analysis sub-module is used to determine that the necessity analysis result of weighted co-covariance matrix construction is necessary if the complexity difference is less than the preset complexity threshold.
Citation Information
Patent Citations
Unsupervised vehicle re-identification method based on adaptive clustering and hard sample weighting
CN116612445B
Subjectterm and descriptor prediction and ordering method based on diagram
CN106682095A
Tensor decomposition-based on-line explosive topic early-discovery method
CN107133219A