Multi-view clustering method, device, equipment and storage medium based on latent variables

By obtaining the latent variable set of the multi-view dataset, mapping it to the feature space using a neural network and performing reconstruction loss and adaptive weight alignment constraints, the problem of neglected independence of latent variables is solved, and a more accurate multi-view clustering effect is achieved.

CN120541549BActive Publication Date: 2025-09-26NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511037374.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-09-26
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

Existing deep multi-view clustering methods ignore the independent characteristics of latent variables, which makes it difficult for the latent representation of multimodal data to accurately capture the intrinsic connections and feature differences between data, reducing the accuracy and reliability of the clustering results.

Method used

By obtaining the latent variable set of the multi-view dataset, mapping it to the feature space using a neural network to construct a reconstruction loss, performing latent variable orthogonalization and modulus constraints, and combining the adaptive weight alignment constraint loss, the global loss function is optimized to improve the independence and discriminability of the latent representation.

Benefits of technology

It improves the clustering performance of multimodal data, can more accurately capture the intrinsic connections of multi-view data, and improve the accuracy and reliability of clustering results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541549B_ABST
    Figure CN120541549B_ABST
Patent Text Reader

Abstract

The present application relates to a multi-view clustering method, apparatus, device and storage medium based on latent variables. The method comprises: obtaining latent variables of multi-view data, obtaining latent representations of samples according to the latent variables; using V A neural network will N The latent representation of each sample is mapped to the feature space of the view data to obtain the reconstructed data representation of each view data. The first loss is constructed based on the distance between each reconstructed data representation and the corresponding view data. The values ​​of the latent variables are determined based on the latent representation. The values ​​of each latent variable are orthogonalized and constrained to have a modulus of 1 to obtain the second loss. The linear kernel matrices of each view data and the corresponding reconstructed data representation are kernel-aligned, and the alignment losses of each view data are weightedly fused to obtain the third loss. A global loss function is derived from these losses. The global loss function is optimized, and the optimized latent representation is used for multi-view clustering. This method can effectively improve the clustering performance of multimodal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a multi-view clustering method, apparatus, device and storage medium based on latent variables. Background Art

[0002] After years of digitalization and informatization, multimodal data is becoming increasingly widespread. For example, in the field of autonomous driving, vehicles need to simultaneously collect data from multiple modalities, such as visual images, lidar data, and GPS, to accurately and in real time judge road conditions and respond accordingly. In the field of advertising recommendations, manufacturers often need to combine information from multiple modalities, such as user profiles, consumption habits, and historical purchase records, to predict user interests and push advertisements. Therefore, how to fuse and analyze the features of these multiple modalities and integrate them into the task scenario to efficiently complete the task has become a key and difficult issue that has attracted much attention.

[0003] Thanks to the rapid development of neural network technology, deep multi-view clustering can leverage the powerful learning capabilities of neural networks to efficiently identify and integrate shared and complementary information across modal data, achieving excellent clustering performance. Existing methods are all based on autoencoder architectures, using encoders to map each modality's data into a separate latent representation, which is then fused and clustered. These two processes typically occur simultaneously, mutually influencing and supporting each other, ultimately learning optimal network parameters and calculating the clustering results for multimodal data.

[0004] However, the latent representation of data is deterministic and unique, while data of different modalities depends on the latent representation of data. Existing deep multi-view clustering methods are all based on the autoencoder architecture, which uses the encoder to map the data of each modality to a latent representation of data. In this process, it is assumed that the latent representation of data depends on the data of each modality, which violates its true relationship. During the mapping process, traditional methods regard latent variables as subordinate entities that are jointly affected by the data of each modality, ignoring the independent characteristics of the latent variables themselves. At the same time, they fail to effectively measure the similarity between the latent variables of different data samples, making it difficult for the generated multimodal data latent representation to accurately capture the intrinsic connections and feature differences between the data, resulting in information loss and feature confusion, and ultimately reducing the accuracy and reliability of the clustering results. Summary of the Invention

[0005] Based on this, it is necessary to provide a multi-view clustering method, device, equipment and storage medium based on latent variables to address the above technical problems.

[0006] A multi-view clustering method based on latent variables, the method comprising:

[0007] Obtain a latent variable set corresponding to a multi-view dataset, and obtain an implicit representation of each sample according to the latent variable set; the multi-view dataset includes N of samples V View data;

[0008] use V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data;

[0009] Determine the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample, orthogonalize the value of each latent variable and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss;

[0010] The linear kernel matrices representing each view data and the corresponding reconstructed data in the multi-view dataset are kernel-aligned, and the inverse numbers of the kernel alignment results of each view data are weighted fused to obtain the adaptive weighted alignment constraint loss.

[0011] Obtaining a global loss function according to the reconstruction loss, the latent variable independence constraint loss, and the adaptive weight alignment constraint loss;

[0012] The global loss function is optimized to obtain an optimized latent representation of each sample, and multi-view clustering is performed based on the optimized latent representation of each sample.

[0013] In one embodiment, constructing the reconstruction loss according to the distance between each reconstructed data representation and the corresponding view data includes: performing weighted averaging of the distance between the reconstructed data representation and the view data of each view data of each sample to obtain the reconstruction loss.

[0014] In one embodiment, the reconstruction loss is:

[0015] ;

[0016] in, is the reconstruction loss, is the number of samples, is the number of view data, For the The feature dimensions of view data, For the The first sample The reconstructed data representation of the view data, For the The first sample View data, is the Euclidean distance.

[0017] In one embodiment, determining the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample includes: determining the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample is:

[0018] ;

[0019] in, For the The hidden variable in The value of the latent variable on the sample, For the The first hidden representation of the sample elements, Indexing operation for vectors.

[0020] In one embodiment, the latent variable independence constraint loss is:

[0021] ;

[0022] in, is the latent variable independence constraint loss, is the number of latent variables, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, is an indicator variable, is the transpose operation.

[0023] In one embodiment, the adaptive weight alignment constraint loss is:

[0024] ;

[0025] ;

[0026] in, is the adaptive weight alignment constraint loss, For the The weight of the view data, For the Alignment loss of the view data.

[0027] In one embodiment, the global loss function is:

[0028] ;

[0029] in, is the global loss function, is the adaptive weight alignment constraint loss, is the latent variable independence constraint loss, and is the loss hyperparameter.

[0030] A multi-view clustering device based on latent variables, comprising:

[0031] The latent representation determination module is used to obtain the latent variable set corresponding to the multi-view dataset and obtain the latent representation of each sample according to the latent variable set; the multi-view dataset includes N of samples V View data;

[0032] Reconstruction loss building block for leveraging V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data;

[0033] An independence constraint loss construction module is used to determine the latent variable value of each latent variable in the latent variable set on the corresponding sample based on the latent representation of each sample, orthogonalize the value of each latent variable and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss;

[0034] An alignment constraint loss construction module is used to perform kernel alignment on the linear kernel matrices representing each view data and the corresponding reconstructed data in a multi-view dataset, and to perform weighted fusion of the opposite numbers of the kernel alignment results of each view data to obtain an adaptive weighted alignment constraint loss.

[0035] A global loss construction module, configured to obtain a global loss function according to the reconstruction loss, the latent variable independence constraint loss, and the adaptive weight alignment constraint loss;

[0036] The multi-view clustering module is used to optimize the global loss function to obtain the optimized latent representation of each sample, and perform multi-view clustering based on the optimized latent representation of each sample.

[0037] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0038] Obtain a latent variable set corresponding to a multi-view dataset, and obtain an implicit representation of each sample according to the latent variable set; the multi-view dataset includesN of samples V View data;

[0039] use V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data;

[0040] Determine the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample, orthogonalize the value of each latent variable and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss;

[0041] The linear kernel matrices representing each view data and the corresponding reconstructed data in the multi-view dataset are kernel-aligned, and the inverse numbers of the kernel alignment results of each view data are weighted fused to obtain the adaptive weighted alignment constraint loss.

[0042] Obtaining a global loss function according to the reconstruction loss, the latent variable independence constraint loss, and the adaptive weight alignment constraint loss;

[0043] The global loss function is optimized to obtain an optimized latent representation of each sample, and multi-view clustering is performed based on the optimized latent representation of each sample.

[0044] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0045] Obtain a latent variable set corresponding to a multi-view dataset, and obtain an implicit representation of each sample according to the latent variable set; the multi-view dataset includes N of samples V View data;

[0046] use V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data;

[0047] Determine the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample, orthogonalize the value of each latent variable and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss;

[0048] The linear kernel matrices representing each view data and the corresponding reconstructed data in the multi-view dataset are kernel-aligned, and the inverse numbers of the kernel alignment results of each view data are weighted fused to obtain the adaptive weighted alignment constraint loss.

[0049] Obtaining a global loss function according to the reconstruction loss, the latent variable independence constraint loss, and the adaptive weight alignment constraint loss;

[0050] The global loss function is optimized to obtain an optimized latent representation of each sample, and multi-view clustering is performed based on the optimized latent representation of each sample.

[0051] The above-mentioned multi-view clustering method, device, equipment and storage medium based on latent variables can mine the potential structure behind the data by obtaining the latent variable set corresponding to the multi-view data set and obtaining the latent representation of each sample, providing essential features for subsequent processing. The latent representation is mapped to the feature space of each view using multiple neural networks to construct a reconstruction loss, which can correct the dependency between multimodal data and latent representation and ensure that the latent representation contains the original data information. The latent variable values ​​are orthogonalized and modulus constrained to obtain an independence constraint loss, which can make the latent variables independent of each other, evenly distribute the clustering information, and improve the quality of the latent representation. The linear kernel matrix kernel of the view data and the reconstructed data is aligned and weighted to form an adaptive weight alignment constraint loss, which can enhance the discriminative ability of the latent representation. Considering the importance of different views, the above losses are combined into a global loss function and optimized. Finally, multi-view clustering is performed using the optimized latent representation to obtain more accurate and reasonable clustering results. The embodiments of the present invention can more accurately capture the intrinsic connection of multi-view data and effectively improve the clustering performance of multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 1 is a flow chart of a multi-view clustering method based on latent variables in one embodiment;

[0053] Figure 2 Schematic diagram of the relationship between data latent variables, latent representations, and multimodal data in one embodiment;

[0054] Figure 3 is a structural block diagram of a multi-view clustering device based on latent variables in one embodiment;

[0055] Figure 4 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0057] In one embodiment, Figure 1 As shown in FIG, a multi-view clustering method based on latent variables is provided, which includes the following steps:

[0058] Step 102: Obtain a latent variable set corresponding to the multi-view dataset, and obtain a latent representation of each sample based on the latent variable set.

[0059] like Figure 2 As shown in the figure, a schematic diagram of the relationship between data latent variables, latent representations and multimodal data is provided, where observation method 1 corresponds to generation , Observation method 2 corresponds to the generation , observation method 3 corresponds to the generation , 、 and For observation data of different modalities, the latent variables used to describe the latent space of a batch of data are fixed, resulting in a unique latent representation of the data, rather than varying with the modality. By observing the data using different observation methods, data of different modalities can be obtained. In short, the latent representation of the data is deterministic and unique, while data of different modalities depends on the latent representation of the data. Therefore, in the method of the present invention, for given multi-view data, the latent variables and latent representation of the data are unique and shared across all modalities.

[0060] Multi-view datasets include N of samples V For example, in image-text multi-view data, a sample may have both a corresponding image and a text describing the image. Here, the image and text are different views.

[0061] It can be understood that by finding the potential influencing factors behind the data and the expression of samples in the latent space, it helps to have a deeper understanding of the intrinsic structure of the data and provide more essential feature information for subsequent processing.

[0062] Step 104, using V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data.

[0063] Multiple neural networks are used to map the latent representation of the data into its feature space, reconstructing multimodal data. Note that neural networks can be of various types, including the common fully connected neural network and deconvolutional neural network. Reconstruction loss measures the difference between the reconstructed data representation and the original view data. The greater the difference, the greater the loss.

[0064] It can be understood that by reconstructing the data and calculating the loss, the model can learn how to accurately restore the original multi-view data from the latent representation, ensure that the latent representation contains enough original data information, correct previous misunderstandings about the relationship between data and latent representation, and allow the model to more accurately grasp the connection between multi-view data.

[0065] Step 106: Determine the latent variable value of each latent variable in the latent variable set on the corresponding sample based on the latent representation of each sample, orthogonalize the latent variable value and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss.

[0066] Orthogonalization means making the values ​​of different latent variables independent and uncorrelated. Latent variable independence constraints are proposed for data latent variables to make them independent of each other, thereby ensuring that data clustering information is distributed as evenly as possible across all latent variables. This prevents a single latent variable from overly dominating the clustering results, helps generate high-quality latent representations of the data, and thus improves the clustering effect of multimodal data.

[0067] Step 108 : performing kernel alignment on the linear kernel matrices representing each view data and the corresponding reconstructed data in the multi-view dataset, and performing weighted fusion on the inverse of the kernel alignment results of each view data to obtain an adaptive weighted alignment constraint loss.

[0068] The linear kernel matrix is ​​a matrix used to describe the similarity relationship between data. The adaptive weight alignment constraint loss is obtained by first inverting the result of the kernel alignment and then performing weighted fusion based on the importance of different views, where the importance is automatically learned by the model during training.

[0069] This method proposes a dynamic weight alignment constraint for data latent representation, taking into account the similarity between different view data and the importance of each view, making the generated data latent representation more discriminative and further improving the clustering performance of multimodal data.

[0070] Step 110 , obtaining a global loss function according to the reconstruction loss, the latent variable independence constraint loss, and the adaptive weight alignment constraint loss.

[0071] In step 112, the global loss function is optimized to obtain an optimized latent representation of each sample, and multi-view clustering is performed based on the optimized latent representation of each sample.

[0072] By continuously adjusting the model's parameters to minimize the global loss function, the global loss function is optimized. Deep learning optimization strategies, such as stochastic gradient descent, are used to optimize the latent representation of the data. After this, k-means clustering is performed on the latent representation to determine the data's classification.

[0073] The optimized latent representation can more accurately reflect the intrinsic characteristics of the data. Clustering based on such latent representation can obtain more reasonable and accurate clustering results.

[0074] In the above-mentioned multi-view clustering method based on latent variables, by obtaining the latent variable set corresponding to the multi-view dataset and obtaining the latent representation of each sample, the potential structure behind the data can be explored, providing essential features for subsequent processing. Multiple neural networks are used to map the latent representation to the feature space of each view to construct a reconstruction loss, which can correct the dependency between multimodal data and the latent representation and ensure that the latent representation contains the original data information. The latent variable values ​​are orthogonalized and modulus constrained to obtain the independence constraint loss, which can make the latent variables independent of each other, evenly distribute the clustering information, and improve the quality of the latent representation. The linear kernel matrix kernel of the view data and the reconstructed data is aligned and weighted to form an adaptive weight alignment constraint loss, which can enhance the discriminative ability of the latent representation. Considering the importance of different views, the above losses are combined into a global loss function and optimized. Finally, multi-view clustering is performed using the optimized latent representation to obtain more accurate and reasonable clustering results. This embodiment of the present invention can more accurately capture the intrinsic connection of multi-view data and effectively improve the clustering performance of multimodal data.

[0075] In one embodiment, constructing the reconstruction loss according to the distance between each reconstructed data representation and the corresponding view data includes: performing weighted averaging on the distance between the reconstructed data representation and the view data of each view data of each sample to obtain the reconstruction loss.

[0076] In this embodiment, the step of obtaining the reconstructed data representation of each view data includes:

[0077] Given multi-view data ,in and Represent the number of data samples and the number of views respectively. The latent variable set of the data can be recorded as (in is the number of latent variables), and The implicit representation of a data sample can be written as:

[0078] ;

[0079] in, For the Data samples in latent variables It is worth noting that the data implicitly represents It is not a constant, but a variable that can be optimized during model training. To ensure that the data implicit representation retains the information of multi-view data more completely, we use A neural network maps it to the feature space of each view data to obtain the reconstructed data representation:

[0080] ;

[0081] in, Indicates the The parameters of a neural network.

[0082] In one embodiment, the reconstruction loss is:

[0083] ;

[0084] in, is the reconstruction loss, is the number of samples, is the number of view data, For the The feature dimensions of view data, For the The first sample The reconstructed data representation of the view data, For the The first sample View data, is the Euclidean distance.

[0085] In one embodiment, determining the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample includes: determining the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample is:

[0086] ;

[0087] in, For the The hidden variable in The value of the latent variable on the sample, For the The first hidden representation of the sample elements, Indexing operation for vectors.

[0088] In this embodiment, for multi-view data, the latent variables used to describe the latent structure of the data are should be independent of each other. However, due to the data distribution corresponding to the latent variables The variables are unknown and difficult to model without sufficient prior knowledge. This paper uses their values ​​on each data sample to approximate the empirical independence between them. The value of is a vector , thus obtaining the above hidden variable value formula.

[0089] In one embodiment, the latent variable independence constraint loss is:

[0090] ;

[0091] in, is the latent variable independence constraint loss, is the number of latent variables, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, is an indicator variable, is the transpose operation.

[0092] In this embodiment, the implicit variables are constrained by the following formula To ensure its independence:

[0093] ;

[0094] In addition, to facilitate subsequent clustering tasks, the present invention sets the weights of all latent variables to be the same and sets their moduli to 1, that is:

[0095] ;

[0096] Since the above two formulas are difficult to guarantee in the optimization process, especially in the gradient descent optimization strategy of neural networks, the present invention intends to approximate them to the above latent variable independence constraint loss formula.

[0097] In one embodiment, the adaptive weight alignment constraint loss is:

[0098] ;

[0099] ;

[0100] in, is the adaptive weight alignment constraint loss, For the The weight of the view data, For the Alignment loss of the view data.

[0101] In this embodiment, given a data sample , the corresponding kernel matrix can be calculated , and its dimension is Similarly, the kernel matrix of the data implicit representation can also be calculated and recorded as In this scenario, kernel alignment can be used to measure the similarity between the two, which can be written as:

[0102] ;

[0103] in, Represents the inner product of two matrices. Although the above formula can effectively fuse the similarity information of multi-view data, the kernel matrix and The calculation and storage of require the square complexity of the sample data volume, which greatly reduces the learning efficiency and increases the consumption of computing memory. Therefore, the present invention reduces the calculation and storage to linear complexity by unifying the kernel matrix in the above formula into a linear kernel matrix, that is:

[0104] ;

[0105] in, Implicit representation The matrix form of , and the size is .at the same time, For data representation The matrix form of , and the size is By bringing the linear kernel matrix into the kernel alignment formula, we can get:

[0106] ;

[0107] It can be seen that the complexity of the above formula is linear. On this basis, The alignment loss of each view can be set to the opposite of the above alignment value, that is:

[0108] ;

[0109] In addition, when integrating the alignment loss of each view, additional parameters are introduced To balance the weight of the view. Therefore, the above adaptive weight alignment loss formula is obtained. It should be noted that the parameter It is not pre-set, but is continuously optimized and adjusted during the model training process to achieve the optimal result.

[0110] In one embodiment, the global loss function is:

[0111] ;

[0112] in, is the global loss function, is the adaptive weight alignment constraint loss, is the latent variable independence constraint loss, and is the loss hyperparameter.

[0113] Compared with existing deep multi-view clustering methods, the present invention proposes that each modality shares a unique data latent representation, and the data of each modality depends on the data latent representation, rather than the latter depending on the former. Therefore, multiple neural networks are used to map the data latent representation to the data feature space, and the data of each modality is reconstructed, correcting the incorrect setting of the dependency relationship between multimodal data and data latent representation in existing methods. The present invention proposes latent variable independence constraints and latent representation adaptive weight alignment constraints, which take into account the mutual independence between latent variables, and also take into account the similarity between the latent representations corresponding to each data sample, thereby improving the generation quality of the latent representation and thus improving the clustering performance of multimodal data. The present invention has been experimentally verified on the public datasets HW and ORL. The comparison algorithms include Multi-level Feature Learning for Contrastive Multi-view Clustering (MFLVC), Global and Cross-view Feature Aggregation for Multi-view Clustering (GCFAGG), Joint Shared-and-Specific Information for Deep Multi-view Clustering (JSSI), Self-weighted Contrastive Fusion for Deep Multi-view Clustering (SCMVC), and Dual Contrastive Multi-view Clustering (DCMVC). The proposed method is denoted as DMCL. The experimental clustering accuracy results are shown in Table 1. As can be seen, the DMCL algorithm achieves the best clustering accuracy, verifying the effectiveness and advanced nature of the proposed method.

[0114] Table 1 Clustering accuracy statistics of various methods

[0115]

[0116] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0117] In one embodiment, Figure 3 As shown, a multi-view clustering device based on latent variables is provided, comprising:

[0118] The latent representation determination module 302 is used to obtain the latent variable set corresponding to the multi-view dataset and obtain the latent representation of each sample according to the latent variable set; the multi-view dataset includes N of samples V View data;

[0119] Reconstruction loss building module 304 is used to utilize V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data;

[0120] An independence constraint loss construction module 306 is used to determine the latent variable value of each latent variable in the latent variable set on the corresponding sample based on the latent representation of each sample, orthogonalize the latent variable value and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss;

[0121] An alignment constraint loss construction module 308 is configured to perform kernel alignment on the linear kernel matrices representing each view data and the corresponding reconstructed data in the multi-view dataset, and perform weighted fusion on the inverse of the kernel alignment results of each view data to obtain an adaptive weighted alignment constraint loss.

[0122] A global loss construction module 310 is used to obtain a global loss function based on the reconstruction loss, the latent variable independence constraint loss and the adaptive weight alignment constraint loss;

[0123] The multi-view clustering module 312 is used to optimize the global loss function to obtain the optimized latent representation of each sample, and perform multi-view clustering based on the optimized latent representation of each sample.

[0124] In one embodiment, constructing the reconstruction loss according to the distance between each reconstructed data representation and the corresponding view data includes: performing weighted averaging of the distance between the reconstructed data representation and the view data of each view data of each sample to obtain the reconstruction loss.

[0125] In one embodiment, the reconstruction loss is:

[0126] ;

[0127] in, is the reconstruction loss, is the number of samples, is the number of view data, For the The feature dimensions of view data, For the The first sample The reconstructed data representation of the view data, For the The first sample View data, is the Euclidean distance.

[0128] In one embodiment, determining the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample includes: determining the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample is:

[0129] ;

[0130] in, For the The hidden variable in The value of the latent variable on the sample, For the The first hidden representation of the sample elements, Indexing operation for vectors.

[0131] In one embodiment, the latent variable independence constraint loss is:

[0132] ;

[0133] in, is the latent variable independence constraint loss, is the number of latent variables, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, is an indicator variable, is the transpose operation.

[0134] In one embodiment, the adaptive weight alignment constraint loss is:

[0135] ;

[0136] ;

[0137] in, is the adaptive weight alignment constraint loss, For the The weight of the view data, For the Alignment loss of the view data.

[0138] In one embodiment, the global loss function is:

[0139] ;

[0140] in, is the global loss function, is the adaptive weight alignment constraint loss, is the latent variable independence constraint loss, and is the loss hyperparameter.

[0141] For the specific definition of the multi-view clustering device based on latent variables, please refer to the definition of the multi-view clustering method based on latent variables above, which will not be repeated here. The various modules in the above-mentioned multi-view clustering device based on latent variables can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0142] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a multi-view clustering method based on latent variables is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0143] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0144] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the method in the above embodiment when executing the computer program.

[0145] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the above embodiment are implemented.

[0146] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0147] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0148] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and such modifications and improvements are intended to fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A multi-view clustering method based on latent variables, characterized in that: The method comprises: Obtain a latent variable set corresponding to a multi-view dataset, and obtain an implicit representation of each sample according to the latent variable set; the multi-view dataset includes N of samples V View data; use V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data; Determine the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample, orthogonalize the value of each latent variable and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss; The linear kernel matrices representing each view data and the corresponding reconstructed data in the multi-view dataset are kernel-aligned, and the inverse numbers of the kernel alignment results of each view data are weighted fused to obtain the adaptive weighted alignment constraint loss. Obtaining a global loss function according to the reconstruction loss, the latent variable independence constraint loss, and the adaptive weight alignment constraint loss; The global loss function is optimized to obtain an optimized latent representation of each sample, and multi-view clustering is performed based on the optimized latent representation of each sample.

2. The method according to claim 1, characterized in that The constructing of the reconstruction loss according to the distance between each reconstructed data representation and the corresponding view data comprises: The reconstruction loss is obtained by taking a weighted average of the reconstructed data representation of each view data of each sample and the distance between the view data.

3. The method according to claim 2, characterized in that The reconstruction loss is: in, is the reconstruction loss, is the number of samples, is the number of view data, For the The feature dimensions of view data, For the The first sample The reconstructed data representation of the view data, For the The first sample View data, is the Euclidean distance.

4. The method according to claim 1, wherein Determining the latent variable value of each latent variable in the latent variable set on the corresponding sample according to the latent representation of each sample includes: According to the latent representation of each sample, the latent variable value of each latent variable in the latent variable set on the corresponding sample is determined as follows: in, For the The hidden variable in The value of the latent variable on the sample, For the The first hidden representation of the sample elements, Indexing operation for vectors.

5. The method according to claim 1, wherein The latent variable independence constraint loss is: in, is the latent variable independence constraint loss, is the number of latent variables, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, For the A vector composed of the values ​​of the latent variables on each sample, is an indicator variable, is the transpose operation.

6. The method according to claim 1, characterized in that The adaptive weight alignment constraint loss is: in, is the adaptive weight alignment constraint loss, For the The weight of the view data, For the Alignment loss of the view data.

7. The method according to claim 1, characterized in that The global loss function is: in, is the global loss function, is the adaptive weight alignment constraint loss, is the latent variable independence constraint loss, and is the loss hyperparameter.

8. A multi-view clustering device based on latent variables, characterized in that: The device comprises: The latent representation determination module is used to obtain the latent variable set corresponding to the multi-view dataset and obtain the latent representation of each sample according to the latent variable set; the multi-view dataset includes N of samples V View data; Reconstruction loss building block for leveraging V The neural networks will N The latent representation of each sample is mapped to the feature space of the corresponding view data to obtain the reconstructed data representation of each view data in the multi-view dataset, and the reconstruction loss is constructed according to the distance between each reconstructed data representation and the corresponding view data; An independence constraint loss construction module is used to determine the latent variable value of each latent variable in the latent variable set on the corresponding sample based on the latent representation of each sample, orthogonalize the value of each latent variable and constrain it with a modulus of 1 to obtain the latent variable independence constraint loss; An alignment constraint loss construction module is used to perform kernel alignment on the linear kernel matrices representing each view data and the corresponding reconstructed data in a multi-view dataset, and to perform weighted fusion of the opposite numbers of the kernel alignment results of each view data to obtain an adaptive weighted alignment constraint loss. A global loss construction module, configured to obtain a global loss function according to the reconstruction loss, the latent variable independence constraint loss, and the adaptive weight alignment constraint loss; The multi-view clustering module is used to optimize the global loss function to obtain the optimized latent representation of each sample, and perform multi-view clustering based on the optimized latent representation of each sample.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-view subspace clustering method based on implicit representation and self-adaptation

    CN109002854A

  • Multi-source heterogeneous medical data multi-view clustering method and device, medium and equipment

    CN116451095A