Multi-modal clustering method and device, equipment, storage medium and product

The unique characterization of multimodal data is extracted by the autoencoder, and a shared characterization is generated by combining adaptive average pooling and expectation maximization algorithms, which solves the problem of low quality of multimodal data clustering labels in the prior art and achieves higher clustering label accuracy.

CN120123792APending Publication Date: 2025-06-10PENG CHENG LAB
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510178959.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has damaged the quality of clustering labels due to extensive fusion strategies and lack of supervision in multimodal data clustering.

Method used

The unique representation of multimodal data is extracted by the autoencoder, and the clustering prototype is obtained using the adaptive average pooling operation and the expectation maximization algorithm, and a shared representation is generated based on the cross attention mechanism. The aligned common representation is obtained through comparative learning, and finally clustered it to obtain high-quality clustering labels.

Benefits of technology

The model data distinction ability is enhanced, the cluster label accuracy is improved, and the cluster label quality is effectively solved in the existing technology due to extensive fusion strategies and lack of supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123792A_ABST
    Figure CN120123792A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing and machine learning, and discloses a multi-modal clustering method, device and equipment, a storage medium and a product, and the method comprises the steps: determining the specific representation of original multi-modal data based on an auto-encoder and reconstruction loss, and carrying out the reconstruction of the original multi-modal data based on the specific representation; obtaining a clustering prototype through an adaptive average pooling operation and an expectation maximization algorithm, generating common characterization of the original multi-modal data according to a cross attention mechanism and the clustering prototype, carrying out comparative learning on the specific characterization and the common characterization to obtain aligned common characterization, and clustering the aligned common characterization to obtain a clustering result; and obtaining a clustering label of the original multi-modal data. According to the method, the unique representation is extracted through the self-encoder, the representation of the clustering prototype is dynamically adjusted by adopting the self-adaptive average pooling operation and the expectation maximization algorithm, and the common representation is generated based on the cross attention mechanism and the clustering prototype, so that the clustering label is obtained, the model data distinguishing capability is enhanced, and the accuracy of the clustering label is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing and machine learning technology, and in particular to a multimodal clustering method, device, equipment, storage medium and product. Background Art

[0002] In the existing technical field, clustering multimodal data through deep learning models is a challenging task. Clustering aims to utilize the similarities between samples, aggregate similar samples and separate dissimilar samples. Clustering tasks in real scenarios usually involve data of different modalities (such as images, text, and speech), and the relationships between different modalities often have strong heterogeneity. However, studies have shown that data of multiple modalities can be sampled from a potential common distribution, and the cross-modal correlation of these data can be characterized by a shared latent semantic space. Therefore, how to effectively eliminate the heterogeneity between data of different modalities and mine their shared latent semantic space is a core issue in multimodal clustering research.

[0003] Existing technical solutions usually learn the potential representations shared by different modalities based on modality or sample level fusion, and there are some overlooked problems: 1) Methods based on modality level, one part simply splices the representations of different modalities, and the other part calculates the weighted sum of the representations of different modalities. The former assumes that the information of each modality is equally important and ignores the difference in data quality between different modalities. The latter adopts the form of weighted sum, and important information in a certain modality may be obscured by the noise information in other modalities. 2) Methods based on sample level assume that the common representation of samples can be improved by other similar samples. However, due to the lack of supervised information and the presence of noise in the data collection process, the similarity relationship between samples is not reliable, and the quality of the common representation will be affected by irrelevant samples. Overall, the existing technical solutions have suffered from the quality of clustering labels due to the rough representation fusion strategy and lack of supervision. Summary of the invention

[0004] The main purpose of this application is to provide a multimodal clustering method, device, equipment, storage medium and product, aiming to solve the technical problem that the quality of clustering labels is impaired due to the rough fusion strategy and lack of supervision in the prior art.

[0005] To achieve the above objectives, the present application proposes a multimodal clustering method, which includes:

[0006] Determine the unique representation of the original multimodal data based on the autoencoder and reconstruction loss;

[0007] Based on the unique representation, a cluster prototype is obtained through an adaptive average pooling operation and an expectation maximization algorithm, and a common representation of the original multimodal data is generated according to a cross attention mechanism and the cluster prototype;

[0008] Perform contrastive learning on the specific representations and the common representations to obtain the aligned common representations, and cluster the aligned common representations to obtain the clustering labels of the original multimodal data.

[0009] In one embodiment, the step of obtaining the common representation of the original multimodal data based on the specific representation by means of an adaptive average pooling operation and an expectation maximization algorithm and generating the common representation according to the cross-attention mechanism and the clustering prototype includes:

[0010] Based on the specific representation, initialize the clustering prototype through adaptive average pooling to obtain the initial clustering prototype;

[0011] Map the initial clustering prototype and the specific representation through a single connection layer to obtain the query matrix of the initial clustering prototype and the key matrix and value matrix of the specific representation;

[0012] Determine the similarity matrix between the initial clustering prototype and the specific representation according to the query matrix, the key matrix, and the value matrix;

[0013] Update the initial clustering prototype based on the expectation maximization algorithm and the similarity matrix to obtain the clustering prototype;

[0014] Determine the assignment matrix of the specific representation, and generate the common representation of the original multimodal data according to the cross-attention mechanism, the clustering prototype, and the assignment matrix.

[0015] In one embodiment, the step of updating the initial clustering prototype based on the expectation maximization algorithm and the similarity matrix to obtain the clustering prototype includes:

[0016] Based on the expectation maximization algorithm and the similarity matrix, partition the specific representation into the corresponding initial clustering prototypes;

[0017] Determine the contribution degree of the specific representation in the corresponding initial clustering prototype, and perform a linear weighted summation on the specific representation based on the contribution degree to obtain the summed specific representation;

[0018] Update the initial clustering prototype according to the summed specific representation and the momentum update strategy to obtain the clustering prototype.

[0019] In one embodiment, the step of determining the assignment matrix of the specific representation and generating the common representation of the original multimodal data according to the cross-attention mechanism, the clustering prototype, and the assignment matrix includes:

[0020] Determine the allocation matrix of the specific features, and based on the cross-attention mechanism, perform a linear weighted sum on the clustering prototypes through the allocation matrix to obtain the cross-attention representation based on the clustering prototypes;

[0021] Measure the quality of the cross-attention representation through the silhouette coefficient, a reference-free evaluation metric;

[0022] Take the specific features as residuals, and perform a weighted sum on the quality and the residuals to obtain the common representation of the original multi-modal data.

[0023] In one embodiment, the step of performing contrastive learning on the specific features and the common representation to obtain the aligned common representation and clustering the aligned common representation to obtain the clustering labels of the original multi-modal data includes:

[0024] Map the specific features and the common representation to a high-order feature space to obtain the same-sample representations in different modalities;

[0025] Perform contrastive learning on the specific features and the common representation according to the same-sample representations to obtain the aligned common representation;

[0026] Cluster the aligned common representation through a clustering analysis algorithm to obtain the clustering labels of the original multi-modal data.

[0027] In one embodiment, the step of performing contrastive learning on the specific features and the common representation according to the same-sample representations to obtain the aligned common representation includes:

[0028] Take the same-sample representations as positive example pairs, and the remaining sample representations as negative example pairs;

[0029] Combine the positive example pairs and the negative example pairs, and perform contrastive learning on the specific features and the common representation to determine the contrastive loss;

[0030] Based on the contrastive loss, iteratively optimize the common representation through the gradient descent algorithm to obtain the aligned common representation.

[0031] In addition, to achieve the above object, the present application also proposes a multi-modal clustering device, and the multi-modal clustering device includes:

[0032] A specific feature determination module, configured to determine the specific features of the original multi-modal data based on an autoencoder and a reconstruction loss;

[0033] A common representation determination module, configured to obtain clustering prototypes based on the specific features through an adaptive average pooling operation and an expectation maximization algorithm, and generate a common representation of the original multi-modal data according to the cross-attention mechanism and the clustering prototypes;

[0034] A characterization clustering module, configured to perform contrastive learning on the specific characterizations and the common characterizations to obtain the aligned common characterizations, and cluster the aligned common characterizations to obtain the clustering labels of the original multimodal data.

[0035] In addition, to achieve the above object, the present application also provides a multimodal clustering device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the multimodal clustering method as described above.

[0036] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the multimodal clustering method as described above.

[0037] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the multimodal clustering method as described above.

[0038] The technical solution proposed by the present application determines the specific characterizations of the original multimodal data based on an autoencoder and a reconstruction loss, obtains clustering prototypes based on the specific characterizations through an adaptive average pooling operation and an expectation maximization algorithm, generates common characterizations of the original multimodal data according to a cross-attention mechanism and the clustering prototypes, performs contrastive learning on the specific characterizations and the common characterizations to obtain the aligned common characterizations, and clusters the aligned common characterizations to obtain the clustering labels of the original multimodal data. By extracting specific characterizations through an autoencoder, dynamically adjusting the representations of the clustering prototypes by using an adaptive average pooling operation and an expectation maximization algorithm, and generating common characterizations based on a cross-attention mechanism and the clustering prototypes, the clustering labels are obtained, enhancing the model's data discrimination ability and improving the accuracy of the clustering labels. Description of the Drawings

[0039] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application and used together with the description to explain the principles of the present application.

[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1Schematic flowchart provided for the first embodiment of the multimodal clustering method of the present application;

[0042] Figure 2 Schematic flowchart provided for the second embodiment of the multimodal clustering method of the present application;

[0043] Figure 3 Schematic diagram of characterization clustering for the multimodal clustering method of the present application;

[0044] Figure 4 Schematic diagram of clustering prototype modeling for the multimodal clustering method of the present application;

[0045] Figure 5 Schematic diagram of feature assignment for the multimodal clustering method of the present application;

[0046] Figure 6 Schematic flowchart provided for the third embodiment of the multimodal clustering method of the present application;

[0047] Figure 7 Schematic diagram of the module structure of the multimodal clustering device according to the embodiment of the present application;

[0048] Figure 8 Schematic diagram of the device structure of the hardware operating environment involved in the multimodal clustering method according to the embodiment of the present application.

[0049] The implementation, functional features, and advantages of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0050] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0051] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and the specific implementation manners.

[0052] Existing technical solutions usually learn the latent representations shared by different modalities based on modality or sample-level fusion, and there are some overlooked problems. For example, it is easy to ignore the data quality differences between different modalities or the important information in a certain modality is easily masked. Or, due to the lack of supervision information and the presence of noise in the data collection process, the similarity relationship between samples is unreliable, and the quality of the common representation will be affected by irrelevant samples, resulting in the deterioration of the quality of clustering labels.

[0053] Therefore, to overcome the above defects, the present application provides a solution. The unique representations are extracted through an autoencoder, the representations of the clustering prototypes are dynamically adjusted by using adaptive average pooling operations and the expectation maximization algorithm, and the common representations are generated based on the cross-attention mechanism and the clustering prototypes, thereby obtaining clustering labels, enhancing the data discrimination ability of the model, and improving the accuracy of clustering labels.

[0054] It should be noted that the execution subject of each embodiment of this application can be a computing service system with data processing, network communication, and program running functions, such as an electronic system, a multimodal clustering system, etc. that can implement the above functions. Hereinafter, taking the multimodal clustering system as an example (hereinafter referred to as the "system"), the following embodiments will be described.

[0055] Based on this, an embodiment of this application provides a multimodal clustering method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the multimodal clustering method of this application.

[0056] In this embodiment, the multimodal clustering method includes steps S10 to S30:

[0057] Step S10, based on the autoencoder and the reconstruction loss, determine the unique representation of the original multimodal data.

[0058] It should be noted that the multimodal clustering method of this application can be applied to, including but not limited to: intelligent security monitoring, e-commerce platform product recommendation, education tutoring, and disease diagnosis and other fields.

[0059] In addition, it should be noted that an autoencoder is an unsupervised learning neural network model that learns a low-dimensional latent space (i.e., the encoding layer) to reconstruct the input data, thereby capturing the internal structure or features of the data. An autoencoder usually consists of two parts: an encoder, which maps the input data to the latent space to generate a unique representation; and a decoder, which maps the representation in the latent space back to the original data space and attempts to reconstruct the input data. In addition, the reconstruction loss is an index to measure the difference between the reconstructed data and the original data. Commonly used loss functions include mean square error (MSE), cross-entropy loss, etc. By minimizing the reconstruction loss, the autoencoder can learn the unique representations that can efficiently reconstruct the input data, and these representations often capture the key features of the data.

[0060] Specifically for multimodal data, this process means that for data from different modalities (such as images, texts, audios, etc.), the autoencoder will separately learn their unique representations, which not only retain the key information of their respective modalities but may also have a certain degree of alignment or correlation in the latent space. Therefore, determining the unique representation of the original multimodal data based on the autoencoder and the reconstruction loss is actually by training the autoencoder network so that it can extract efficient, compact, and information-rich feature representations from the original data with the minimum reconstruction error.

[0061] Step S20: Based on the specific representation, obtain the clustering prototypes through adaptive average pooling operation and the expectation maximization algorithm, and generate the common representation of the original multi-modal data according to the cross-attention mechanism and the clustering prototypes.

[0062] It should be noted that the multi-modal clustering method proposed in this application overcomes the defects of the existing modal- or sample-level multi-modal fusion methods by introducing the intermediate representation of clustering prototypes. The multi-modal clustering method of this application can be roughly divided into two steps: clustering prototype modeling and feature assignment. In clustering prototype modeling, the method updates the clustering prototypes based on the Expectation Maximization (EM) algorithm. The expectation maximization algorithm is mainly divided into two steps: adjusting the model according to the parameters (E-step) and adjusting the parameters according to the model (M-step). That is, in the E-step, the posterior probability distribution between the sample representation and the clustering prototype is calculated, that is, the similarity matrix between the sample and the clustering prototype, and the contribution degree of each sample to the clustering prototype is recorded. In the M-step, the clustering prototype is weighted and updated through the posterior probability, that is, based on the similarity matrix, the clustering prototype is linearly weighted and updated with the sample representations.

[0063] In the feature assignment step, the system enhances the learning of its common representation through the clustering prototypes of the samples. However, considering the quality of the clustering prototype representation in the unsupervised case, a cross-attention mechanism based on clustering prototypes is proposed. That is, first, use the clustering prototypes estimated in prototype modeling to calculate the similarity matrix between the prototypes and the sample representations to obtain the assignment matrix of the samples; subsequently, perform a linear weighted sum of the clustering prototypes according to the assignment matrix to obtain the cross-attention representation based on the clustering prototypes; use the silhouette coefficient, a reference-free evaluation index, to set the weight parameter to measure the information content of the attention representation, and perform a weighted addition with the original input residual to obtain the fused multi-modal common representation.

[0064] In this step, the adaptive average pooling operation downsamples the specific representation in the spatial dimension to extract global features or aggregate information; the expectation maximization algorithm is an iterative algorithm used to estimate the parameters in a statistical model in the presence of latent variables. The specific representation is used to initialize the clustering prototypes, and these prototypes are iteratively updated to maximize the likelihood of the data. In each iteration, the E-step evaluates the probability that each data point belongs to each cluster (i.e., soft assignment), and the M-step updates the clustering prototypes according to these probabilities. This process is iterated until the clustering prototypes converge or reach the preset number of iterations.

[0065] In addition, the cross-attention mechanism is an information fusion technique that allows the model to dynamically focus on different parts of the input data and generate outputs based on the information of these focus points. In this step, the cross-attention mechanism is used to combine the clustering prototypes and the specific representations to generate a common representation of the original multimodal data. Specifically, the model assigns attention weights according to the similarity (or correlation) between the specific representations and the clustering prototypes, and these weights are then used to weighted-sum the clustering prototypes to generate a common representation that fuses information from different modalities.

[0066] In step S30, contrastive learning is performed on the specific representations and the common representations to obtain the aligned common representations, and the aligned common representations are clustered to obtain the clustering labels of the original multimodal data.

[0067] It should be noted that contrastive learning is an unsupervised learning method that learns the feature representations of data by comparing the similarities between different samples, fully exploiting the complementary and consistency information between different modalities under unsupervised conditions, realizing the grouping of different modality data, and greatly reducing the dependence on the annotation scale.

[0068] In this step, contrastive learning is used to align the specific representations and the common representations. Specifically, the model tries to bring closer the specific representations and the common representations from the same original data point (i.e., regarded as a positive example pair), while pushing away the representations from different data points (i.e., regarded as a negative example pair). This process usually involves a contrastive loss function that encourages the representations between positive example pairs to be more similar while making the representations between negative example pairs more different. By minimizing the contrastive loss, the model can learn more robust and consistent feature representations. The aligned common representations are the result of contrastive learning, which represents the alignment state of the specific representations and the common representations in the feature space. These aligned representations not only retain the key information of the original data but also achieve cross-modal consistency and alignment.

[0069] In addition, it should be noted that clustering analysis is also an unsupervised learning method that aims to divide data points into several clusters or categories such that the data points within the same cluster are similar to each other while the data points between different clusters are quite different. In this step, the aligned common representations are used as the input to the clustering algorithm. Commonly used clustering algorithms include K-means, hierarchical clustering, DBSCAN, etc. These algorithms divide the clusters according to the similarity metrics (such as Euclidean distance, cosine similarity, etc.) between data points and generate corresponding clustering labels.

[0070] The finally generated clustering labels represent the classification results of the original multimodal data. These labels not only reflect the similarities and differences between data points, but also provide an important basis for subsequent data analysis and processing. For example, in tasks such as image classification, text clustering, and cross-modal retrieval, the clustering labels can be used to evaluate the performance of the model, guide the subsequent data processing flow, or serve as input features for other machine learning tasks.

[0071] In this embodiment, based on the autoencoder and the reconstruction loss, the unique representation of the original multimodal data is determined. Based on the unique representation, the clustering prototypes are obtained through the adaptive average pooling operation and the expectation maximization algorithm. According to the cross-attention mechanism and the clustering prototypes, the common representation of the original multimodal data is generated. The unique representation and the common representation are subjected to contrastive learning to obtain the aligned common representation, and the aligned common representation is clustered to obtain the clustering labels of the original multimodal data. The unique representation is extracted by the autoencoder, the representation of the clustering prototypes is dynamically adjusted by the adaptive average pooling operation and the expectation maximization algorithm, and the common representation is generated based on the cross-attention mechanism and the clustering prototypes, thereby obtaining the clustering labels, enhancing the model's data discrimination ability, and improving the accuracy of the clustering labels.

[0072] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , the step S20 may include steps S201 to S205:

[0073] Step S201, based on the unique representation, initialize the clustering prototypes through adaptive average pooling to obtain the initial clustering prototypes.

[0074] It can be understood that initializing the clustering prototypes is the process of converting the global feature vectors into clustering center points. Specifically, each global feature vector can be regarded as an initial clustering prototype, representing a potential clustering center. These initial clustering prototypes will be continuously updated and optimized in the subsequent clustering iteration process to better reflect the true distribution of the data.

[0075] Step S202, map the initial clustering prototypes and the unique representation through a single connection layer to obtain the query matrix of the initial clustering prototypes and the key matrix and value matrix of the unique representation.

[0076] It should be noted that mapping the initial clustering prototypes and the unique representation through a single connection layer (also known as a fully connected layer or a linear layer) is a common feature transformation step in deep learning to map the initial clustering prototypes and the unique representation to the same feature space. Specifically, the single connection layer will perform a linear transformation on the input feature vectors, that is, multiply by a weight matrix and add a bias vector.

[0077] When it should be understood, the initial clustering prototype and the specific representation are used as input feature vectors respectively and are mapped through different single connection layers. For the initial clustering prototype, the single connection layer maps it to the query matrix. Each row in the query matrix represents the query vector of a clustering prototype, which is used to calculate the similarity with the key matrix of the specific representation in the subsequent steps. For the specific representation, it is mapped to the key matrix and the value matrix respectively through two different single connection layers. Each row in the key matrix represents the key vector of a specific representation, which is used to match with the query matrix; and each row in the value matrix represents the value vector corresponding to the specific representation, and these value vectors will be weighted and summed according to the attention weights in the subsequent steps.

[0078] Step S203: Determine the similarity matrix between the initial clustering prototype and the specific representation according to the query matrix, the key matrix, and the value matrix.

[0079] Step S204: Update the initial clustering prototype based on the expectation maximization algorithm and the similarity matrix to obtain the clustering prototype.

[0080] It can be understood that the similarity matrix provides the similarity information between the initial clustering prototype and the specific representation, and can be regarded as the basis for evaluating the potential belonging of data points in the E step of the expectation maximization algorithm. Specifically, each element in the similarity matrix represents the similarity between an initial clustering prototype and a specific representation, and this similarity can be interpreted as the probability that the data point belongs to this cluster. Next, enter the M step of the expectation maximization algorithm. Based on the probability information provided by the similarity matrix, the new position of each clustering prototype can be calculated. By alternately executing the E step (evaluating potential belonging using the similarity matrix) and the M step (updating the clustering prototype), the expectation maximization algorithm will iterate continuously until the clustering prototype converges or reaches the preset number of iterations.

[0081] In this embodiment, the above step S204 may include: dividing the specific representation into the corresponding initial clustering prototypes based on the expectation maximization algorithm and the similarity matrix; determining the contribution degree of the specific representation in the corresponding initial clustering prototype, and based on the contribution degree, performing a linear weighted sum on the specific representation to obtain the weighted sum of the specific representation; updating the initial clustering prototype according to the weighted sum of the specific representation and the momentum update strategy to obtain the clustering prototype.

[0082] Specifically, first, using the similarity information in the similarity matrix, each unique feature is assigned to the initial clustering prototype that is most similar to it. Then, the contribution degree of each unique feature to its belonging clustering prototype is determined, which can be achieved through the similarity in the similarity matrix. The greater the similarity, the greater the contribution of the unique feature to the clustering prototype. Based on these contribution degrees (i.e., weights), a linear weighted sum of all unique features is performed to obtain the new position of each clustering prototype. This step can be regarded as an update of the clustering prototype, which takes into account the influence of all data points on the clustering prototype and performs weighted processing according to the similarity. In this way, each clustering prototype will be updated to the weighted average position of all its potential belonging data points.

[0083] Finally, according to the sum of the unique features (actually the new position of the clustering prototype) and the momentum update strategy, the initial clustering prototype is updated. The momentum update strategy is an optimization technique that uses the update direction of the previous iteration and the gradient information of the current iteration to update the parameters, aiming to accelerate convergence and reduce oscillations. In this step, the sum of the unique features can be regarded as the new position of the clustering prototype, and the initial clustering prototype is updated in combination with the momentum term. Through continuous iteration, the clustering prototype can be gradually optimized to better reflect the true distribution of the data, and the finally obtained clustering prototype will be able to more accurately represent the clusters to which each data point belongs.

[0084] Step S205, determine the assignment matrix of the unique features, and generate the common representation of the original multimodal data according to the cross-attention mechanism, the clustering prototype, and the assignment matrix.

[0085] It should be noted that the assignment matrix records the probability or weight of each unique feature being assigned to each clustering prototype, usually obtained by calculating the similarity or distance between the unique feature and the clustering prototype. Each row represents a unique feature, each column represents a clustering prototype, and the elements in the matrix represent the probability or weight of the unique feature belonging to the corresponding clustering prototype.

[0086] It can be understood that the clustering prototype is regarded as a "query", while the unique feature is used as the "key" and "value". By calculating the similarity between the query and the key, an attention weight matrix can be obtained, which reflects the degree of association between the clustering prototype and the unique feature. Then, this weight matrix is used to perform a weighted sum of the unique features to obtain the weighted representation of each clustering prototype. Finally, the common representation of the original multimodal data is generated.

[0087] In this embodiment, the above step S205 may include: determining the allocation matrix of the specific representation, and based on the cross-attention mechanism, linearly weighting and summing the clustering prototypes through the allocation matrix to obtain a cross-attention representation based on the clustering prototypes; measuring the quality of the cross-attention representation through the reference-free evaluation index silhouette coefficient; using the specific representation as a residual, and performing weighted summation on the quality and the residual to obtain the common representation of the original multi-modal data.

[0088] It should be noted that, in order to measure the quality of the cross-attention representation, the reference-free evaluation index silhouette coefficient is introduced. The silhouette coefficient is a commonly used clustering effect evaluation index, which combines information from two aspects: compactness and separation, and is used to evaluate the quality of the clustering result.

[0089] The quality of the cross-attention representation is evaluated by calculating its silhouette coefficient, and the specific representation is used as a residual, and is weighted and summed with the quality of the cross-attention representation to obtain the common representation of the original multi-modal data. Among them, using the specific representation as a residual ensures that the common representation not only contains the information of the clustering prototypes, but also retains the uniqueness and details of the original data.

[0090] For ease of understanding, refer to Figures 3 to 5 for illustration. Figure 3 FIG. is a schematic diagram of the representation clustering of the multi-modal clustering method of the present application. Figure 4 FIG. is a schematic diagram of the clustering prototype modeling of the multi-modal clustering method of the present application. Figure 5 FIG. is a schematic diagram of the feature allocation of the multi-modal clustering method of the present application. The multi-modal clustering method framework can include three parts. The feature extraction module realizes the reconstruction of multi-modal data through an autoencoder and extracts the specific representations of each modality; the feature fusion method based on clustering prototypes (PAM) makes full use of the complementary information between different modalities to learn the common representation of modalities; the alignment module based on contrastive learning further realizes the distribution alignment of the common representation of modalities and the specific representation of modalities.

[0091] Specifically, first, use the encoder and the decoder constituting an autoencoder network to map the original multi-modal data

[0092] to the feature space, which will obtain the specific representation

[0093]

[0094] In the formula, L is the loss function; X v is the input data; is the output data reconstructed by the autoencoder network. In this way, by constraining the consistency between the original input and the output reconstructed by the autoencoder network, the information unique to each modality can be largely retained.

[0095] Step 2. Feature fusion module based on clustering prototypes: The main task of this step is to fuse the unique representations of different modalities and learn the common representations of multiple modalities.

[0096] Prototype modeling: Establish a partitioning clustering method based on the expectation-maximization algorithm, model the process of updating clustering prototypes, and implement it through the cross-attention mechanism. First, concatenate the unique representations of all modalities to obtain Z, initialize the clustering prototype C through adaptive average pooling, and the number of clustering prototypes is predefined as the number of categories K of the dataset. Through a single fully connected layer, map C and Z to the query matrix Q C , the key matrix K Z and the value matrix V Z .

[0097] In the E step, calculate the similarity matrix M C between the sample representation and the clustering prototype, and the samples are partitioned into the corresponding clustering prototypes. In the M step, according to the contribution degree of the sample representation to the clustering prototype, linearly weighted sum the sample representations to update the clustering prototype. Due to inconsistent network updates caused by different batches of training data, the proposed method adopts a momentum update strategy to maintain the consistency of the clustering prototype during the training phase. Specifically, define a memory to store the clustering prototype at the previous moment, and the clustering prototype at the current moment is updated as:

[0098]

[0099] In the formula, C is the clustering prototype; η is the momentum update parameter.

[0100] Feature assignment: After modeling the clustering prototype based on the common representation of modalities, establish a cross-attention mechanism based on the clustering prototype, and further enhance the learning of the common representation of modalities through the semantic information introduced by the clustering prototype. Specifically, first map Z and C to the query matrix Q Z , the key matrix K C and the value matrix V C , calculate the assignment matrix M Z to assign the sample features to the clustering prototypes. Subsequently, according to the assignment matrix M Z , linearly weight the clustering prototypes to obtain the cross-attention representation Z' based on the clustering prototype.

[0101] Since the clustering prototype representation directly affects the quality of the cross-attention representation Z', the silhouette coefficient, a reference-free evaluation metric, is introduced to measure the quality of Z', and the original sample representation Z is used as the residual, and the weighted sum is used to obtain the fused multi-modal common representation. Specifically, let s represent the silhouette coefficient score of Z', and scale it to the interval [0, 1] as the weight coefficient of Z', then we have:

[0102]

[0103] Furthermore, the fused common representation is:

[0104]

[0105] In the formula, α 1 +α 2 = 1.

[0106] Step 3. Contrastive learning-based alignment module: The main task of this step is to align the distributions of the modal common representation and the modal-specific representation using contrastive learning. First, map and to the high-order feature space to obtain and Select the same sample representations under different modalities as positive example pairs, and the rest as negative example pairs for contrastive learning. The contrastive loss used in this step is:

[0107]

[0108] In the formula, g(·,·) is the cosine similarity function; h is the sample augmented data representation; τ is the temperature coefficient.

[0109] The entire training process involves the modal reconstruction loss and the contrastive loss, and is iteratively optimized through gradient descent, and finally outputs the common representation

[0110] Step 4. Clustering: Use the K-means algorithm to cluster the fused common representation to obtain the predicted labels.

[0111] In this embodiment, based on the unique representation, the clustering prototype is initialized by adaptive average pooling to obtain the initial clustering prototype. The initial clustering prototype and the unique representation are mapped through a single connection layer to obtain the query matrix of the initial clustering prototype, and the key matrix and value matrix of the unique representation. According to the query matrix, the key matrix, and the value matrix, the similarity matrix between the initial clustering prototype and the unique representation is determined. Based on the expectation maximization algorithm and the similarity matrix, the initial clustering prototype is updated to obtain the clustering prototype. The assignment matrix of the unique representation is determined, and according to the cross-attention mechanism, the clustering prototype, and the assignment matrix, the common representation of the original multimodal data is generated, thereby effectively integrating the multimodal data features, improving the accuracy of the clustering prototype and the representativeness of the common representation, and enhancing the clustering effect.

[0112] Based on the first embodiment of the present application, in the third embodiment of the present application, the content that is the same as or similar to the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 , step S30 may include steps S301 to S303:

[0113] Step S301, map the unique representation and the common representation to a high-order feature space to obtain the same sample representation in different modalities.

[0114] It can be understood that in order to make the data in different modalities comparable and fusible at the feature level, the unique representation and the common representation are mapped to a high-order feature space to obtain the high-order unified representation of the same sample in different modalities. Specifically, a deep neural network (such as a convolutional neural network CNN, a recurrent neural network RNN, or a variational autoencoder VAE, etc.) or other complex non-linear transformation methods can be used to map the unique representation and the common representation to the high-order feature space respectively. In this process, the data of each modality will pass through the corresponding transformation network to obtain its high-order feature representation. Since these high-order feature representations are obtained in the same feature space, they have the same dimension and semantic meaning and can be directly compared and fused.

[0115] Step S302, perform contrastive learning on the unique representation and the common representation according to the same sample representation to obtain the aligned common representation.

[0116] It can be understood that by using the same sample representation as a bridge, contrastive learning is performed on the unique representation and the common representation, which can optimize and align the common representation to make it consistent with the unique representation on the same sample, thereby obtaining the aligned common representation.

[0117] Specifically, in this embodiment, the above step S302 may include: taking the same sample representation as a positive example pair and the remaining sample representations as a negative example pair; combining the positive example pair and the negative example pair, determining the contrast loss by performing comparative learning on the unique representation and the common representation; based on the contrast loss, iteratively optimizing the common representation through a gradient descent algorithm to obtain the aligned common representation.

[0118] It should be understood that the same sample representations are considered as positive pairs, which represent the common features of the same or similar data points, while the remaining sample representations are considered as negative pairs, which represent the features of different or dissimilar data points. By constructing positive and negative pairs, similar and dissimilar data points can be clearly distinguished. Combining positive and negative pairs, unique representations and common representations are contrastively learned. The core of contrastive learning is to shorten the distance between positive pairs and push the distance between negative pairs. By calculating the similarity or distance between unique representations and common representations, the degree of association between them can be quantified, and the contrastive loss can be determined accordingly. Contrastive loss is a measure of how well a model performs in distinguishing similar and dissimilar data points, reflecting the degree of understanding of the relationship between unique representations and common representations.

[0119] Based on contrast loss, the common representation is iteratively optimized through the gradient descent algorithm, and the common representation is updated using the gradient descent algorithm to more accurately reflect the common features of the data points. Through continuous iterative optimization, the aligned common representation is finally obtained. These representations not only maintain the common features of the data points, but also achieve alignment with the unique representations.

[0120] Step S303: clustering the aligned common representations using a cluster analysis algorithm to obtain cluster labels of the original multimodal data.

[0121] After obtaining the common representations after alignment, clustering analysis algorithms can be further used to cluster these representations to obtain cluster labels for the original multimodal data. Specifically, first, select a suitable clustering analysis algorithm. When selecting a clustering algorithm, factors such as the characteristics of the data, the goal of clustering, and the algorithm's computational efficiency and interpretability need to be considered. Then, the common representations after alignment are used as input data, and the selected clustering analysis algorithm is applied for clustering. The clustering analysis algorithm divides them into different clusters based on the similarity or distance between the representations. In this process, the algorithm continuously adjusts the boundaries and number of clusters to maximize the similarity of data points within the cluster and the difference of data points between clusters, thereby obtaining cluster labels for the original multimodal data.

[0122] It should be noted that the multi-modal clustering method of the present application shows excellent performance in multi-modal clustering on standard multi-modal data sets. Specifically, it will be described with reference to Tables 1 to 3. The performance of the multi-modal clustering method of the present application was tested on six multi-modal data sets, and the detailed information of the data sets is shown in Table 1.

[0123] Table 1 Details of multi-modal data sets

[0124] Datasets Samples Views Clusters Hdigit 10000 2 10 CCV 6773 3 20 Prokaryotic 551 3 4 VOC 5649 2 20 RGB-D 1449 2 13 Wikipedia 2866 2 10

[0125] The clustering accuracy (ACC), normalized mutual information (NMI), and purity (PUR) were used as performance evaluation indicators, and the experimental results are shown in Tables 2 and 3.

[0126] Table 2 Clustering results of the present application on the Hdigit, CCV, and Prokaryotic data sets

[0127]

[0128] Table 3 Clustering results of the present application on the VOC, RGB-D, and Wikipedia data sets

[0129]

[0130] From the performance comparison in Tables 2 and 3, it can be seen that the multi-modal clustering method of the present application has achieved good clustering performance on half of the data sets, far superior to other comparison methods. On the remaining data sets, the method also achieved similar experimental results, which effectively demonstrates the effectiveness of the multi-modal clustering method of the present application.

[0131] In this embodiment, by mapping the unique representation and the common representation to a high-order feature space, the same sample representation in different modalities is obtained. According to the same sample representation, the unique representation and the common representation are compared and learned to obtain the aligned common representation. The aligned common representation is clustered through a clustering analysis algorithm to obtain the clustering labels of the original multi-modal data, thereby enabling the deep alignment of multi-modal data and improving the accuracy and consistency of the clustering results.

[0132] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the multi-modal clustering method of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.

[0133] The present application also provides a multi-modal clustering device. Please refer to Figure 7 , the multi-modal clustering device includes:

[0134] A unique representation determination module, configured to determine the unique representation of the original multi-modal data based on an autoencoder and a reconstruction loss;

[0135] A common feature determination module, configured to obtain a clustering prototype based on the specific features through an adaptive average pooling operation and an expectation maximization algorithm, and generate a common feature of the original multimodal data according to a cross-attention mechanism and the clustering prototype;

[0136] A feature clustering module, configured to perform contrastive learning on the specific features and the common features to obtain an aligned common feature, and cluster the aligned common feature to obtain a clustering label of the original multimodal data.

[0137] The multimodal clustering device provided by this application adopts the multimodal clustering method in the above embodiment, and can solve the technical problem that the quality of clustering labels is damaged in the prior art due to rough fusion strategies and lack of supervision. Compared with the prior art, the beneficial effects of the multimodal clustering device provided by this application are the same as those of the multimodal clustering method provided by the above embodiment, and other technical features in the multimodal clustering device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0138] This application provides a multimodal clustering device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the multimodal clustering method in the first embodiment above.

[0139] Refer to the following Figure 8 , which shows a schematic structural diagram of a multimodal clustering device suitable for implementing the embodiments of this application. The multimodal clustering device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Desctions), PMPs (Portable Media Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The multimodal clustering device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0140] As Figure 8As shown, the multimodal clustering device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the multimodal clustering device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the multimodal clustering device to communicate with other devices wirelessly or wiredly to exchange data. Although the multimodal clustering device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems can be alternatively implemented or had.

[0141] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0142] The multimodal clustering device provided by the present application adopts the multimodal clustering method in the above embodiments, and can solve the technical problem that the quality of clustering labels is damaged in the prior art due to the rough fusion strategy and the lack of supervision. Compared with the prior art, the beneficial effects of the multimodal clustering device provided by the present application are the same as those of the multimodal clustering method provided by the above embodiments, and other technical features in the multimodal clustering device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0143] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0144] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0145] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the multi-modal clustering method in the above embodiments.

[0146] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0147] The above computer-readable storage medium can be included in the multi-modal clustering device; it can also exist alone without being assembled into the multi-modal clustering device.

[0148] The above computer-readable storage medium carries one or more programs, which when executed by a multimodal clustering device, cause the multimodal clustering device to: determine the unique representation of the original multimodal data based on an autoencoder and a reconstruction loss; obtain clustering prototypes through an adaptive average pooling operation and an expectation maximization algorithm based on the unique representation; generate a common representation of the original multimodal data according to a cross-attention mechanism and the clustering prototypes; perform contrastive learning on the unique representation and the common representation to obtain an aligned common representation; and cluster the aligned common representation to obtain the clustering labels of the original multimodal data.

[0149] Computer program code for performing the operations of this application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, execute as a stand-alone software package, execute partly on the user's computer and partly on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0151] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0152] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above multi-modal clustering method, which can solve the technical problem that the quality of clustering labels is damaged due to the extensive fusion strategy and lack of supervision in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the multi-modal clustering method provided by the above embodiments, and will not be elaborated here.

[0153] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the multi-modal clustering method as described above are implemented.

[0154] The computer program product provided by the present application can solve the technical problem that the quality of clustering labels is damaged due to the extensive fusion strategy and lack of supervision in the prior art. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the multi-modal clustering method provided by the above embodiments, and will not be elaborated here.

[0155] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the description and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A multimodal clustering method, characterized in that: The method comprises the following steps: Determine the unique representation of the original multimodal data based on the autoencoder and reconstruction loss; Based on the unique representation, a cluster prototype is obtained through an adaptive average pooling operation and an expectation maximization algorithm, and a common representation of the original multimodal data is generated according to a cross attention mechanism and the cluster prototype; The unique representation and the common representation are compared and learned to obtain an aligned common representation, and the aligned common representation is clustered to obtain cluster labels of the original multimodal data.

2. The multimodal clustering method according to claim 1, characterized in that: The step of obtaining a cluster prototype based on the unique representation through an adaptive average pooling operation and an expectation maximization algorithm, and generating a common representation of the original multimodal data according to a cross attention mechanism and the cluster prototype includes: Based on the unique representation, the cluster prototype is initialized by adaptive average pooling to obtain an initial cluster prototype; Mapping the initial cluster prototype and the unique representation through a single connection layer to obtain a query matrix of the initial cluster prototype and a key matrix and a value matrix of the unique representation; Determining a similarity matrix between the initial cluster prototype and the unique representation according to the query matrix, the key matrix and the value matrix; Based on the expectation maximization algorithm and the similarity matrix, the initial cluster prototype is updated to obtain a cluster prototype; Determine the allocation matrix of the unique representation, and generate a common representation of the original multimodal data based on the cross-attention mechanism, the clustering prototypes and the allocation matrix.

3. The multimodal clustering method according to claim 2, characterized in that: The step of updating the initial cluster prototype based on the expectation maximization algorithm and the similarity matrix to obtain the cluster prototype includes: Based on the expectation maximization algorithm and the similarity matrix, the unique representations are divided into corresponding initial cluster prototypes; Determine the contribution degree of the unique representation to the corresponding initial cluster prototype, and based on the contribution degree, perform linear weighted summation on the unique representation to obtain a summed unique representation; According to the post-summation unique representation and momentum update strategy, the initial cluster prototype is updated to obtain a cluster prototype.

4. The multimodal clustering method according to claim 2, characterized in that: The step of determining the allocation matrix of the unique representation and generating a common representation of the original multimodal data according to the cross-attention mechanism, the clustering prototype and the allocation matrix includes: Determine a distribution matrix of the unique representation, and based on a cross-attention mechanism, perform linear weighted summation on the cluster prototypes through the distribution matrix to obtain a cross-attention representation based on the cluster prototype; The quality of cross-attention representation is measured by the no-reference evaluation metric silhouette coefficient; The unique representation is used as a residual, and a weighted sum is performed on the quality and the residual to obtain a common representation of the original multimodal data.

5. The multimodal clustering method according to any one of claims 1 to 4, characterized in that: The step of performing comparative learning on the unique representation and the common representation to obtain the aligned common representation, and clustering the aligned common representation to obtain cluster labels of the original multimodal data includes: Mapping the unique representation and the common representation to a high-order feature space to obtain the same sample representation under different modalities; Performing comparative learning on the unique representation and the common representation according to the same sample representation to obtain an aligned common representation; The aligned shared representations are clustered using a cluster analysis algorithm to obtain cluster labels for the original multimodal data.

6. The multimodal clustering method according to claim 5, characterized in that: The step of performing comparative learning on the unique representation and the common representation according to the same sample representation to obtain the aligned common representation includes: The same sample representation is used as a positive example pair, and the other sample representations are used as negative example pairs; Combining the positive example pair and the negative example pair, determining a contrast loss by performing contrastive learning on the unique representation and the common representation; Based on the contrast loss, the shared representation is iteratively optimized by a gradient descent algorithm to obtain an aligned shared representation.

7. A multimodal clustering device, characterized in that: The multimodal clustering device comprises: A unique representation determination module, used to determine the unique representation of the original multimodal data based on the autoencoder and the reconstruction loss; A common representation determination module, used to obtain a cluster prototype based on the unique representation through an adaptive average pooling operation and an expectation maximization algorithm, and generate a common representation of the original multimodal data according to a cross attention mechanism and the cluster prototype; The representation clustering module is used to perform comparative learning on the unique representation and the common representation to obtain the aligned common representation, and cluster the aligned common representation to obtain cluster labels of the original multimodal data.

8. A multimodal clustering device, characterized in that: The multimodal clustering device comprises: a memory, a processor, and a multimodal clustering program stored in the memory and executable on the processor, wherein the multimodal clustering program implements the multimodal clustering method as claimed in any one of claims 1 to 6 when executed by the processor.

9. A storage medium, characterized in that: The storage medium stores a multimodal clustering program, and when the multimodal clustering program is executed by the processor, the multimodal clustering method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that The computer program product comprises a multimodal clustering program, and when the multimodal clustering program is executed by a processor, the multimodal clustering method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Acute myelogenous leukemia subtype typing method based on prototype contrast learning

    CN120277548A