A multi-view clustering method and device based on hierarchical hint guided alignment

By constructing a main layer structure and an embedded layer structure through a hierarchical prompt-guided alignment multi-view clustering method, and combining attention mechanism and pseudo-label generation, the challenges of fine-grained semantic mining within views and global consistency modeling in multi-view clustering are solved, thereby improving clustering accuracy and stability.

CN120976590BActive Publication Date: 2026-03-24BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing multi-view clustering technologies struggle to effectively balance fine-grained semantic mining within views with global consistency modeling between views, resulting in performance limitations in complex data environments. In particular, in smart manufacturing scenarios, the identification of abnormal clusters in equipment health evolution is lagging or has a high false alarm rate.

Method used

A hierarchical prompt-guided alignment-based multi-view clustering method is adopted. By combining the main layer structure, the embedding layer structure and the label layer, a semantic feature representation of multimodal data is constructed. Fine-grained semantic enhancement and cross-view semantic alignment are performed through attention mechanism and pseudo-label generation mechanism. The clustering results are optimized by contrastive learning and transition mapping modules.

Benefits of technology

It improves clustering accuracy and stability, effectively solves the problems of fine-grained semantic mining within views and global consistency modeling between views, and improves the accuracy and consistency of multi-view clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976590B_ABST
    Figure CN120976590B_ABST
Patent Text Reader

Abstract

The application provides a multi-view clustering method and device based on hierarchical prompt guidance alignment, acquires multi-view data to be clustered for a target scene, constructs a main layer structure to generate preliminary semantic feature representation; then constructs an embedded layer structure, obtains local and global high-level semantic representation through a local and global prompt generation network and a double mapping encoder, introduces an attention mechanism to fuse local high-level semantic representation of each mode to obtain fused semantic representation; then constructs a label layer, generates multiple pseudo labels by using a pseudo label generation module, selects a maximum value to obtain a global clustering label; finally, a joint loss is constructed by using contrast learning and the like to update parameters of related structures and modules, so as to generate a clustering result. Through hierarchical prompt guidance and multi-level alignment operation, the method fully excavates effective information of each mode of multi-view data and performs deep fusion, effectively improves the precision and robustness of multi-view clustering, and makes the clustering result more accurately reflect the real semantic structure of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-view data processing technology, and in particular to a multi-view clustering method and apparatus based on hierarchical prompting and alignment. Background Technology

[0002] While multi-view clustering technology is widely used in image retrieval, social network analysis, and other fields, it has significant limitations. Traditional methods, such as those based on subspace assumptions or tensor decomposition models, often rely on linear assumptions or tensor decomposition models when integrating information from different views. This makes them ineffective at handling high-dimensional and nonlinear features, resulting in performance limitations in complex data environments. For example, in smart manufacturing scenarios, different types of sensor data exhibit significant differences, making it difficult for traditional methods to simultaneously characterize the fine-grained dynamics of equipment health evolution and cross-view data. Figure 1 This inconsistency can lead to delayed identification of abnormal clusters or a high false alarm rate.

[0003] Furthermore, existing technologies struggle to achieve effective cross-modal alignment regarding distribution differences at the feature layer and clustering result levels. Some methods focus on global distribution alignment or subspace consistency between views, neglecting the balance between "fine-grained semantics" within views and "local differences" between views, thus failing to fully exploit the unique information of each view. Regarding cue learning, while it has demonstrated powerful guidance capabilities in natural language processing and computer vision, its application in multi-view clustering tasks remains in its early stages, particularly lacking a systematic approach to design hierarchical cue structures to capture fine-grained semantics and shared features.

[0004] Existing technologies still struggle to balance global consistency and local differences between views. Some methods focus only on global alignment, which can easily lead to the loss of detailed view information; while others, although introducing local feature enhancements, fail to effectively establish global consistency between views, affecting the stability and accuracy of clustering results. Therefore, a new approach is urgently needed to handle multi-view clustering tasks. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a multi-view clustering method and apparatus based on hierarchical prompting and alignment, so as to eliminate or improve one or more defects existing in the prior art, and solve the problem that the prior art cannot effectively take into account both fine-grained semantic mining within the view and global consistency modeling between views in multi-view clustering.

[0006] On the one hand, the present invention provides a multi-view clustering method based on hierarchical prompt guidance alignment, the method comprising the following steps:

[0007] For a target scenario, obtain multi-view data to be clustered. The multi-view data contains multiple samples, and each sample has multiple modal views.

[0008] A main layer structure is constructed. A main layer cue encoder is used to take the sample as input and output the main layer cue vector of each modality. The main layer cue vector of each modality is added to the original data and input into the shallow encoder to obtain the preliminary semantic feature representation of each modality.

[0009] An embedding layer structure is constructed. For each modality, a local prompt generation network is established, taking the preliminary semantic feature representation of the corresponding modality as input and outputting local prompts. A global prompt generation network is established, taking the preliminary semantic feature representations fused based on adaptive weight parameters as input and outputting global prompts. The preliminary semantic feature representation, local prompts, and global prompts corresponding to each modality in the sample are fused and projected onto the semantic clustering space through a first mapping encoder to obtain the corresponding local high-level semantic representation. The global prompts are projected onto the semantic clustering space through a second mapping encoder to obtain the corresponding global high-level semantic representation. An attention mechanism is introduced to fuse the local high-level semantic representations of each modality in the sample to obtain a fused semantic representation.

[0010] A label layer is constructed by inputting the local high-level semantic representation of each modality of the sample into the representation-level pseudo-label generation module to obtain the local pseudo-labels corresponding to each modality, and inputting the fused semantic representation into the representation-level pseudo-label generation module to obtain the fused pseudo-labels; the global high-level semantic representation is input into the prompt-level pseudo-label generation module to obtain the global pseudo-labels; for each classification category, the maximum value among the local pseudo-labels, the fused pseudo-labels, and the global pseudo-labels is selected and normalized to obtain the global clustering label;

[0011] A contrastive learning method is used to align the embedding layer structure and the label layer to establish a contrastive loss. The probability distribution difference between the global clustering label and the fused pseudo-label is calculated, and the information entropy of the local pseudo-label and the global pseudo-label is suppressed to construct a forward alignment loss. A transition mapping module is used to perform transition mapping on the local pseudo-label and the global pseudo-label, and the probability distribution difference between them and the original global pseudo-label and global clustering label is calculated to construct a back feedback loss. Based on the contrastive loss, the forward alignment loss, and the back feedback loss, a joint loss is constructed to update the parameters of the embedding layer structure, the label layer, and the transition mapping module. Clustering results are generated based on the updated global clustering label.

[0012] In some embodiments, the main layer cue encoder is a multilayer perceptron encoder, and the shallow layer encoder is a multilayer perceptron encoder or a graph convolution encoder; the representation-level pseudo-label generation module and the cue-level pseudo-label generation module are nonlinear support vector machines with the same structure.

[0013] In some embodiments, a global prompt generation network is established to take the preliminary semantic features fused based on adaptive weight parameters as input and output global prompts, and the calculation formula is as follows:

[0014] ;

[0015] in, This indicates the global suggestion. This refers to the global suggestion generation network. This represents the adaptive weight parameter for the v-th modality data of the sample. The preliminary semantic feature representation of the v-th modality data of the sample.

[0016] In some embodiments, the preliminary semantic feature representation, the local cue, and the global cue corresponding to each modality in the sample are fused and projected onto the semantic clustering space through a first mapping encoder to obtain the corresponding local high-level semantic representation, as calculated below:

[0017] ;

[0018] in, This represents the local high-level semantic representation corresponding to the v-th modality data of the sample. This refers to the first mapping encoder. This represents the v-th modal data of the sample. These are preset parameters. This indicates the local cue corresponding to the vth modal data of the sample.

[0019] In some embodiments, the local high-level semantic representations of each modality of the sample are fused using an attention mechanism to obtain a fused semantic representation, including:

[0020] The weight coefficients are calculated based on attention mechanism networks and nonlinear feedforward neural networks, using the following formula:

[0021] ;

[0022] Based on the weighting coefficients, the local high-level semantic representations corresponding to v modal data in the sample are fused, and the calculation formula is as follows:

[0023] ;

[0024] in, This represents element-wise multiplication. This represents the weight matrix of the v-th modality data in the sample. This represents the attention mechanism network. This refers to the nonlinear feedforward neural network. This represents the fused semantic representation.

[0025] In some embodiments, contrastive learning is used to align the embedding layer structure and the label layer to establish a contrastive loss, calculated as follows:

[0026] ;

[0027] in, This represents the contrast loss. The local high-level semantic representation corresponding to the k-th modality data of the i-th sample is represented by the following. The modality category number is represented by m and n, which represent the sample number. This represents the label matrix for each sample in the k-th modality, where m and n represent the class numbers; N represents the number of samples, and C represents the number of classes. This represents the cosine similarity function.

[0028] In some embodiments, the global clustering label is obtained by normalizing the maximum value among the local pseudo-label, the fused pseudo-label, and the global pseudo-label for each classification category, including:

[0029] For the i-th sample with respect to classification m, the maximum value of the label is selected from the local pseudo-label, the fused pseudo-label, and the global pseudo-label, expressed as:

[0030] ;

[0031] in, Let represent the global pseudo-label of the i-th sample with respect to classification m. Let represent the local pseudo-label of the v-th modality in the i-th sample with respect to classification m. The fused pseudo-label of the i-th sample with respect to classification m;

[0032] The global cluster label for a single sample is calculated using the following formula:

[0033] ;

[0034] in, This represents the value corresponding to category p in the global clustering label. This represents the maximum value of the label with respect to category p. Let C represent the maximum value of the label for category q, where C is the number of categories;

[0035] The formula for calculating the forward alignment loss is:

[0036] ;

[0037] in, This represents the forward alignment loss. This represents the local pseudo-label matrix of each sample with respect to mode v. Represents the global pseudo-label matrix. Represents the global clustering label matrix. This represents the fused pseudo-label matrix; Indicates KL divergence; This represents information entropy.

[0038] In some embodiments, a transition mapping module is used to perform transition mapping on the local pseudo-tags and the global pseudo-tags, as expressed in the following expression:

[0039] ;

[0040] ;

[0041] in, Represents the ReLU activation function. This indicates the origin from the global pseudo-label matrix. The intermediate representation obtained from learning. This indicates the fusion pseudo-label matrix The intermediate representation obtained from learning; , , and This is the weight matrix; , , and It is a fixed value;

[0042] The formula for calculating the reverse feedback loss is:

[0043] ;

[0044] in, This represents the reverse feedback loss;

[0045] The formula for calculating the joint loss is:

[0046] ;

[0047] in, For the aforementioned joint loss, and For weights.

[0048] On the other hand, the present invention also provides a multi-view clustering method apparatus based on hierarchical prompt guidance alignment, including a processor, a memory, and a computer program / instructions stored in the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the apparatus implements the steps of the above method.

[0049] On the other hand, the present invention also provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0050] The multi-view clustering method and apparatus based on hierarchical cue-guided alignment described in this invention constructs a main layer structure and an embedding layer structure to extract semantic features from multimodal data, and optimizes the clustering results through a label layer and a bidirectional alignment mechanism. The main layer cue encoder generates a main-level cue vector for each modality and adds it to the original data. This vector is then passed through a shallow encoder to obtain a preliminary semantic feature representation, which enhances the original input data and preserves and mines the unique semantic information of each view. The local cue generation network and the global cue generation network in the embedding layer structure output local and global cuees, respectively. After fusion, these cuees are projected onto the semantic clustering space through a mapping encoder to obtain local and global high-level semantic representations. An attention mechanism is then used to fuse the local high-level semantic representations of each modality, achieving fine-grained semantic enhancement and cross-view semantic alignment optimization. The label layer generates local pseudo-labels, fused pseudo-labels, and global pseudo-labels for each modality. The global clustering label is obtained by normalizing the maximum value, strengthening the most discriminative signal and weakening interference from low-confidence views. A contrastive loss is established for the embedding layer structure and the label layer through contrastive learning. A forward alignment loss is constructed by combining the probability distribution difference between global clustered labels and fused pseudo-labels. At the same time, a back feedback loss is constructed using the transition mapping module. Finally, the parameters are updated based on the joint loss to generate clustering results, taking into account both fine-grained semantic mining within the view and global consistency between views, thereby improving the clustering accuracy and stability.

[0051] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.

[0052] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description

[0053] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, are not intended to limit the scope of the invention. In the drawings:

[0054] Figure 1 This is a flowchart illustrating a multi-view clustering method based on hierarchical prompting and alignment according to an embodiment of the present invention.

[0055] Figure 2 This is a schematic diagram of the logical structure of a multi-view clustering method based on hierarchical prompt guidance alignment according to another embodiment of the present invention.

[0056] Figure 3 This is a schematic diagram illustrating the sensitivity of different values ​​of hyperparameters β and γ in the joint loss based on Caltech7 data according to an embodiment of the present invention.

[0057] Figure 4 This is a schematic diagram illustrating the sensitivity of different values ​​of hyperparameters β and γ in the joint loss based on Citeseer data according to an embodiment of the present invention.

[0058] Figure 5 This is a schematic diagram of the original features of the BBCSport dataset.

[0059] Figure 6 The HiPMVC method used in an embodiment of the present invention Figure 5 A diagram illustrating the clustering results of the BBCSport dataset.

[0060] Figure 7 This is a schematic diagram of the original features of the Caltech7 dataset.

[0061] Figure 8 The HiPMVC method used in an embodiment of the present invention Figure 7 A schematic diagram illustrating the clustering results of the Caltech7 dataset. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention. It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only structures and / or processing steps closely related to the solutions according to the invention are shown in the accompanying drawings, while other details not closely related to the invention are omitted.

[0063] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0064] In large-scale, multi-source data environments, the complementarity of information from different modalities or feature views provides rich support for clustering tasks. Multi-View Clustering (MVC) has therefore gained widespread attention in fields such as image retrieval, social network analysis, and bioinformatics. Traditional MVC methods are mostly based on subspace assumptions or tensor decomposition, seeking shared subspaces or low-rank tensor representations in the feature spaces of each view to achieve collaborative learning among views. Although these methods can integrate complementary information from different views to some extent, they often rely on linear assumptions or tensor decomposition models, which limits their performance when dealing with high-dimensional, nonlinear features.

[0065] Multi-view clustering is an unsupervised learning method for handling multi-source, multi-modal data. In many real-world scenarios, data objects can be described from multiple perspectives or modalities. These perspectives may be different types of data sources (such as images, text, audio, etc.) or different feature representations of the same data. The goal of multi-view clustering is to integrate this data from different perspectives, uncover the inherent structure and patterns of the data, and group similar data objects into the same cluster. Traditional clustering methods typically assume that data samples have only one feature representation, while multi-view clustering considers the multimodality of the data and can comprehensively utilize information from multiple perspectives to obtain more comprehensive and accurate data partitioning results. This application introduces cue learning, a learning paradigm that guides the model to complete downstream tasks. By injecting learnable cue vectors into the input, it activates the existing knowledge in the pre-trained model, reduces the amount of parameter updates, and has been widely applied in fields such as natural language processing and computer vision. Its core idea is to introduce additional information at the input end, guiding the model to focus on task-related features and achieving efficient representation learning.

[0066] This invention provides a multi-view clustering method based on hierarchical prompt guidance alignment, such as... Figure 1 As shown, the method includes the following steps S101~S105:

[0067] Step S101: Obtain multi-view data to be clustered for the target scenario. The multi-view data contains multiple samples, and the samples exist in multiple modal views.

[0068] Step S102: Construct the main layer structure, use the main layer cue encoder to take the sample as input and output the main layer cue vector of each modality, add the main layer cue vector of each modality to the original data and input it into the shallow encoder to obtain the preliminary semantic feature representation of each modality.

[0069] Step S103: Construct an embedding layer structure. For each modality, establish a local prompt generation network, taking the preliminary semantic feature representation of the corresponding modality as input and outputting local prompts. Establish a global prompt generation network, taking the preliminary semantic feature representations fused based on adaptive weight parameters as input and outputting global prompts. After fusing the preliminary semantic feature representation, local prompts, and global prompts corresponding to each modality in the sample, project them onto the semantic clustering space through a first mapping encoder to obtain the corresponding local high-level semantic representation. Project the global prompts onto the semantic clustering space through a second mapping encoder to obtain the corresponding global high-level semantic representation. Introduce an attention mechanism to fuse the local high-level semantic representations of each modality in the sample to obtain the fused semantic representation.

[0070] Step S104: Construct a label layer. Input the local high-level semantic representation of each modality of the sample into the representation-level pseudo-label generation module to obtain the local pseudo-labels corresponding to each modality. Input the fused semantic representation into the representation-level pseudo-label generation module to obtain the fused pseudo-labels. Input the global high-level semantic representation into the prompt-level pseudo-label generation module to obtain the global pseudo-labels. For each classification category, select the maximum value among the local pseudo-labels, fused pseudo-labels, and global pseudo-labels and normalize them to obtain the global clustering label.

[0071] Step S105: Use contrastive learning to align the embedding layer structure and the label layer to establish a contrastive loss; calculate the difference in distribution probability between the global cluster label and the fused pseudo-label and suppress the information entropy of the local pseudo-label and the global pseudo-label to construct a forward alignment loss; use the transition mapping module to perform transition mapping on the local pseudo-label and the global pseudo-label, and calculate the difference in distribution probability with the original global pseudo-label and the global cluster label to construct a back feedback loss; construct a joint loss based on the contrastive loss, the forward alignment loss and the back feedback loss to update the parameters of the embedding layer structure, the label layer and the transition mapping module, and generate clustering results based on the updated global cluster label.

[0072] In step S101, the multi-view data to be clustered for the target scene includes multiple samples, and each sample has multiple modal views. For example, in a smart manufacturing scenario, various heterogeneous sensors deployed on the production line, such as vibration, temperature, current, and acoustic sensors, will generate high-speed streaming data with different sampling frequencies. These different modal data constitute multi-view data. Multi-view data can be represented as... ,in, Indicates the number of samples. Indicates the number of modes. For the first Data dimensions of each modality.

[0073] In step S102, a master layer structure is constructed. A master layer cue encoder takes each modality data of each sample as input and outputs a master layer cue vector for each modality. The master layer cue vector for each modality is added to the original data and input into the shallow encoder to obtain the preliminary semantic feature representation of each modality. This process generates cue vectors for each modality's data through the master layer cue encoder, aiming to enhance the original input data and preserve and mine the unique semantic information of each view.

[0074] The master-level cue encoder generates modality-adaptive master-level cue vectors for each modality view in the sample. By adding the cue vectors element-wise to the original data, the shallow encoder is forced to prioritize local features strongly correlated with the clustering task, suppressing irrelevant noise. Compared to traditional subspace methods that force views to share the same linear space, this application preserves the physical characteristics of each view before feature mapping. For example, in bearing monitoring, the high-frequency resonance component (5-10kHz) of the vibration signal and the slow variation trend of the temperature signal are enhanced separately, avoiding the blurring of heterogeneous features in early fusion. The output features serve as input for subsequent processing, and their cluster separability is significantly improved through cue guidance.

[0075] In some embodiments, the main layer prompt encoder is a multilayer perceptron encoder, and the shallow layer encoder is a multilayer perceptron encoder or a graph convolution encoder.

[0076] In some embodiments, a global prompt generation network is established to take the preliminary semantic features fused based on adaptive weight parameters as input and output global prompts, and the calculation formula is as follows:

[0077] ;

[0078] in, This indicates a global suggestion. This indicates that a global suggestion is generated for the network. This represents the adaptive weight parameters for the v-th modality of the sample. This represents the preliminary semantic feature representation of the v-th modality data of the sample.

[0079] In step S103, a local-global dual-cue collaboration mechanism is used to achieve fine-grained view semantics and cross-view functionality during the feature fusion stage. Figure 1 The dynamic balancing of consistency fundamentally solves the problem of complementary information loss caused by coarse-grained fusion in traditional methods. Specifically, fine-grained local semantic enhancement is performed, and the local cue generation network extracts view-specific clustering cues from the preliminary semantic feature representations of each view. Features are projected into a more discriminative subspace through weighted fusion and a mapping encoder. Cross-view semantic alignment optimization is performed, and global cues dynamically learn the correlation patterns between views by adaptively fusion of multi-view features.

[0080] In some embodiments, the preliminary semantic feature representation, local cue, and global cue corresponding to each modality in the sample are fused and projected onto the semantic clustering space through a first mapping encoder to obtain the corresponding local high-level semantic representation, as calculated below:

[0081] ;

[0082] in, This represents the local high-level semantic representation corresponding to the v-th modality data of the sample. Indicates the first mapping encoder. This represents the v-th modal data of the sample. These are preset parameters. This indicates the local cue corresponding to the vth modal data of the sample.

[0083] In some embodiments, the local high-level semantic representations of each modality of the sample are fused using an attention mechanism to obtain a fused semantic representation, including:

[0084] The weight coefficients are calculated based on attention mechanism networks and nonlinear feedforward neural networks, using the following formula:

[0085] ;

[0086] The local high-level semantic representations corresponding to v modal data in the sample are fused based on weighting coefficients, and the calculation formula is as follows:

[0087] ;

[0088] in, This represents the weight matrix for the v-th modality in the sample. This represents an attention mechanism network. This represents a nonlinear feedforward neural network. This represents the fused semantic representation.

[0089] In step S104, this step achieves fine-grained view cues and cross-view functionality at the label level through three-layer pseudo-label generation and maximum value normalization aggregation. Figure 1 The deep integration of consistency-based decision-making effectively solves the misjudgment problem caused by label ambiguity or view conflict in traditional methods. Specifically, hierarchical labels exhibit semantic complementarity, local pseudo-labels mine fine-grained features within views, fused pseudo-labels integrate multiple views through an attention mechanism, and global pseudo-labels are generated based on shared semantics across views. By taking the maximum value of local / fused / global labels and normalizing it to generate global cluster labels, the most discriminative signals are strengthened, while interference from low-confidence views is weakened. In some embodiments, the representation-level pseudo-label generation module and the prompt-level pseudo-label generation module employ a nonlinear support vector machine with the same structure.

[0090] In step S105, this step achieves cross-view semantic alignment and distribution balance simultaneously at the feature layer and label layer through contrastive learning bidirectional loss constraints, completely resolving the clustering performance degradation problem caused by view distribution offset or training instability in traditional methods. Multi-level alignment is performed, and contrastive loss forces the feature layer and pseudo-label layer to cross views. Figure 1 Consistency is achieved. Forward loss strengthens decision confidence, KL divergence constraints minimize the distribution difference between global cluster labels and fused pseudo-labels, and information entropy penalties are added to suppress low-confidence predictions. Backward feedback stabilization is implemented, and smooth intermediate labels are generated through a transition mapping module to construct backward loss constraints. Finally, a joint loss is constructed based on contrastive loss, forward alignment loss, and backward feedback loss for end-to-end training.

[0091] In some embodiments, contrastive learning is used to align the embedding layer structure and the label layer to establish a contrastive loss, calculated as follows:

[0092] ;

[0093] in, Indicates comparative loss, This represents the local high-level semantic representation corresponding to the k-th modality data of the i-th sample. The modality category number is represented by m and n, which represent the sample number. This represents the label matrix for each sample in the k-th modality, where m and n represent the class numbers; N represents the number of samples, and C represents the number of classes. This represents the cosine similarity function.

[0094] In some embodiments, the maximum value among the local pseudo-label, fused pseudo-label, and global pseudo-label for each classification category is selected and normalized to obtain the global clustering label, including steps S201~S203:

[0095] Step S201: For the i-th sample with respect to classification m, select the maximum value of the label among the local pseudo-label, the fused pseudo-label, and the global pseudo-label. The expression is:

[0096] ;

[0097] in, This represents the global pseudo-label of the i-th sample with respect to category m. Let represent the local pseudo-label of the v-th modality in the i-th sample with respect to classification m. This represents the fusion pseudo-label of the i-th sample with respect to classification m.

[0098] Step S202: Calculate the global cluster label for a single sample, using the following formula:

[0099] ;

[0100] in, This represents the value corresponding to category p in the global cluster label. This represents the maximum value of the label with respect to category p. This represents the maximum number of labels for category q, where C is the number of categories.

[0101] Step S203: The formula for calculating the forward alignment loss is:

[0102] ;

[0103] in, This represents the forward alignment loss. This represents the local pseudo-label matrix of each sample with respect to mode v. Represents the global pseudo-label matrix. Represents the global clustering label matrix. This represents the fused pseudo-label matrix; Indicates KL divergence; This represents information entropy.

[0104] In some embodiments, a transition mapping module is used to perform transition mapping between local pseudo-tags and global pseudo-tags, and the expression is:

[0105] ;

[0106] ;

[0107] in, Represents the ReLU activation function. Indicates from the global pseudo-label matrix The intermediate representation obtained from learning. Indicates the fusion pseudo-label matrix The intermediate representation obtained from learning; , , and This is the weight matrix; , , and It is a fixed value.

[0108] The formula for calculating the back feedback loss is:

[0109] ;

[0110] in, This represents the reverse feedback loss;

[0111] The formula for calculating the joint loss is:

[0112] ;

[0113] in, For joint losses, and For weights.

[0114] On the other hand, the present invention also provides a multi-view clustering method apparatus based on hierarchical prompt guidance alignment, including a processor, a memory, and a computer program / instructions stored in the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the apparatus implements the steps of the above method.

[0115] On the other hand, the present invention also provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0116] The present invention will now be described with reference to a specific embodiment:

[0117] This embodiment provides a scheme for achieving fine-grained semantic alignment and consistent clustering in multi-view scenarios—Hierarchical Prompt-Guided Alignment for Multi-View Clustering (HiPMVC). First, a learnable prompting mechanism is introduced to enhance the original input data at the primary level, preserving and mining the unique semantic information of each view, thereby improving the expressive power of basic feature extraction. Second, at the embedding level, local prompts and global prompts are designed respectively. The former focuses on fine-grained clustering cues within a view, while the latter captures shared semantics between different views, thus compensating for the shortcomings of traditional coarse-grained fusion strategies in capturing complementary information between views. The features obtained after the input is processed by the encoder are further enhanced by fusing modality-specific embedded prompts and global prompts. Subsequently, these processed features and the global prompt are jointly input into the mapping encoder and mapped into a unified feature space. Finally, based on the multi-level labels generated by the pseudo-label classifier, a two-way cue-guided alignment operation is performed. The most representative cluster pseudo-labels are highlighted through forward alignment, and the global and local distributions are adaptively balanced by a feedback mechanism, thereby achieving collaborative alignment at the feature level and the pseudo-label level.

[0118] This embodiment introduces three types of cues: primary level, local embedding level, and global embedding level, to fully extract semantic information from the input data. The input data and global cues are mapped together into a unified latent space by a mapping encoder. Subsequently, high-level semantic features generate pseudo-labels through pseudo-label classifiers at the data level and cue level, and are aligned using a bidirectional cue-guided alignment mechanism, thereby improving multi-level semantic consistency and clustering performance.

[0119] Reference Figure 2 The specific implementation method of this embodiment is as follows:

[0120] 1. Primary level prompts for learning methods

[0121] To effectively extract semantic features from the input data, a learnable prompt is introduced for each sample in each modality. Let the multimodal input data be... ,in Indicates the number of samples. Indicates the number of modes. For the first The data dimension has multiple modalities. A multilayer perceptron (MLP) network is designed as the cue encoder. To extract the principal-level cue vector for each modality, the calculation formula is as follows:

[0122] ;

[0123] in, Maintain consistency with the original input The same number of samples and feature dimensions. This cue vector represents a cue operation based on the raw data, used to guide the extraction of semantic information. Subsequently, the input features... Its corresponding main-level cue vector The sums are then fed into a feature embedding network to obtain preliminary semantic feature representations. The calculation method is as follows:

[0124] ;

[0125] in, This represents the initial semantic feature representation after fusing the prompt information. To verify the universality of the prompt mechanism, in The design employs both MLP encoders and graph convolutional encoders (GCNs).

[0126] 2. Embedded Hierarchical Hint Learning Method

[0127] In most existing methods, feature processing can generally be divided into single-modal processing and cross-modal multimodal fusion. However, existing fusion mechanisms (such as attention mechanisms) often struggle to provide sufficiently fine-grained guidance for the extraction of shared semantics, primarily due to limitations in model capacity and the inherent modeling difficulties in unsupervised tasks. To address these shortcomings, this paper proposes introducing a modality-specific local cue generation network at the embedding level. Global suggestion generation network .in, Used to capture key clustering information in each modality, This is used to model shared clustering semantics across samples, enhancing global consistency among different instances.

[0128] Similar to the primary-level cue learning mechanism, a lightweight local cue generation network is designed. Global suggestion generation network These are used to generate modality-specific local cues. With global hints The method for extracting local prompts is as follows:

[0129] ;

[0130] in, , Indicates the first Preliminary semantic feature representations across modalities. To generate global cues. This requires cross-modal fusion of features from all modalities. To address this, a set of learnable adaptive weight parameters is introduced. This is used to dynamically adjust the contribution of each modal feature to the fusion result. The fusion calculation formula is as follows:

[0131] ;

[0132] Among them, weight Control Mode The importance of this in the fusion process. In obtaining local cues. With global hints Then, the preliminary semantic features are represented. This information is then fused with prompts and guided to be projected into a higher-level semantic clustering space. The process is defined as follows:

[0133] ;

[0134] in, To adjust the hyperparameters of local and global cue weights, The local high-level semantic representation corresponding to the v-th modal data of the sample is the final projected local high-level semantic representation.

[0135] Furthermore, inspired by the DealMVC method, an attention-based fusion strategy is introduced to represent features across multiple modalities. Aggregate the features to generate a unified fusion feature. The fusion process is defined as follows:

[0136] The weight coefficients are calculated based on attention mechanism networks and nonlinear feedforward neural networks, using the following formula:

[0137] ;

[0138] The local high-level semantic representations corresponding to v modal data in the sample are fused based on weighting coefficients, and the calculation formula is as follows:

[0139] ;

[0140] in, This represents the Hadamard product (element-by-element multiplication). This represents the weight matrix for the v-th modality in the sample. This represents an attention mechanism network. This represents a nonlinear feedforward neural network.

[0141] Due to global prompts It inherently contains clustering semantic information shared across samples, and further utilizes global prompts, leveraging... Projection modules with the same structure This maps them to a higher semantic level clustering space. This process can be represented as:

[0142] ;

[0143] 3. Two-way prompting and alignment mechanism

[0144] Given the distributional differences of multimodal data at the feature layer and clustering result level, cross-modal alignment needs to be performed simultaneously at both levels. To this end, a two-layer alignment strategy based on a pseudo-label classifier is proposed to effectively unify the feature representations of different modalities.

[0145] For each high-level feature representation and its mapping results in the unified semantic clustering space A representation-level pseudo-label generation module defined using a nonlinear multilayer perceptron (MLP) structure. and prompt-level pseudo-label generation module This results in the following pseudo-label representation:

[0146] ;

[0147] in, This represents the local pseudo-label matrix of each sample with respect to mode v. This represents the fused pseudo-label matrix.

[0148] Meanwhile, the global prompts are represented after semantic mapping. , The global pseudo-label matrix is ​​defined as follows:

[0149] ;

[0150] in, and Using the same non-linear MLP architecture, the output pseudo-label set is... .

[0151] After obtaining the representations of the feature layer and the pseudo-label layer, a contrastive learning strategy is employed to perform alignment operations simultaneously at these two layers. Specifically, the constructed contrastive loss function is defined as follows:

[0152] ;

[0153] in, This represents the contrast loss. The local high-level semantic representation corresponding to the k-th modality data of the i-th sample is represented by the following. The modality category number is represented by m and n, which represent the sample number. This represents the label matrix for each sample in the k-th modality, where m and n represent the class numbers; N represents the number of samples, and C represents the number of classes. This represents the cosine similarity function. Where, " "Used to identify hidden data dimensions."

[0154] The first term is used to align the feature layer representation, and the second term is used to align the semantic information of the pseudo-label layer, thereby achieving multimodal semantic alignment and consistent modeling. To enhance the consistency of clustering results and highlight the most representative category labels, a novel feedback mechanism is introduced, comprising two stages: forward alignment and backward feedback.

[0155] The core of forward alignment lies in highlighting effective clustering results and suppressing irrelevant or incorrect category labels. Specifically, it involves maximizing all generated labels to obtain the pseudo-labels for the final representation.

[0156] ;

[0157] in, Let represent the global pseudo-label of the i-th sample with respect to classification m. This represents the local pseudo-label of the v-th modality in the i-th sample with respect to classification m. This represents the fusion pseudo-label of the i-th sample with respect to classification m.

[0158] To further enhance representative labels and mitigate the influence of irrelevant labels, a global clustering label is constructed based on the above results. The expression for the p-th class is as follows:

[0159] ;

[0160] in, This represents the value corresponding to category p in the global clustering label. This represents the maximum value of the label with respect to category p. Let q represent the maximum value of the label for category q, where C is the number of categories.

[0161] Finally, the formula for calculating the forward alignment loss is:

[0162] ;

[0163] in, This represents the forward alignment loss. This represents the local pseudo-label matrix of each sample with respect to mode v. Represents the global pseudo-label matrix. Represents the global clustering label matrix. This represents the fused pseudo-label matrix; Indicates KL divergence; Represents information entropy. For example, its entropy is calculated as follows: .

[0164] To achieve reverse feedback, a transition mapping module is designed to perform a transition mapping on the global pseudo-label matrix. With fused pseudo-label matrix A balance is established between them to mitigate the instability caused by direct alignment, and a relatively stable clustering representation is constructed in the intermediate representation space. The calculation method of the transition mapping is as follows:

[0165] ;

[0166] ;

[0167] in, Represents the ReLU activation function. Indicates from the global pseudo-label matrix The intermediate representation obtained from learning. Indicates the fusion pseudo-label matrix The intermediate representation obtained from learning; , , and This is the weight matrix; , , and These are fixed values. Although they exist in the same class probability space, their distribution shape can be flexibly adjusted after parameterization. Especially in the early stages of model training, this transition mapping module acts as a buffer against direct alignment, contributing to training stability.

[0168] The formula for calculating the back feedback loss is defined as follows:

[0169] ;

[0170] The formula for calculating the joint loss is:

[0171] ;

[0172] in, For joint losses, and The weights are used as the base values. By jointly optimizing this objective function, all network parameters can be automatically learned in an end-to-end manner, and the final clustering results can be directly derived from the global labels. generate.

[0173] This embodiment can be applied to smart manufacturing scenarios, where multiple heterogeneous sensors are often deployed simultaneously on the production line: a triaxial vibration probe near the high-speed bearing samples at 1 kHz to capture minute mechanical impacts; thermocouples and infrared thermometers update temperature changes every second; current transformers report the main motor load in real time; an acoustic microphone array listens for abnormal howling at 48 kHz; and simultaneously, the PLC and SCADA logs of the CNC system asynchronously write key process parameters and alarm flags. The aforementioned vibration, temperature, current, acoustic, and control logs can be viewed as five complementary perspectives, continuously flowing to the production line edge gateway at different frequencies ranging from milliseconds to seconds.

[0174] To identify evolutionary clusters of equipment health status and predict potential failures under unsupervised conditions, the factory embedded a hierarchical cue-guided alignment multi-view streaming clustering method into its edge-cloud collaborative system. The cloud first aggregates months of historical multi-view data, aligning timestamps to a minimum of 10 ms and standardizing amplitude, temperature, and current dimensions. Then, based on the proposed "dual-layer cueing" mechanism: a learnable master cue vector is injected into each viewpoint in the main layer, emphasizing its unique signal characteristics; local cues and cross-view global cues are simultaneously generated in the embedding layer, and feature fusion is performed on the five-view features using an attention mechanism. After model training converges, the system fine-tunes the obtained pseudo-labels using k-means, solidifies them into a weight file, and periodically distributes them to the production line edge server.

[0175] Furthermore, to evaluate the performance of the method, experiments were conducted on seven multi-view datasets: BBCSport, Nuswide, Citeseer, Wiki, Caltech7, and Caltech-all. Brief information on these datasets is summarized in Table 1. The proposed method was compared with nine baseline multi-view clustering methods, including COMIC, CoMVC, SiMVC, MFLVC, MVD, DealMVC, GCFAgg, ICMVC, and VITAL.

[0176] In terms of implementation details, the HiPMVC method described in this embodiment is implemented on the Linux platform using PyTorch and trained on an NVIDIA RTX A6000 GPU. Training consisted of 500 epochs with a fixed learning rate of 0.001. Temperature parameters were set as follows: , The hyperparameters β and γ were tuned using a grid search on {0.01, 0.1, 1, 10, 100}. The results in the parameter analysis section show that the model is insensitive to these parameters; therefore, β = 1.0 and γ = 1.0 were ultimately chosen.

[0177] Table 1 Seven different multi-view datasets

[0178]

[0179] The experimental results are shown in Table 2, comparing the performance differences between the method described in this embodiment and existing multi-view clustering (MVC) methods. Clustering quality was evaluated on seven commonly used datasets, using metrics including accuracy (ACC), normalized mutual information (NMI), and purity. The results show that the method used in this embodiment achieves superior performance on all seven datasets. For example, on the BBCSport dataset, this model significantly outperforms other methods, exceeding the second-place method by 6-7%. Furthermore, for large-scale datasets such as YouTubeFace, the multi-level cue design enables the model to better capture cluster associations. Further observation reveals that the GCN-based variant consistently outperforms its MLP-based counterpart.

[0180] Table 2 shows the comparison of HiPMVC's clustering performance on various datasets with other models:

[0181]

[0182] like Figure 3 and 4As shown, the hyperparameters β and γ in the joint loss function of the method are set to range from [0.01, 0.1, 1, 10, 100], and experiments were conducted on the Citeseer and Caltech7 datasets. For the Caltech7 dataset, parameter γ exhibits high sensitivity in the range of 1 to 100; in contrast, for the Citeseer dataset with a smaller number of views, the sensitivity of both β and γ is relatively stable.

[0183] To visually demonstrate the clustering effect of HiPMVC, the t-SNE algorithm is used for visualization. For example... Figures 5 to 8 The diagram illustrates the clustering performance of the model proposed in this embodiment. Compared with the original features ( Figure 5 and Figure 7 The comparison reveals that without multi-view fusion, the distribution of categories is rather chaotic, with significant overlap or loose distribution between some categories. Conversely, after applying our method ( Figure 6 and Figure 8 The distribution of sample points in the low-dimensional space becomes more compact and the boundaries are clearer, forming well-structured and mutually separated clusters.

[0184] Corresponding to the above method, the present invention also provides an apparatus / system including a computer device, the computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the apparatus / system performs the steps of the method as described above.

[0185] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0186] In summary, the multi-view clustering method and apparatus based on hierarchical cue-guided alignment described in this invention constructs a main layer structure and an embedding layer structure to extract semantic features from multimodal data, and optimizes the clustering results through a label layer and a bidirectional alignment mechanism. The main layer cue encoder generates a main-level cue vector for each modality and adds it to the original data. This vector is then passed through a shallow encoder to obtain a preliminary semantic feature representation, which enhances the original input data and preserves and mines the unique semantic information of each view. The local cue generation network and the global cue generation network in the embedding layer structure output local and global cuees, respectively. After fusion, these cuees are projected onto the semantic clustering space through a mapping encoder to obtain local and global high-level semantic representations. An attention mechanism is then used to fuse the local high-level semantic representations of each modality, achieving fine-grained semantic enhancement and cross-view semantic alignment optimization. The label layer generates local pseudo-labels, fused pseudo-labels, and global pseudo-labels for each modality. The global clustering label is obtained by normalizing the maximum value, strengthening the most discriminative signal and weakening interference from low-confidence views. A contrastive loss is established for the embedding layer structure and the label layer through contrastive learning. A forward alignment loss is constructed by combining the probability distribution difference between global clustered labels and fused pseudo-labels. At the same time, a back feedback loss is constructed using the transition mapping module. Finally, the parameters are updated based on the joint loss to generate clustering results, taking into account both fine-grained semantic mining within the view and global consistency between views, thereby improving the clustering accuracy and stability.

[0187] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.

[0188] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.

[0189] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0190] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-view clustering method based on hierarchical prompts and alignment, characterized in that, The method includes the following steps: For a target scenario, obtain multi-view data to be clustered. The multi-view data contains multiple samples, and each sample has multiple modal views. A main layer structure is constructed. A main layer cue encoder is used to take the sample as input and output the main layer cue vector of each modality. The main layer cue vector of each modality is added to the original data and input into the shallow encoder to obtain the preliminary semantic feature representation of each modality. An embedding layer structure is constructed. For each modality, a local prompt generation network is established, taking the preliminary semantic feature representation of the corresponding modality as input and outputting local prompts. A global prompt generation network is established, taking the preliminary semantic feature representations fused based on adaptive weight parameters as input and outputting global prompts. The preliminary semantic feature representation, local prompts, and global prompts corresponding to each modality in the sample are fused and projected onto the semantic clustering space through a first mapping encoder to obtain the corresponding local high-level semantic representation. The global prompts are projected onto the semantic clustering space through a second mapping encoder to obtain the corresponding global high-level semantic representation. An attention mechanism is introduced to fuse the local high-level semantic representations of each modality in the sample to obtain a fused semantic representation. A label layer is constructed by inputting the local high-level semantic representation of each modality of the sample into the representation-level pseudo-label generation module to obtain the local pseudo-labels corresponding to each modality, and inputting the fused semantic representation into the representation-level pseudo-label generation module to obtain the fused pseudo-labels; the global high-level semantic representation is input into the prompt-level pseudo-label generation module to obtain the global pseudo-labels; for each classification category, the maximum value among the local pseudo-labels, the fused pseudo-labels, and the global pseudo-labels is selected and normalized to obtain the global clustering label; A contrastive learning method is used to align the embedding layer structure and the label layer to establish a contrastive loss. The probability distribution difference between the global clustering label and the fused pseudo-label is calculated, and the information entropy of the local pseudo-label and the global pseudo-label is suppressed to construct a forward alignment loss. A transition mapping module is used to perform transition mapping on the local pseudo-label and the global pseudo-label, and the probability distribution difference between them and the original global pseudo-label and global clustering label is calculated to construct a back feedback loss. Based on the contrastive loss, the forward alignment loss, and the back feedback loss, a joint loss is constructed to update the parameters of the embedding layer structure, the label layer, and the transition mapping module. Clustering results are generated based on the updated global clustering label.

2. The multi-view clustering method based on hierarchical prompt guidance alignment according to claim 1, characterized in that, The main layer prompt encoder is a multilayer perceptron encoder, and the shallow layer encoder is a multilayer perceptron encoder or a graph convolution encoder; the representation-level pseudo-label generation module and the prompt-level pseudo-label generation module are nonlinear support vector machines with the same structure.

3. The multi-view clustering method based on hierarchical prompt guidance alignment according to claim 2, characterized in that, A global prompt generation network is established, taking the preliminary semantic features fused based on adaptive weight parameters as input and outputting global prompts. The calculation formula is as follows: ; in, This indicates the global suggestion. This refers to the global suggestion generation network. This represents the adaptive weight parameter for the v-th modality data of the sample. The preliminary semantic feature representation of the v-th modality data of the sample.

4. The multi-view clustering method based on hierarchical prompting and alignment according to claim 3, characterized in that, The preliminary semantic feature representation, the local cue, and the global cue corresponding to each modality in the sample are fused and projected onto the semantic clustering space through the first mapping encoder to obtain the corresponding local high-level semantic representation. The calculation formula is as follows: ; in, This represents the local high-level semantic representation corresponding to the v-th modality data of the sample. This refers to the first mapping encoder. This represents the v-th modal data of the sample. These are preset parameters. This indicates the local cue corresponding to the vth modal data of the sample.

5. The multi-view clustering method based on hierarchical prompt guidance alignment according to claim 4, characterized in that, The local high-level semantic representations of each modality of the sample are fused using an attention mechanism to obtain a fused semantic representation, including: The weight coefficients are calculated based on attention mechanism networks and nonlinear feedforward neural networks, using the following formula: ; Based on the weighting coefficients, the local high-level semantic representations corresponding to v modal data in the sample are fused, and the calculation formula is as follows: ; in, This represents element-wise multiplication. This represents the weight matrix of the v-th modality data in the sample. This represents the attention mechanism network. This refers to the nonlinear feedforward neural network. This represents the fused semantic representation.

6. The multi-view clustering method based on hierarchical prompt guidance alignment according to claim 1, characterized in that, Contrastive learning is used to align the embedding layer structure and the label layer to establish a contrastive loss, calculated as follows: ; in, This represents the contrast loss. The local high-level semantic representation corresponding to the k-th modality data of the i-th sample is represented by the following. The modality category number is represented by m and n, which represent the sample number. This represents the label matrix for each sample in the k-th modality, where m and n represent the class numbers; N represents the number of samples, and C represents the number of classes. Represents the cosine similarity function. For the temperature parameters of the embedded structural layer, This refers to the temperature parameters for the label layer.

7. The multi-view clustering method based on hierarchical prompt guidance alignment according to claim 6, characterized in that, For each classification category, the maximum value among the local pseudo-label, the fused pseudo-label, and the global pseudo-label is selected and normalized to obtain the global clustering label, including: For the i-th sample with respect to classification m, the maximum value of the label is selected from the local pseudo-label, the fused pseudo-label, and the global pseudo-label, expressed as: ; in, Let represent the global pseudo-label of the i-th sample with respect to classification m. Let represent the local pseudo-label of the v-th modality in the i-th sample with respect to classification m. The fused pseudo-label of the i-th sample with respect to classification m; The global cluster label for a single sample is calculated using the following formula: ; in, This represents the value corresponding to category p in the global clustering label. This represents the maximum value of the label with respect to category p. Let C represent the maximum value of the label for category q, where C is the number of categories; The formula for calculating the forward alignment loss is: ; in, This represents the forward alignment loss. This represents the local pseudo-label matrix of each sample with respect to mode v. Represents the global pseudo-label matrix. Represents the global clustering label matrix. This represents the fused pseudo-label matrix; Indicates KL divergence; This represents information entropy.

8. The multi-view clustering method based on hierarchical prompt guidance alignment according to claim 7, characterized in that, The transition mapping module is used to perform transition mapping on the local pseudo-tags and the global pseudo-tags, and the expression is: ; ; in, Represents the ReLU activation function. This indicates the origin from the global pseudo-label matrix. The intermediate representation obtained from learning. This indicates the fusion pseudo-label matrix The intermediate representation obtained from learning; , , and This is the weight matrix; , , and It is a fixed value; The formula for calculating the reverse feedback loss is: ; in, This represents the reverse feedback loss; The formula for calculating the joint loss is: ; in, For the aforementioned joint loss, and For weights.

9. A multi-view clustering method apparatus based on hierarchical prompting and alignment, comprising a processor, a memory, and a computer program / instructions stored in the memory, characterized in that, The processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal clustering method and system based on feature fusion and label alignment

    CN118760913A

  • Infrared and visible light image data association method based on multi-view clustering

    CN119723134A