An uncertainty enhanced unsupervised multi-view data representation method

CN122839286APending Publication Date: 2026-09-29QINGDAO INNOVATION & DEV CENT OF HARBIN ENG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611126905.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]本发明提供了一种不确定性增强的无监督多视图数据表征方法,旨在针对应用场景中基于采集的多视图数据构建多视角学习时,在无监督学习条件下如何有效应对不同视图数据间噪声水平不一致以及样本层面视图质量差异显著的问题,对不同视图数据的置信程度进行动态建模,实现多视图信息的自适应融合,从而提升下游聚类等无监督任务的性能与鲁棒性

Benefits of technology

[0014]本发明的优点与积极效果在于:(1)本发明方法通过为每个视图分别构建独立的编码器与聚类分配网络,以提取潜在表征并完成视图自重构,保留各视图的特异性信息;通过在各个视图中独立执行软标签分配,实现跨视图的粗粒度结构对齐,引导各视图向一致的语义空间收敛;通过跨视图邻域信息互蒸馏进一步完成样本层面细粒度对齐,利用局部邻域相似性进行双向软标签约束;通过引入不确定性权重分配机制,根据各视图的表征置信度动态调节其融合贡献。本发明方法在保证各视图表示能力的同时,有效提升了多视图聚类的一致性,能够获得更稳定、判别性更强的融合表征。(2)本发明方法实现对应用场景采集的多视图数据的多粒度对齐与不确定性建模机制,能够处理典型的无配对、多视图、异构多视图数据中普遍存在的结构偏差与噪声扰动,生成更稳定、更具判别性的样本级表示,从而提升下游聚类等无监督任务的性能与鲁棒性。(3)对于实际应用场景,例如单细胞多组学数据融合问题中普遍存在的高维稀疏、掉落噪声严重以及不同组学数据质量不均衡的问题,本发明方法通过对各组学表征置信度进行不确定性建模并自适应分配融合权重,有效降低低质量组学对融合表示的干扰,提升融合后的表征一致性与判别性,从而显著增强下游无监督聚类任务的稳定性与准确性,实现更可靠的细胞类型划分。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839286A_ABST
    Figure CN122839286A_ABST
Patent Text Reader

Abstract

This invention discloses an unsupervised multi-view data representation method with enhanced uncertainty, belonging to the field of data mining technology, and capable of realizing biomedical data mining. Addressing the characteristics of high noise, large missing rate, and inconsistent distribution of data across different biomedical views, this invention, under unsupervised learning conditions, constructs an independent encoder and clustering assignment network for each view to extract latent representations and complete view self-reconstruction. It achieves coarse-grained structural alignment across views by independently performing soft label assignment in each view, and achieves fine-grained alignment at the sample level through cross-view neighborhood information inter-distillation. An uncertainty weight allocation mechanism is introduced to dynamically adjust the fusion contribution of each view based on its representation confidence. This invention effectively improves the consistency of multi-view clustering while ensuring the representational capabilities of each view, resulting in more stable and discriminative fusion representations, thereby improving the performance and robustness of downstream unsupervised tasks such as clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and data mining technology, and relates to biomedical data mining and clustering, specifically to an unsupervised multi-view data representation method with enhanced uncertainty. Background Technology

[0002] In real-world scenarios, data often comes from diverse sources and takes various forms, exhibiting significant heterogeneity. With the rapid development of sensing devices, network systems, and information acquisition technologies, data from different sources such as images, text, voice, sensor signals, or biological sequencing are used together to describe the same object, forming a typical multi-view data structure. Multi-view data, through the complementary features of multiple angles and perspectives, helps to more comprehensively characterize the potential properties of samples; therefore, multi-view learning has gradually become an important research direction in the fields of machine learning and data mining.

[0003] However, in the actual acquisition and processing of multi-view data, different views are often inevitably affected by various uncertainties. For example, sensor equipment may experience temporary malfunctions or performance degradation, data transmission may be incomplete or erroneous due to bandwidth limitations or interference, and data preprocessing may involve perturbation to protect privacy, resulting in varying degrees of noise pollution at the sample level. This problem is particularly prominent in the field of bioinformatics. For instance, in the fusion of multi-omics data and the joint analysis of medical images and clinical texts, data from different views often exhibit high noise, large missing rates, and inconsistent distribution. In other multi-view applications, such as video-audio analysis, image-text description generation, and remote sensing multi-source information fusion, significant differences in noise levels and uneven data quality between views are also common. These factors all lead to obvious and inconsistent noise fluctuations in multi-view data at the sample level, posing a significant challenge to traditional multi-view learning methods.

[0004] Clustering, as a typical unsupervised learning method, aims to uncover the latent structure of data without providing labels and divide samples into groups with similar characteristics. In multi-view scenarios, the quality differences of data from different perspectives directly affect the accuracy of sample representation, and thus impact clustering performance. If we consider all views... Figure 1 Simply merging views indiscriminately can easily allow noisy views to interfere with the overall representation, making it difficult to reveal the underlying structure. Therefore, under unsupervised conditions, how to dynamically model the confidence levels of different views based on the sample level, and then adaptively adjust the view weights to suppress the negative impact of noisy views and enhance the contribution of effective views, has become an important problem that urgently needs to be solved in multi-view representation learning.

[0005] Data collected from different perspectives in the biomedical field often exhibits characteristics such as high noise, large missing rates, and inconsistent distribution. For example, single-cell multi-omics data suffers from high-dimensional sparsity, severe dropout noise, and uneven quality among different omics datasets. How to effectively utilize the information between multiple omics datasets, reduce the impact of noise, and generate more reliable features to learn the global structure of single cells is a problem that needs to be solved. Summary of the Invention

[0006] This invention provides an unsupervised multi-view data representation method with enhanced uncertainty. It aims to address the challenges of inconsistent noise levels and significant differences in view quality among different view data under unsupervised learning conditions when constructing multi-view learning based on collected multi-view data in application scenarios. The method dynamically models the confidence level of different view data to achieve adaptive fusion of multi-view information, thereby improving the performance and robustness of downstream unsupervised tasks such as clustering.

[0007] This invention provides an unsupervised multi-view data representation method with enhanced uncertainty, comprising:

[0008] Step 1, obtain the application scenario Individuals in the study Datasets of class views;

[0009] Step 2: Construct an independent encoder, decoder, and clustering assignment network for each view. Input the data of each sample in the corresponding view into the encoder to extract the latent representation in its own latent semantic space. Map the latent representation to the shared latent space of all views through the clustering assignment network. The decoder reconstructs the data of each view corresponding to the sample based on the multi-view fusion representation of the sample. Construct the reconstruction loss for all samples. The reconstruction loss for each sample in each class view is defined as the mean square error between the data of that sample in that class view and the reconstruction data of the decoder corresponding to that class view.

[0010] Step 3, let the number of target categories be... By normalizing the mapping, the common representation of each class of view data for each sample in the shared latent space is mapped to its affiliation. Clustering soft labels for target classes are used to obtain the probability distribution of all samples for each target class in each view. To ensure that the probability distribution of the same target class is consistent across different views, and to distinguish between different target classes, a cross-view contrast loss is designed. The clustering-level alignment loss is obtained by summing the cross-view contrast losses for all ordered view pairs. ;

[0011] Step 4: Design a one-way distillation loss between views based on the cross-view neighborhood information inter-distillation mechanism, and average the one-way distillation losses of all samples to obtain the neighborhood inter-distillation loss. Simultaneously calculate the soft label consistency loss for view pairs. The sample-level alignment loss is obtained by combining the neighborhood inter-distillation loss and the soft label consistency loss. ;

[0012] Step 5: Perform weighted fusion of the common representations of various views for the sample based on the representation uncertainty; the representation uncertainty of each sample under each view class is obtained by calculating the Shannon entropy of the cluster soft label of the sample in that view; then determine the sample-level view weight vector based on the representation uncertainty. , Indicates sample In view The weights are then determined; the common representations of the sample under each view are then weighted and summed to obtain the multi-view fusion representation of the sample, which is used for subsequent target classification.

[0013] Step 6: Optimize and train the encoder, decoder, and clustering assignment network for all views in advance based on the total loss function; the total loss function is determined by... , and The weighted sum is obtained; the optimized encoder and clustering assignment network are used to convert the data of each class view of each sample into a common representation in the shared latent space, and then the multi-view fusion representation of the sample is calculated in step 5 for subsequent target clustering.

[0014] The advantages and positive effects of the present invention are as follows: (1) The method of the present invention extracts potential representations and completes view self-reconstruction by constructing an independent encoder and clustering assignment network for each view, and retains the specific information of each view; by independently performing soft label assignment in each view, it realizes coarse-grained structural alignment across views and guides each view to converge toward a consistent semantic space; by inter-distillation of cross-view neighborhood information, it further completes fine-grained alignment at the sample level and uses local neighborhood similarity for bidirectional soft label constraint; by introducing an uncertainty weight allocation mechanism, it dynamically adjusts the fusion contribution of each view according to the representation confidence. The method of the present invention effectively improves the consistency of multi-view clustering while ensuring the representation capability of each view, and can obtain a more stable and discriminative fusion representation. (2) The method of the present invention realizes a multi-granularity alignment and uncertainty modeling mechanism for multi-view data collected in application scenarios, which can handle the structural deviations and noise disturbances that are common in typical unpaired, multi-view, and heterogeneous multi-view data, and generate a more stable and discriminative sample-level representation, thereby improving the performance and robustness of downstream unsupervised tasks such as clustering. (3) For practical application scenarios, such as the common problems of high-dimensional sparsity, severe drop noise and uneven quality of different omics data in the problem of single-cell multi-omics data fusion, the method of the present invention effectively reduces the interference of low-quality omics on the fusion representation by performing uncertainty modeling on the confidence of each omics representation and adaptively allocating fusion weights, thereby improving the consistency and discriminability of the fused representation, significantly enhancing the stability and accuracy of downstream unsupervised clustering tasks, and achieving more reliable cell type classification. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the implementation of the uncertainty-enhanced unsupervised multi-view data representation method according to an embodiment of the present invention. Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0017] This invention provides an unsupervised multi-view data representation method with enhanced uncertainty, whose overall framework mainly includes the following four core modules:

[0018] (1) Construct an independent encoder and clustering assignment network for each view to extract latent representations and complete view self-reconstruction while preserving the specific information of each view;

[0019] (2) Perform soft label assignment independently in each view to achieve coarse-grained structural alignment across views and guide each view to converge toward a consistent semantic space;

[0020] (3) Further refine the fine-grained alignment at the sample level by cross-view neighborhood information inter-distillation, and use local neighborhood similarity to perform bidirectional soft label constraints;

[0021] (4) Introduce an uncertainty weight allocation mechanism to dynamically adjust the fusion contribution of each view based on the representation confidence of each view.

[0022] This invention effectively improves the consistency of multi-view clustering while ensuring the representational capabilities of each view, resulting in a more stable and discriminative fused representation of each object. This embodiment uses single-cell multi-omics data from the biomedical field as a specific scenario to illustrate the implementation of this invention. Each single cell represents a research individual, a sample.

[0023] Step one: First, acquire multi-omics data for each single cell, such as transcriptome scRNA-seq, chromatin accessibility scATAC-seq, proteomics scADT-seq, etc. Each type of omics can be considered as a view, and these are combined to form a... Heterogeneous multi-view datasets , Representing the A view dataset; assuming a total of Each sample is in the view. The data / features belong to the feature space , Represents the set of real numbers. For feature dimension, the first The view dataset is represented as follows:

[0024] (1)

[0025] in, Indicates the first The sample at the th Data for each view.

[0026] These view samples may exhibit significant differences in noise levels, data quality, and feature distributions, and are aligned only at the sample level; that is, different views... Describing the same sample .

[0027] Step two, in order to extract discriminative potential features from each view that can be used for cross-view alignment, embodiments of the present invention provide each view Design an independent encoder and decoder For the sample In view Data First through the encoder This invention extracts view features and maps the encoded features to the latent representation space of each view. To maintain the integrity of the input features and ensure the stability of the latent space learning process, a decoder corresponding to each view is introduced. The encoded features before projection are reconstructed using multi-view fusion representation. The first reconstructed sample was obtained View features :

[0028] (2)

[0029] The reconstruction loss function is defined as the mean square error between the input samples and the reconstructed samples. It is used to constrain the encoder to maintain the structure of the view content and suppress information loss caused by over-compression. The reconstruction loss is calculated. as follows:

[0030] (3)

[0031] Step 3: Perform coarse-grained cross-view structure alignment. In order to characterize the latent clustering structure of samples without real labels and quantify the representation confidence of different views, the method of this invention constructs an additional view-specific clustering assignment network for each view in the shared latent space. , will the Each view feature is projected by its individual encoder and then transmitted through a dedicated network. Mapped into the shared potential space of all views.

[0032] For the Each view constructs a normalized mapping. :

[0033] (4)

[0034] in For dimension The probability simplex, This indicates mapping features in the view feature space to... This transforms it into a valid discrete probability distribution. The dimension is The real number space. Total number of target categories.

[0035] Let the first The first sample The common representation of each view's data in the shared latent space is: ,as follows:

[0036] (5)

[0037] After normalization mapping, Corresponding clustering soft tags Represented as:

[0038] (6)

[0039] in, Indicates the first The first view The sample belongs to the first The probability of a target class.

[0040] For each view and each class index The probability of class belonging to any given class across all samples is expressed as:

[0041] (7)

[0042] in, For the first Class targets in the The probability distribution over all samples in the nth view has a shape that directly reflects the nth view. The semantic outline of the class; Indicates the first The sample at the th The view belongs to the first Probability of class; superscript This indicates transpose.

[0043] Given two views and They are calculated in the target class using cosine similarity. similarity for:

[0044] (8)

[0045] in, It is the first Class targets are respectively in the view and The probability distribution over all samples. This indicates the calculation of the L2 norm.

[0046] To ensure that classes with the same semantics behave consistently across different views, while also distinguishing between classes with different semantics, this invention designs the following cross-view comparison loss. (Based on view pairs...) For example, its cross-view contrast loss Defined as:

[0047] (9)

[0048] in The temperature coefficient is represented by the numerator of formula (9), which corresponds to the positive sample pair and is the view. With View For classes with the same index, other items in the denominator are treated as negative sample pairs, including different classes within the same view and different classes across views. In this embodiment of the invention, the default base parameter of the log function is 10.

[0049] The total cluster-level alignment loss is obtained by summing the cross-view contrast losses for all ordered view pairs. :

[0050] (10)

[0051] Step four involves fine-grained cross-view structure alignment. This invention further designs a cross-view neighborhood information inter-distillation mechanism at the sample level to achieve fine-grained consistency of local structures. For each view... Cosine similarity is calculated in the latent semantic space to construct a sample similarity matrix. For each sample... Record in view Its recent The set of neighbor indexes is:

[0052] (11)

[0053] Number of neighbors As a hyperparameter.

[0054] With two views For example, consider the sample In view soft tags in and its view Middle Neighborhood A neighbor randomly selected from the middle Define sample From view To view One-way distillation loss for:

[0055] (12)

[0056] in, This indicates that cosine similarity is used for calculation. and Similarity; the first term on the right side of formula (12) encourages samples In view The soft label and its view The cluster soft labels of the closest neighboring samples are close, and the latter two terms suppress false matches and in-view collapse by treating other samples as negative samples. Symmetrically, samples can be calculated From view To view Distillation loss The neighborhood inter-distillation loss is obtained by averaging the unidirectional distillation loss over all samples. :

[0057] (13)

[0058] for In general, updates can be accumulated across all view pairs or a rotational update strategy can be adopted, both of which are covered in this invention.

[0059] Simply using random neighbor distillation might overlook the consistency requirements of samples across different views. Therefore, this invention further introduces soft label consistency constraints for samples with the same index. (Calculate view pairs) soft tag consistency loss for:

[0060] (14)

[0061] Combining neighborhood inter-distillation and consistency regularization, the fine-grained sample-level alignment loss of this invention Defined as:

[0062] (15)

[0063] in It is used to control the strength of consistency constraints.

[0064] Through fine-grained sample-level alignment constraints, weak views can obtain semantic structure guidance from strong views, while strong views can also use the supplementary information from weak views for fine-grained corrections, thus achieving true bidirectional complementarity.

[0065] Step five involves implementing multi-view adaptive fusion based on uncertainty enhancement. After completing multi-granularity alignment of sample multi-views through their respective encoders and clustering assignment networks, this invention performs sample-level weighted fusion of different views based on representational uncertainty to explicitly suppress the interference of high-noise views on the clustering structure. (For the view...) The following samples The uncertainty of this characteristic is defined as the Shannon entropy of the soft clustering distribution. (Labeled samples) The uncertainty of each view is ,as follows:

[0066] (16)

[0067] This invention determines the sample-level view weight vector by constructing and solving the following convex optimization problem. :

[0068] (17)

[0069] in, for 3D probabilistic simplex, constraints ,parameter The second term is the negative entropy regularization, used to avoid "view collapse" caused by extreme concentration of weights in a single view. Solving the above equation yields a unique closed-form solution for the Gibbs distribution:

[0070] (18)

[0071] This shows that uncertainty Smaller views, more reliable representations, have higher weights. The larger the value, the lower the weight; conversely, the lower the value, the lower the weight. Based on the obtained weights, this invention linearly weights the potential representations of each view to obtain a sample-level fused representation:

[0072] (19)

[0073] Fusion characterization By integrating the information contributions of multiple views at the sample level and explicitly mitigating the influence of unreliable views through uncertainty modeling, it provides a more stable and discriminative representation that can improve the performance of subsequent target clustering.

[0074] Step six: Combining the aforementioned reconstruction constraints, coarse-grained cluster-level alignment constraints, and fine-grained sample-level alignment constraints, this invention performs overall training on all models. All models requiring training include the encoder for each view. decoder Clustering assignment network for each view Overall loss function for:

[0075] (20)

[0076] in, This is a hyperparameter used to balance the losses of each component. It utilizes the overall loss function. Train all models, then use the optimized model to analyze the samples. Alignment of view data is performed by mapping the data of the corresponding view of the sample to the encoder and clustering assignment network obtained by optimizing each view. Then, in the fundamental step 5, multi-view fusion is performed to obtain a more discriminative multi-view representation for downstream unsupervised tasks such as clustering.

[0077] This embodiment compares the clustering performance of the proposed uncertainty-enhanced unsupervised multi-view data representation method (scUML) with several representative multi-view clustering baseline methods, including DealMVC, scDRAE, scEMC, scMCs, and MoClust, on multiple publicly available single-cell multi-omics datasets. In the experiments, each single cell serves as a sample to be characterized and clustered. Measurement data obtained from the same cell at different omics levels constitute different views, such as gene expression profiles, chromatin accessibility profiles, or cell surface protein abundance. To ensure fairness in the comparison, all methods use ARI (Adjusted Rand Index), AMI (Adjusted Mutual Information), and NMI (Normalized Mutual Information) as evaluation metrics to measure the consistency between the clustering results and the actual cell type labels. As shown in Table 1, scUML demonstrates good clustering performance on the MA-53, MA-54, MA-55, MA-56, and Human cell line mixture datasets.

[0078] Table 1. Performance comparison of different methods for clustering on multiple public single-cell multi-omics datasets.

[0079]

[0080] The experimental results above show that the method of the present invention can effectively integrate multi-omics observation information from the same cell, has good effectiveness and robustness, and can significantly improve the performance of downstream clustering tasks such as cell type identification.

[0081] Except for the technical features described in the specification, all other technologies are known to those skilled in the art. Descriptions of well-known components and technologies are omitted in this invention to avoid redundancy and unnecessary limitation. The embodiments described above do not represent all embodiments consistent with this application. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this invention are still within the protection scope of this invention.

[0082] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

Claims

1. An unsupervised multi-view data representation method with enhanced uncertainty, characterized in that, Includes the following steps: Step 1, obtain the application scenario Individuals in the study Datasets of class views; Step 2: Construct an independent encoder, decoder, and clustering assignment network for each view. Input the data of each sample in the corresponding view into the encoder to extract the latent representation in its own latent semantic space. Then, map the latent representation to the shared latent space of all views through the clustering assignment network. The decoder reconstructs the corresponding view data of each sample based on the multi-view fusion representation of the sample; and constructs the reconstruction loss for all samples. The reconstruction loss for each sample in each class view is defined as the mean square error between the data of that sample in that class view and the reconstruction data of the decoder corresponding to that class view. Step 3, let the number of target categories be... By normalizing the mapping, the common representation of each class of view data for each sample in the shared latent space is mapped to its affiliation. Clustering soft labels for target classes are used to obtain the probability distribution of all samples for each target class in each view. To ensure that the probability distribution of the same target class is consistent across different views, and to distinguish between different target classes, a cross-view contrast loss is designed. The clustering-level alignment loss is obtained by summing the cross-view contrast losses for all ordered view pairs. ; Step 4: Design a one-way distillation loss between views based on the cross-view neighborhood information inter-distillation mechanism, and average the one-way distillation losses of all samples to obtain the neighborhood inter-distillation loss. Simultaneously calculate the soft label consistency loss for view pairs. The sample-level alignment loss is obtained by combining the neighborhood inter-distillation loss and the soft label consistency loss. ; Step 5: Perform weighted fusion of the common representations of various views of the sample based on representation uncertainty; The representation uncertainty of each sample under each class view is obtained by calculating the Shannon entropy of the cluster soft label of that sample in that view; Then, based on the aforementioned representation uncertainty, the sample-level view weight vector is determined. , Indicates sample In view The weights; Then, the common representations of the sample under each view are weighted and summed to obtain the multi-view fusion representation of the sample; Step 6: Optimize and train the encoder, decoder, and clustering assignment network for all views in advance based on the total loss function; the total loss function is determined by... , and We get the result by weighted summation; The optimized encoder and clustering assignment network are used to transform the data of each class view of the input samples into a common representation in a shared latent space; Then, step 5 is used to calculate the multi-view fusion representation of the samples for subsequent target clustering.

2. The method according to claim 1, characterized in that, Step 3 includes: Let the sample In view Data via view After mapping the encoder and clustering assignment network, the representation in the shared latent space is obtained. , Cluster soft labels are obtained through normalized mapping. ; Indicates sample In view Belongs to the The probability of class target; the probability of the first target. Class target in view The probability distribution of all samples is expressed as: ,in Indicates sample In view Belongs to the The probability of class targets, Superscript Indicates transpose; Then calculate the view pair Cross-view contrast loss as follows: ; in, For temperature coefficient, ; It is the first Class targets are respectively in the view ,view The probability distribution of all samples on the dataset; This indicates that the view is calculated using cosine similarity. and view In the Similarity between target classes; The cluster-level alignment loss is obtained by summing the cross-view contrast losses for all ordered view pairs. .

3. The method according to claim 1 or 2, characterized in that, Step 4 includes: For each sample, calculate the cosine similarity with other samples in the latent semantic space of each view, and find the nearest neighbor set. Let the sample... In view The nearest neighbor set in is , Number of neighbors; For the sample Two different views and From the sample In view nearest neighbor set A neighbor sample was randomly selected from the samples. Calculate samples From view To view One-way distillation loss for: ; in, For temperature coefficient, ; It is a sample In view Clustering soft tags in It is a sample In view Clustering soft tags in; and These are samples In view ,view Clustering soft tags in; This indicates that cosine similarity is used for calculation. and Similarity; The neighborhood inter-distillation loss is obtained by averaging the unidirectional distillation loss for all samples. as follows: ; in, For the sample From view To view One-way distillation loss; Calculate view pairs soft tag consistency loss for: ; Further, sample-level alignment loss is obtained. for: ;in This is used to control the strength of soft tag consistency constraints.

4. The method according to claim 1, characterized in that, Step 5 includes: Let the sample In view Clustering soft tags , Indicates sample In view Belongs to the The probability of a target class is calculated from the sample. In view Uncertainty ; Construct the following convex optimization problem and solve it to determine the samples. View weight vector ; ; in, for 3D probabilistic simplex, constraints ,parameter The solution yields: ; Furthermore, samples were obtained. Multi-view fusion representation ;in It is a sample In view Data via view The common representation is obtained after the encoder and clustering assignment network are mapped.

5. The method according to claim 1, characterized in that, Step 6 involves calculating the total loss function. for: ; in, To balance the hyperparameters of loss for each component, .

6. The method according to claim 1 or 5, characterized in that, In step 2, the reconstruction loss of all samples is calculated. as follows: ; in, It is a sample In view Data; assuming sample Multi-view fusion is characterized as , It is a view The encoder according to Reconstructed samples In view The data.

7. The method according to claim 1, characterized in that, In step 1, single-cell multi-omics data are acquired. Each single cell is a sample, and each omics is a view. The omics types acquired for single cells include transcriptomics, chromatin accessibility profiles, and cell surface proteomics.