Single-cell multi-omics cell type annotation method based on distribution and knowledge alignment

Through a multi-omics variational autoencoder model based on distribution and knowledge alignment, the high-dimensional sparsity and noise problems in single-cell data annotation are solved, efficient and accurate multi-omics data integration and annotation are achieved, and the adaptability to diverse experimental conditions is enhanced.

CN120636558APending Publication Date: 2025-09-12CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510440159.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing single-cell data annotation methods have difficulty dealing with high-dimensional sparsity and technical noise, ignore multi-omics information, resulting in information loss and reduced accuracy, and make the integration and analysis of multi-omics data difficult. Traditional methods have high computational complexity.

Method used

A multi-omics variational autoencoder model (scLTH) based on distribution and knowledge alignment is adopted, combined with variational autoencoders and knowledge distillation technology. The distribution characteristics of single-cell transcriptome data and chromatin accessibility sequencing data are aligned through the multi-omics variational autoencoder model, the self-attention mechanism is used to capture shared features, and knowledge is transferred through the teacher-student model framework to reduce computing resources.

Benefits of technology

It achieves efficient and accurate multi-omics data annotation and integration, enhances adaptability to diverse experimental conditions, and improves the ability to capture and generalize cross-omics features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005350812450000041
    Figure BDA0005350812450000041
  • Figure BDA0005350812450000051
    Figure BDA0005350812450000051
  • Figure BDA0005350812450000052
    Figure BDA0005350812450000052
Patent Text Reader

Abstract

The invention provides a single-cell multi-omics cell type annotation method based on distribution and knowledge alignment, and belongs to the technical field of single-cell type annotation, the method comprises the following steps: obtaining single-cell transcriptome data and single-cell chromatin accessibility sequencing data, and pre-training and training a multi-omics variation auto-encoder model, the multi-omics variational auto-encoder model is combined with a variational auto-encoder and a knowledge distillation technology, and multi-omics single cell data is integrated and annotated through distribution and knowledge alignment. And performing cell type prediction on the single cell transcriptome data and the single cell chromatin accessibility sequencing data which are input at the same time by using the trained multi-omics variational auto-encoder model. According to the method, the problem of limitation of a method only depending on single omics is solved, the synergistic effect between the omics is enhanced, the accuracy of annotation is improved, and the calculation overhead is reduced through knowledge distillation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of single cell type annotation, and in particular relates to a single cell multi-omics cell type annotation method based on distribution and knowledge alignment. Background Art

[0002] Single-cell type annotation is crucial for understanding the tumor microenvironment. Advances in single-cell sequencing technologies have enabled researchers to analyze the functional characteristics of different cell types at the single-cell level. However, existing annotation methods struggle to handle the high-dimensional sparsity and technical noise inherent in single-cell data. Traditional methods primarily rely on transcriptome data, neglecting complementary multi-omics information, leading to information loss and reduced accuracy. Furthermore, existing multi-omics methods, due to the complexity of their models, often incur high training and prediction costs.

[0003] With the emergence of emerging technologies such as scATAC-seq, multi-omics sequencing methods have gained significant attention, providing unprecedented opportunities for the integrated analysis of multidimensional data. By combining different omics dimensions such as transcriptomics and epigenomics, multi-omics data can provide a more comprehensive understanding of cellular regulatory networks and more accurately characterize cell states and functions. However, the inherent high dimensionality, sparsity, and distribution heterogeneity of multi-omics data pose difficulties for integrated analysis. Addressing these difficulties requires the development of effective computational models that can coordinate the distributions between multi-omics datasets while capturing their shared features. This remains a focus of current research.

[0004] Existing multi-omics data integration strategies can be roughly divided into two paradigms. The first is data-level integration, which involves connecting or mapping different omics data into a shared feature space, and then applying traditional machine learning or deep learning techniques for dimensionality reduction, clustering or classification. Although this approach is intuitive, it often ignores the specificity and interrelationships between each omics, resulting in suboptimal integration results. The second paradigm is model-level integration, which uses deep neural networks to encode individual omics datasets and map them into a shared latent space to achieve cross-omics integration. Although these methods perform well in extracting shared features across omics, they are usually accompanied by high computational complexity and rely heavily on large-scale datasets. Summary of the Invention

[0005] In response to the above-mentioned deficiencies in the prior art, the present invention provides a single-cell multi-omics cell type annotation method based on distribution and knowledge alignment, which reduces the dependence of traditional methods on a single modality and the large amount of computing resources required for multimodal models, and enhances adaptability to diverse experimental conditions.

[0006] To achieve the above objectives, the present invention adopts a technical solution: a single-cell multi-omics cell type annotation method based on distribution and knowledge alignment, comprising the following steps:

[0007] S1. Obtain single-cell transcriptome data and single-cell chromatin accessibility sequencing data, and pre-train and train a multi-omics variational autoencoder model. The multi-omics variational autoencoder model combines variational autoencoders and knowledge distillation techniques to integrate and annotate multi-omics single-cell data through distribution and knowledge alignment.

[0008] S2. Use the trained multi-omics variational autoencoder model to predict cell types based on the simultaneously input single-cell transcriptome data and single-cell chromatin accessibility sequencing data.

[0009] The beneficial effects of the present invention are: the present invention proposes a novel multi-omics variational autoencoder (VAE) model, named scLTH. The model combines latent space representation learning with unsupervised knowledge distillation, aiming to achieve efficient multi-omics integration and annotation. Through joint pre-training, the multi-omics variational autoencoder model scLTH constructs a shared latent space to align the distribution characteristics of single-cell transcriptome data (scRNA-seq) and single-cell chromatin accessibility sequencing data (scATAC-seq). Subsequently, knowledge distillation prompts the teacher model to transfer knowledge to the student model. The teacher model aims to capture complex feature representations, while the student model inherits the optimized latent space representation. While significantly reducing computational complexity and parameterization, the inference efficiency is improved. Finally, the encoder and student models are used for training and prediction tasks to achieve efficient and accurate multi-omics data annotation and integration.

[0010] Furthermore, the pre-training process of the multi-omics variational autoencoder model is as follows:

[0011] The single-cell transcriptome data and single-cell chromatin accessibility sequencing data are respectively passed through their respective encoders to extract the potential representation Z of the modal features. RNA and Z ATAC , and integrate the modal feature potential representation Z RNA and Z ATAC , get the fused potential representation Z fused ;

[0012] Using KL divergence loss, we can ensure that the potential representation Z of different omics RNA and Z ATAC Align in a shared latent space to minimize the latent representation Z RNA and Z ATAC Distribution differences in the latent space, and the use of reconstruction loss to ensure that the reconstructed latent representation is aligned with the original data in the latent space;

[0013] The fused latent representation Z fused The data is fed into the teacher model to extract cross-omics features from high-dimensional data. The self-attention mechanism is used to capture the shared features between single-cell transcriptome data and single-cell chromatin accessibility sequencing data.

[0014] Use the student model to receive the shared potential representation y passed by the teacher model teacher , learning shared features, where the student model receives the shared latent representation y teacher Includes cross-omics features;

[0015] The output of the student model is compared with the output of the teacher model, and the student model learns features similar to those of the teacher model by minimizing the difference between the two. distill , completing the pre-training of the multi-omics variational autoencoder model.

[0016] The beneficial effect of the above further scheme is that the present invention combines single-cell transcriptome data (scRNA-seq) and single-cell chromatin accessibility sequencing data (scATAC-seq) by designing two alignment strategies (inter-modality feature alignment and inter-model knowledge alignment). The inter-modality feature alignment strategy corrects the distribution differences between different modalities, enabling seamless mapping in the latent space and more accurately capturing multi-omics features. The inter-model knowledge alignment strategy transfers structured knowledge from the teacher model to the student model, balancing computational efficiency and annotation performance. This approach reduces the dependence on a single modality in traditional methods and the large amount of computing resources required for multimodal models, and enhances adaptability to diverse experimental conditions. Experimental results on public datasets show that the proposed method surpasses the most advanced single-omics models in multiple benchmarks, demonstrating strong annotation capabilities and excellent generalization capabilities for rare cell types.

[0017] Furthermore, the potential representation Z RNA Generated by the following formula:

[0018] Z RNA =μ+σ·ε

[0019] μ,logσ 2 =Encoder(X RNA )

[0020] The potential representation Z ATAC Generated by the following formula:

[0021] Z ATAC =μ+σ·ε

[0022] μ,logσ 2 =Encoder(X ATAC )

[0023] where μ and σ represent the mean and standard deviation of the latent representation of the encoder output, ε represents the noise sampled from the standard normal distribution, and logσ 2 Represents logarithmic variance, Encoder() represents encoding operation, X RNA represents the input single-cell transcriptome data, X ATAC Represents the input single-cell chromatin accessibility sequencing data.

[0024] The beneficial effect of this further approach is that, through a multi-layered nonlinear mapping structure, the present invention reduces high-dimensional gene expression data to a lower dimension while preserving key information. Through layer-by-layer transformations, the final latent representation effectively captures the complex relationships between transcriptional signatures and gene expression within the cellular transcriptome.

[0025] Furthermore, the formula of the KL divergence loss is as follows:

[0026] L KL =D KL (p RNA (Z RNA )||p fused (Z fused ))+D KL (p ATAC (Z ATAC )||p fused (Z fused ))

[0027] Among them, L KL represents the KL divergence loss, D KL represents KL divergence, p RNA and p ATAC Represent the potential distribution of single-cell transcriptome data and single-cell chromatin accessibility sequencing data, respectively, p fused represents the distribution of the target shared latent space;

[0028] The reconstruction loss formula is as follows:

[0029]

[0030] Among them, L reconstruction represents the reconstruction loss function, N i represents the data dimension of mode i, ω ij represents the weight of the jth feature in modality i, x ij and x ij represent the jth component of the original data and reconstructed data of mode i respectively.

[0031] The beneficial effects of the above further scheme are: by minimizing the KL divergence, the multi-omics variational autoencoder model of the present invention optimizes the data alignment in the latent space, ensuring that data from different omics have similar distribution characteristics in the shared latent space; by minimizing the reconstruction error, the multi-omics variational autoencoder model can fuse data from different omics while retaining the original cellular information and ensuring their consistency in the latent space.

[0032] Furthermore, the formula of the cross-omics feature is as follows:

[0033]

[0034]

[0035] F = ReLU(A multi-head W1+b1)W2+b2

[0036] A multi-head =Concat(Attention1,Attention2,...,Attention h )W out

[0037]

[0038] Z fused =Combiner(Z RNA ,Z ATAC )

[0039] Among them, y teacher Represents the cross-omics features output by the teacher model, including all learned cross-omics feature representations, FCout represents the fully connected layer, Represents the features output by the self-attention mechanism and the feedforward neural network, F represents the output of the feedforward neural network, ReLU represents the ReLU activation function, and A multi-head Represents the concatenation result of multiple attention heads after linear transformation processing, W1 and W2 both represent the weights of the feedforward neural network, b1 and b2 both represent the bias of the feedforward neural network, Concat represents the concatenation operation, Attention h represents the output of the h-th attention head, W out Represents the linear transformation matrix, A represents the final output of the single-head attention in the self-attention mechanism, Attention represents the self-attention operation, Q, K and V represent Query, Key and Value respectively, and softmax represents the normalization function. represents the similarity between the query and the key, d k represents the dimension of the key vector, T represents transpose, Represents the potential representation Z fused The high-dimensional representation mapped to by the latent layer, W Q 、W k and W v Represent the weight matrices of query, key and value respectively, W embedding and b embedding They represent the weight matrix and bias term of the embedding layer respectively, and Combiner() represents the fusion operation, which includes the connection operation and a fully connected layer.

[0040] The beneficial effect of the above further scheme is that the teacher model extracts complex cross-omics features from high-dimensional data through the Transformer network and uses the self-attention mechanism to capture shared features between single-cell transcriptome data (scRNA-seq) and single-cell chromatin accessibility sequencing data (scATAC-seq).

[0041] Furthermore, the feature L distill The formula is as follows:

[0042] L distill =MSE(y student ,y teacher )

[0043] Among them, MSE represents the mean square error calculation, y student represents the output of the student model, y teacher A shared latent representation representing the output of the teacher model.

[0044] The beneficial effect of the above further solution is that the present invention compares the output of the student model with the output of the teacher model. By minimizing the difference between the two, the student model learns features similar to those of the teacher model. This process enables the student model to efficiently learn cross-omics features and align these features to a shared latent space under low-parameter conditions. Through knowledge distillation, the student model inherits the advantages of the teacher model in cross-omics feature learning and achieves efficient feature alignment through a lightweight design. With low computing resources, the student model is able to process large-scale multi-omics data and perform similarly to the teacher model on tasks.

[0045] Furthermore, the training process of the multi-omics variational autoencoder model includes:

[0046] Based on the pre-training results, the fused latent representation processed by the encoder, the aggregator, and the student model is used as the input of the classifier;

[0047] The output of the student model is mapped to the predicted cell type using a classifier. There are two encoders, corresponding to single-cell transcriptome data and single-cell chromatin accessibility sequencing data of different omics.

[0048] The beneficial effect of the above further solution is that in the present invention, during the optimization and alignment training process, the model parameters θ are optimized to minimize the total loss function. This optimization ensures that the predicted labels are as close as possible to the true labels, thereby improving the classification accuracy. At the same time, the fusion representation Z fused It captures transcriptional and epigenetic signatures, providing rich contextual information for cell type annotation.

[0049] Furthermore, the loss function formula of the multi-omics variational autoencoder model during training is as follows:

[0050]

[0051] Where θ* represents the minimum loss function of the multi-omics variational autoencoder model during training, θ represents the parameters of the multi-omics variational autoencoder model, and L CE represents the cross entropy loss, y pred represents the predicted probability of cell type, y true represents the true label, C represents the number of cell types, and y i' represents the true label, y i' represents the predicted probability of cell category i'. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Flow chart of the method of the present invention.

[0053] Figure 2 Schematic diagram of the training and prediction model structure of the present invention.

[0054] Figure 3 This is a framework diagram of the pre-training model of the present invention.

[0055] Figure 4 This is a comparison chart between the present invention and expert classification.

[0056] Figure 5 Schematic diagram of clustering scores for original data and shared integrated examples in four datasets. DETAILED DESCRIPTION

[0057] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0058] Example

[0059] This paper introduces a multi-omics variational autoencoder model, scLTH, which combines variational autoencoders (VAEs) and knowledge distillation (KD) techniques to efficiently integrate and annotate multi-omics single-cell data through distribution and knowledge alignment. The input of this model is single-cell transcriptome data (scRNA-seq) and single-cell chromatin accessibility sequencing data (scATAC-seq). These two data are independently processed by two encoders and then fused in a shared latent space. In this latent space, the teacher model uses Transformer to extract shared features, optimize the distribution consistency across omics datasets, generate global feature representations, and pass this knowledge to the student model. The student model inherits the knowledge of the teacher model while reducing the number of parameters and improving inference efficiency, thereby enhancing the model's generalization ability. The framework combines reconstruction loss and KL divergence loss to ensure consistency between input data and reconstructed data, while mean squared error (MSE) loss helps the student model learn latent space features and further align knowledge between different omics datasets. This method improves the performance of cell type annotation and downstream tasks. Through self-supervised learning, the multi-omics variational autoencoder model scLTH effectively integrates single-cell transcriptome data (scRNA-seq) and single-cell chromatin accessibility sequencing data (scATAC-seq), optimizing the distribution and knowledge alignment of cross-omics datasets. It significantly enhances the interpretation and annotation capabilities of single-cell data and provides a new solution for multi-omics data integration. Figure 1-Figure 3 As shown, Figure 2 (a) Integrating single-cell transcriptome data scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq, the encoder and prediction model are used to predict cell types in a shared latent space. Figure 2 (b) is the multi-omics distribution alignment, Figure 2 (c) is knowledge distillation alignment. The present invention provides a single-cell multi-omics cell type annotation method based on distribution and knowledge alignment, which is implemented as follows:

[0060] S1. Obtain single-cell transcriptome data and single-cell chromatin accessibility sequencing data, and pre-train and train a multi-omics variational autoencoder model. The multi-omics variational autoencoder model combines variational autoencoders and knowledge distillation techniques to integrate and annotate multi-omics single-cell data through distribution and knowledge alignment.

[0061] S2. Use the trained multi-omics variational autoencoder model to predict cell types based on the simultaneously input single-cell transcriptome data and single-cell chromatin accessibility sequencing data.

[0062] In this embodiment, the pre-training and training process of the multi-omics variational autoencoder model is as follows:

[0063] The single-cell transcriptome data and single-cell chromatin accessibility sequencing data are respectively passed through their respective encoders to extract the potential representation Z of the modal features. RNA and Z ATAC , and integrate the modal feature potential representation Z RNA and Z ATAC , get the fused potential representation Z fused ;

[0064] Using KL divergence loss, we can ensure that the potential representation Z of different omics RNA and Z ATAC Align in a shared latent space to minimize the latent representation Z RNA and Z ATAC Distribution differences in the latent space, and the use of reconstruction loss to ensure that the reconstructed latent representation is aligned with the original data in the latent space;

[0065] The fused latent representation Z fused The data is fed into the teacher model to extract cross-omics features from high-dimensional data. The self-attention mechanism is used to capture the shared features between single-cell transcriptome data and single-cell chromatin accessibility sequencing data.

[0066] Use the student model to receive the shared potential representation y passed by the teacher model teacher , learning shared features, where the student model receives the shared latent representation y teacher Includes cross-omics features;

[0067] The output of the student model is compared with the output of the teacher model, and the student model learns features similar to those of the teacher model by minimizing the difference between the two. distill , complete the pre-training of the multi-omics variational autoencoder model;

[0068] The training process of the multi-omics variational autoencoder model includes:

[0069] Based on the pre-training results, the fused latent representation processed by the encoder, the aggregator, and the student model is used as the input of the classifier;

[0070] The output of the student model is mapped to the predicted cell type using a classifier. There are two encoders, corresponding to single-cell transcriptome data and single-cell chromatin accessibility sequencing data of different omics.

[0071] In this embodiment, the pre-training process is as follows:

[0072] Step 1: The input single-cell transcriptome data scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq are processed through their respective encoders to extract modality-specific latent representations, ZRNA and Z ATAC ,Then, these representations are connected through a fuser module to form the fused latent representation Z fused .

[0073] The multi-omics variational autoencoder model (scLTH) uses a specially designed distributed encoder architecture to extract features from multi-omics data, efficiently processing data from different omics layers and extracting latent features. Each encoder maps the input data into a low-dimensional latent space through multiple nonlinear transformation layers, capturing key features related to cellular function and state.

[0074] Single-cell transcriptome data, scRNA-seq, is primarily used to extract the transcriptional state of cells. This data is typically high-dimensional, containing the expression levels of numerous genes and reflecting the gene expression profile of the cell. To process this data, the encoder designs a multi-layer nonlinear mapping structure, including fully connected layers (FC), batch normalization, and ReLU activation.

[0075] Through these transformation layers, high-dimensional gene expression data is gradually reduced in dimensionality while retaining key information. Through layer-by-layer transformation, the final latent representation effectively captures the complex relationship between transcriptional features and gene expression in the cell transcriptome. At the end of the single-cell transcriptome data scRNA-seq encoder, the output is the mean μ and logarithmic variance logσ of the latent representation 2 , which is calculated as follows:

[0076] μ,logσ 2 =Encoder(X RNA )

[0077] Then, using the reparameterization technique, the latent representation Z RNA Generated by the following formula:

[0078] Z RNA =μ+σ·ε

[0079] Potential representation Z ATAC Generated by the following formula:

[0080] Z ATAC =μ+σ·ε

[0081] μ,logσ 2 =Encoder(X ATAC )

[0082] where μ and σ represent the mean and standard deviation of the latent representation of the encoder output, respectively, and logσ 2Represents logarithmic variance, Encoder() represents encoding operation, X RNA represents the input single-cell transcriptome data, X ATAC represents the input single-cell chromatin accessibility sequencing data, and ε~N(0,I) represents the noise sampled from the standard normal distribution N(0,I). In this way, the model is able to learn key information about the transcriptome in the latent space, providing a foundation for subsequent multi-omics feature fusion. The computational process of each layer of the scATAC-seq encoder for single-cell chromatin accessibility sequencing data is the same as that of the scRNA-seq encoder for single-cell transcriptome data, as shown below:

[0083] h (l) =ReLU(BN(W (l) h (l-1) +b (l) ))

[0084] Among them, h (l) represents the output of the l-th layer encoder, W (l) and b (l) Represent the weight and bias of the layer respectively, h (l-1) Represents the output of the previous layer, BN represents batch normalization, and ReLU represents the activation function.

[0085] Ultimately, the encoder maps the input data into a latent space and generates a latent representation Z that captures the accessible chromatin regions and their regulatory roles in cells. This latent representation provides a foundation for further understanding the mechanisms regulating gene expression.

[0086] Step 2: Use Kullback Leibler (KL) divergence to ensure that the data of different omics are aligned in the shared latent space and minimize the distribution difference of the data of different omics in the latent space. Assume that the latent representations of single-cell transcriptome data scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq obtained by the encoder are Z RNA and Z ATAC , the goal of the model is to make the distributions of these two latent spaces as close as possible. The KL divergence loss function is as follows:

[0087] L KL =D KL (p RNA (Z RNA )||p fused (Z fused ))+D KL (p ATAC (Z ATAC )||p fused (Z fused ))

[0088] Among them, L KL represents the KL divergence loss, D KL represents KL divergence, p RNA and p ATAC Represent the potential distribution of single-cell transcriptome data and single-cell chromatin accessibility sequencing data, respectively, p fused represents the distribution of the target shared latent space.

[0089] By minimizing the KL divergence, the model optimizes data alignment in the latent space, ensuring that data from different omics have similar distribution characteristics in the shared latent space. The reconstruction task in the multi-omics variational autoencoder model scLTH is not just a simple reconstruction task, but a complex process that includes cross-omics data alignment.

[0090] During the reconstruction phase, the multi-omics variational autoencoder model not only reconstructs the single-cell transcriptome data scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq in each omics data, but also ensures that the reconstructed latent representation is aligned with the original data in the latent space. Specifically, the multi-omics variational autoencoder model scLTH optimizes the reconstruction through the following loss function:

[0091]

[0092] Among them, L reconstruction represents the reconstruction loss function, N i Represents the data dimension of mode i, x ij and x ij Represent the jth component of the original data and reconstructed data of mode i, ω ij Represents the weight of the jth feature in modality i, emphasizing the importance of certain features. This loss function is combined with the enhanced mean squared error (EMSE) to calculate the reconstruction error and align the consistency of the reconstructed representation with the original data in the latent space.

[0093] By minimizing these reconstruction errors, the multi-omics variational autoencoder model scLTH is able to fuse data from different omics while preserving the original cellular information and ensuring their consistency in the latent space.

[0094] Step 3: In the multi-omics variational autoencoder model scLTH, the teacher model is equipped with a six-layer Transformer architecture to mine shared features in multi-omics data under high-dimensional and complex data distributions. Due to its larger number of parameters and more complex structure, it can effectively capture cross-omics information. The teacher model receives data from single-cell transcriptome scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq, which are mapped to the latent space through their respective encoders. Each dataset is processed by its encoder to obtain a low-dimensional latent representation Z RNA and Z ATAC The teacher model does not directly connect these representations, but instead passes Z through a fuser module. RNA and Z ATAC Fused into a joint representation Z fused (latent representation), which captures the integrated latent features of the two data modalities. The fused representation is then input into the Transformer network of the teacher model to further explore the deep relationships between the data. This process can be represented as follows:

[0095] Z fused =Combiner(Z RNA ,Z ATAC )

[0096] Next, the fused latent representation Z fused Processed through the Transformer network of the teacher model to reveal cross-omics relationships. The latent representation Z of the input fused First, it is mapped to a new high-dimensional space through the embedding layer:

[0097]

[0098] The embedded representation is then processed through the self-attention mechanism. The self-attention mechanism can be expressed as:

[0099]

[0100] The output of the self-attention mechanism is:

[0101]

[0102] To enhance the expressiveness of the model, Transformer uses a multi-head attention mechanism. Multi-head attention calculates multiple attention heads in parallel, concatenates the results and performs a linear transformation:

[0103] A multi-head =Concat(Attention1,Attention2,...,Attention h )W out

[0104] The output of the multi-head attention is further processed by a feedforward neural network. The calculation of the feedforward network is:

[0105] F = ReLU(A multi-head W1+b1)W2+b2

[0106] After self-attention and feedforward neural network, the final output feature is expressed as:

[0107]

[0108] Output Passed to the next layer or output layer.

[0109] The teacher model finally generates output through a fully connected layer FCout:

[0110]

[0111] Among them, y teacher Represents the cross-omics features output by the teacher model, including all learned cross-omics feature representations, FCout represents the fully connected layer, Represents the features output by the self-attention mechanism and the feedforward neural network, F represents the output of the feedforward neural network, ReLU represents the ReLU activation function, and A multi-head Represents the concatenation result of multiple attention heads after linear transformation processing, W1 and W2 both represent the weights of the feedforward neural network, b1 and b2 both represent the bias of the feedforward neural network, Concat represents the concatenation operation, Attention h represents the output of the h-th attention head, W out Represents the linear transformation matrix, A represents the final output of the single-head attention in the self-attention mechanism, Attention represents the self-attention operation, Q, K and V represent Query, Key and Value respectively, and softmax represents the normalization function. represents the similarity between the query and the key, d k represents the dimension of the key vector, T represents transpose, Represents the potential representation Z fused The high-dimensional representation mapped to by the latent layer, W Q 、W k and W v Represent the weight matrices of query, key and value respectively, W embedding and b embedding They represent the weight matrix and bias term of the embedding layer respectively, and Combiner() represents the fusion operation, which includes the connection operation and a fully connected layer.

[0112] The teacher model extracts complex cross-omics features from high-dimensional data through the Transformer network and uses the self-attention mechanism to capture the shared features between single-cell transcriptome data scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq. Finally, the latent feature representation y generated by the teacher model teacher It is passed to the student model to help it learn more efficient cross-omics features, thereby improving the performance of multi-omics tasks.

[0113] Step 4: The student model learns deep features extracted from the teacher model in the shared latent space of multi-omics data and aligns these features as much as possible while maintaining low computational complexity. To achieve this goal, the student model adopts a lightweight Transformer architecture with far fewer parameters than the teacher model, significantly reducing computational complexity.

[0114] The architecture of the student model is similar to the teacher model, but smaller in size (fewer layers and attention heads) to reduce the number of parameters. Its goal is to remain efficient during inference while inheriting the shared features extracted by the teacher model. The student model receives the shared latent representation y passed from the teacher model. teacher , which contains shared features across omics. The student model further learns these shared features through a smaller Transformer network:

[0115] Z student =Transformer student (Z fused )

[0116] Due to the smaller number of parameters, the student model is able to learn the feature representations of the teacher model with lower computational complexity. This means that the student model can be trained with limited resources while maintaining similar performance to the teacher model.

[0117] The core of knowledge distillation is to effectively transfer the deep features learned by the teacher model to the student model. teacher By calculating the difference between the two models, the student model gradually approaches the deep feature representation extracted by the teacher model in the shared latent space.

[0118] The output y of the student model student and the output y of the teacher model teacher By minimizing the difference between the two, the student model learns similar features to the teacher model:

[0119] L distill =MSE(y student ,y teacher )

[0120] Among them, MSE represents the mean square error calculation, y student represents the output of the student model, y teacher A shared latent representation representing the output of the teacher model.

[0121] This process enables the student model to efficiently learn cross-omics features and align them to a shared latent space with minimal parameters. Through knowledge distillation, the student model inherits the strengths of the teacher model in cross-omics feature learning while achieving efficient feature alignment through a lightweight design. With minimal computational resources, the student model can process large-scale multi-omics data and achieve similar performance on tasks as the teacher model.

[0122] In this embodiment, step 5: After the above pre-training process, two encoders with good encoding capabilities are obtained, corresponding to the input data of two different omics, scRNA-seq and scATAC-seq, respectively, a fusion unit with feature fusion capabilities, and a student model capable of mining the potential representation of the data.

[0123] The fused latent representation y processed by the encoder, combiner, and student modules teacher As input to the classifier. The classifier is implemented as a fully connected layer that maps the output of the student module to the predicted cell type. During training, the cross entropy loss function is used to convert the cell type prediction probability y pred and the true label y true Compare to ensure accurate predictions:

[0124]

[0125] By minimizing this loss, the model learns to align predictions with the true cell type labels. Optimization and Alignment During training, the model parameters θ are optimized to minimize the overall loss function:

[0126]

[0127] Where θ* represents the minimum loss function of the multi-omics variational autoencoder model during training, θ represents the parameters of the multi-omics variational autoencoder model, and L CE represents the cross entropy loss, y pred represents the predicted probability of cell type, y true represents the true label, C represents the number of cell types, and y i' represents the true label, y i' represents the predicted probability of cell category i'.

[0128] This optimization ensures that the predicted labels are as close as possible to the true labels, thereby improving the classification accuracy. fusedIt captures transcriptional and epigenetic signatures, providing rich contextual information for cell type annotation.

[0129] At this point, the model training is complete, and then single-cell transcriptome data scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq can be simultaneously input for prediction.

[0130] In this embodiment, Figure 4 As shown in the figure, t-SNE plots of the three datasets are plotted, with the left panel colored according to the expert classification labels in the original experiment and the right panel colored according to the best prediction results of the multi-omics variational autoencoder model scLTH. Figure 5 As shown, the figure compares the clustering scores of the original data and the shared integrated data in the four datasets, green represents NMI, purple represents ARI, and yellow represents Purity.

[0131] Compared with the prior art, the advantages of the present invention are:

[0132] Deep fusion of multiple omics: Dedicated encoders are used to extract latent representations for single-cell transcriptome data scRNA-seq and single-cell chromatin accessibility sequencing data scATAC-seq, respectively. The data are then efficiently fused through the Combiner, effectively reducing the distribution differences between different omics data, thereby capturing key information on cellular transcription and epigenetic regulation.

[0133] Transformer module and self-attention mechanism: The multi-head self-attention mechanism using the Transformer structure can capture long-distance dependencies and complex nonlinear relationships in the fusion features, improving the model's ability to mine cross-omics information.

[0134] Teacher-student knowledge distillation framework: Build a high-parameter teacher model and efficiently transfer its deep features to a lightweight student model through knowledge distillation technology, achieving high prediction accuracy while significantly reducing computational complexity.

[0135] Pre-training and self-supervised learning strategy: Pre-training using variational autoencoders enables the model to automatically learn stable latent representations when faced with high noise and sparse data, laying a solid foundation for subsequent end-to-end training.

[0136] End-to-end joint optimization: The overall system adopts an end-to-end training strategy, and each module is collaboratively optimized, reducing the error accumulation caused by intermediate steps, and significantly improving the overall prediction performance and robustness.

Claims

1. A single-cell multi-omics cell type annotation method based on distribution and knowledge alignment, characterized by: The following steps are involved: S1. Obtain single-cell transcriptome data and single-cell chromatin accessibility sequencing data, and pre-train and train a multi-omics variational autoencoder model. The multi-omics variational autoencoder model combines variational autoencoders and knowledge distillation techniques to integrate and annotate multi-omics single-cell data through distribution and knowledge alignment. S2. Use the trained multi-omics variational autoencoder model to predict cell types based on the simultaneously input single-cell transcriptome data and single-cell chromatin accessibility sequencing data.

2. The single-cell multi-omics cell type annotation method based on distribution and knowledge alignment according to claim 1, characterized in that The pre-training process of the multi-omics variational autoencoder model is as follows: The single-cell transcriptome data and single-cell chromatin accessibility sequencing data are respectively passed through their respective encoders to extract the potential representation Z of the modal features. RNA and Z ATAC , and integrate the modal feature potential representation Z RNA and Z ATAC , get the fused potential representation Z fused ; Using KL divergence loss, we can ensure that the potential representation Z of different omics RNA and Z ATAC Align in a shared latent space to minimize the latent representation Z RNA and Z ATAC Distribution differences in the latent space, and the use of reconstruction loss to ensure that the reconstructed latent representation is aligned with the original data in the latent space; The fused latent representation Z fused The data is fed into the teacher model to extract cross-omics features from high-dimensional data. The self-attention mechanism is used to capture the shared features between single-cell transcriptome data and single-cell chromatin accessibility sequencing data. Use the student model to receive the shared potential representation y passed by the teacher model teacher , learning shared features, where the student model receives the shared latent representation y teacher Includes cross-omics features; The output of the student model is compared with the output of the teacher model, and the student model learns features similar to those of the teacher model by minimizing the difference between the two. distill , completing the pre-training of the multi-omics variational autoencoder model.

3. The single-cell multi-omics cell type annotation method based on distribution and knowledge alignment according to claim 2, characterized in that: The potential representation Z RNA Generated by the following formula: Z RNA =μ+σ·e μ,logσ 2 =Encoder(X RNA ) The potential representation Z ATAC Generated by the following formula: Z ATAC =μ+σ·e μ,logσ 2 =Encoder(X ATAC ) where μ and σ represent the mean and standard deviation of the latent representation of the encoder output, ε represents the noise sampled from the standard normal distribution, and logσ 2 Represents logarithmic variance, Encoder() represents encoding operation, X RNA represents the input single-cell transcriptome data, X ATAC Represents the input single-cell chromatin accessibility sequencing data.

4. The single-cell multi-omics cell type annotation method based on distribution and knowledge alignment according to claim 2, characterized in that The formula for the KL divergence loss is as follows: L KL =D KL (p RNA (Z RNA )||p fused (Z fused ))+D KL (p ATAC (Z ATAC )||p fused (Z fused )) Among them, L KL represents the KL divergence loss, D KL represents KL divergence, p RNA and p ATAC Represent the potential distribution of single-cell transcriptome data and single-cell chromatin accessibility sequencing data, respectively, p fused represents the distribution of the target shared latent space; The reconstruction loss formula is as follows: Among them, L reconstruction represents the reconstruction loss function, N i represents the data dimension of mode i, ω ij represents the weight of the jth feature in modality i, x ij and x ij represent the jth component of the original data and reconstructed data of mode i respectively.

5. According to the single-cell multi-omics cell type annotation method based on distribution and knowledge alignment according to claim 2, the formula of the cross-omics feature is as follows: <h2 style=";text-align:left;direction:ltr">F = ReLU(A<h2 style=";text-align:left;direction:ltr"> multi-head <h2 style=";text-align:left;direction:ltr"> W1+b1)W2+b2 A multi-head =Concat(Attention1,Attention2,...,Attention h )W out WITH fused =Combiner(Z RNA ,WITH ATAC ) in, y teacher Represents the cross-omics features output by the teacher model, including all learned cross-omics feature representations, FCout represents the fully connected layer, Represents the features output by the self-attention mechanism and the feedforward neural network, F represents the output of the feedforward neural network, ReLU represents the ReLU activation function, and A multi-head Represents the concatenation result of multiple attention heads after linear transformation processing, W1 and W2 both represent the weights of the feedforward neural network, b1 and b2 both represent the bias of the feedforward neural network, Concat represents the concatenation operation, Attention h represents the output of the h-th attention head, W out Represents the linear transformation matrix, A represents the final output of the single-head attention in the self-attention mechanism, Attention represents the self-attention operation, Q, K and V represent Query, Key and Value respectively, and softmax represents the normalization function. represents the similarity between the query and the key, d k represents the dimension of the key vector, T represents transpose, Represents the potential representation Z fused The high-dimensional representation mapped to by the latent layer, W Q 、W k and W v Represent the weight matrices of query, key and value respectively, W embedding and b embedding They represent the weight matrix and bias term of the embedding layer respectively, and Combiner() represents the fusion operation, which includes the connection operation and a fully connected layer.

6. The single-cell multi-omics cell type annotation method based on distribution and knowledge alignment according to claim 2, characterized in that: The feature L distill The formula is as follows: L distill =MSE(y student ,y teacher ) Among them, MSE represents the mean square error calculation, y student represents the output of the student model, y teacher A shared latent representation representing the output of the teacher model.

7. The single-cell multi-omics cell type annotation method based on distribution and knowledge alignment according to claim 2, characterized in that: The training process of the multi-omics variational autoencoder model includes: Based on the pre-training results, the fused latent representation processed by the encoder, the aggregator, and the student model is used as the input of the classifier; The output of the student model is mapped to the predicted cell type using a classifier. There are two encoders, corresponding to single-cell transcriptome data and single-cell chromatin accessibility sequencing data of different omics.

8. The single-cell multi-omics cell type annotation method based on distribution and knowledge alignment according to claim 7, characterized in that: The loss function formula of the multi-omics variational autoencoder model during training is as follows: Where θ* represents the minimum loss function of the multi-omics variational autoencoder model during training, θ represents the parameters of the multi-omics variational autoencoder model, and L CE represents the cross entropy loss, y pred represents the predicted probability of cell type, y true represents the true label, C represents the number of cell types, and y i' represents the true label, y i' represents the predicted probability of cell category i'.

Citation Information

Cited By

  • Single-cell multi-modal data integration method based on attention mechanism and graph variation auto-encoder

    CN121393546A

  • A single-cell multi-modal data integration method based on attention mechanism and graph variational autoencoder

    CN121393546B

  • Method and device for predicting non-coding mutation function based on single cell multi-omics and medium

    CN121583327A