Pairing-based single-cell multi-omics data integration method and system

CN116629123BActive Publication Date: 2026-08-11NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-25
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

发明人发现这类方法虽然可以解决更广泛的多组学联合嵌入问题,但是对于配对的多组学数据往往无法做到更好,因为其没有利用细胞标签的对应关系

Benefits of technology

[0051]本发明使得配对的单细胞多组学可以嵌入在同一特征空间下,并且尽可能地消除了原本数据中的批次效应;保留了原本配对单细胞多组学数据的细胞一一对应关系,使得在嵌入维度上,有同一细胞标签的不同组学的数据尽可能在低维空间下欧氏距离相近,保护了生物学信息,使得在嵌入维度下不同组学的相同细胞类型的细胞可以克服技术和批次效应聚集在一起。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116629123B_ABST
    Figure CN116629123B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of single-cell multi-omics analysis and provides a method and system for pair-based single-cell multi-omics data integration. The method includes acquiring paired single-cell multi-omics data and preprocessing it to obtain expression matrices for different omics. Based on these expression matrices, a pre-trained pseudo-Siamese neural network model is used to embed the expression matrices into the same dimensional space for data integration, resulting in integrated single-cell multi-omics data. During the training phase, different cell expression matrices are generated using different variational autoencoders based on the expression matrices of different omics. This data helps to obtain a better pre-trained Siamese neural network model. This invention eliminates the batch effect problem when performing paired cell joint embedding, preserves a large amount of biological information, makes the cell type distribution in the low-dimensional space more obvious, and maintains a high level of cell alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of single-cell multi-omics analysis technology, specifically relating to a pairwise single-cell multi-omics data integration method and system. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] In recent years, single-cell multi-omics sequencing technology refers to the technology of performing multiple omics measurements on the same cell. With the continuous development and improvement of this technology, it has overcome the problem that a single omics may not be able to accurately explain cell state and heterogeneity, and provides more refined molecular analysis at the cellular level. It has also become the data foundation for understanding the cellular function of organisms and exploring the regulatory mechanisms of organisms.

[0004] Many existing machine learning methods attempt to fully integrate multi-omics data through joint embedding, but most of them are unsupervised learning methods. The inventors found that while these methods can address a broader range of multi-omics joint embedding problems, they often fall short when dealing with paired multi-omics data because they don't utilize the correspondence between cell tags. With the continuous development of sequencing technology, paired single-cell multi-omics data will become increasingly abundant, and previous methods have limited effectiveness in processing this type of data, exhibiting generally poor overall performance in cell alignment, batch effect removal, and cell type-based clustering. Therefore, it is necessary to develop an integration method focused on processing paired single-cell multi-omics data to address these issues. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a pairwise single-cell multi-omics data integration method and system. This invention embeds single-cell multi-omics data jointly into the same feature space, while minimizing batch effects and preserving a significant amount of biological information. This provides data support for downstream multi-omics analysis.

[0006] According to some embodiments, the first aspect of the present invention provides a pairwise single-cell multi-omics data integration method, employing the following technical solution:

[0007] Pair-based single-cell multi-omics data integration methods include:

[0008] Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics;

[0009] Based on the expression matrices of different omics, a pre-trained pseudo-Twin neural network model is used to embed the expression matrices of different omics into the same dimensional space for data integration, resulting in integrated single-cell multi-omics data.

[0010] The training process of the pseudo-twin neural network model is as follows:

[0011] Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics;

[0012] Different cell expression matrices are generated using different pre-trained variational autoencoders for expression matrices of different omics.

[0013] By using a pseudo-twin neural network model, cell expression matrices and expression matrices from pairwise omics of different dimensions are embedded in the same dimensional space to obtain integrated single-cell multi-omics data;

[0014] The classification loss and triple loss are calculated in the embedding dimension to continuously optimize the pseudo-twin neural network model, resulting in a well-trained pseudo-twin neural network model.

[0015] Furthermore, the acquisition and preprocessing of paired single-cell multi-omics data to obtain expression matrices for different omics is specifically as follows:

[0016] Acquire paired single-cell multi-omics data;

[0017] Cells with high mitochondrial gene content, shallow cell count depth, and low gene expression data in different omics groups were filtered out.

[0018] The expression matrices of different omics were obtained.

[0019] Furthermore, for the expression matrices of different omics, different cell expression matrices are generated using different pre-trained variational autoencoders, specifically as follows:

[0020] The learned representation matrix is ​​encoded and mapped for the first time to obtain the encoded data;

[0021] By simultaneously performing two different secondary encodings on the first-encoded data, two secondary-encoded data are obtained;

[0022] Latent variables are obtained by sampling based on two secondary encoded data through reparameterization.

[0023] The latent variables are decoded twice to obtain the cell expression matrix.

[0024] Furthermore, the expression matrices of the different omics correspond to trained variational autoencoders with different parameters;

[0025] The variational autoencoders have the same structure, including a first encoder consisting of a fully connected layer, a first hidden layer, a second encoder consisting of two fully connected layers, a first decoder consisting of a fully connected layer, a second hidden layer, and a second decoder consisting of a fully connected layer.

[0026] Furthermore, by utilizing a pseudo-twin neural network model, cell expression matrices and expression matrices from different dimensions of pairwise omics are embedded in the same dimensional space to obtain integrated single-cell multi-omics data, including:

[0027] Based on the expression matrices and cell expression matrices of different omics, input triples of different omics are constructed;

[0028] By using a pseudo-twin neural network model, the input triples of pairwise omics from different dimensions are scaled to the same dimensional space;

[0029] By using common embedding units, the encoding results in the same dimensional space are embedded into the required common low-dimensional space to obtain integrated single-cell multi-omics data.

[0030] Given that the cell type corresponding to the input triple is known, a classifier consisting of fully connected layers is used to classify the cell type in this dimension in order to learn the features of the cell type.

[0031] Furthermore, the construction of input triples based on expression matrices and cell expression matrices from different omics is specifically as follows:

[0032] Anchor cells are selected from the row expression matrix of any row in the first omics;

[0033] Then, the positive example cells are selected from the row expression matrix in the second omics that corresponds one-to-one with the row labels of the anchor cells;

[0034] The negative example cells are selected from the row expression matrix in the second omics, which has completely different row labels from the anchor cells.

[0035] Based on anchor cells, positive example cells, and negative example cells, different omics input triples are constructed.

[0036] Furthermore, the structure of the pseudo-Twin neural network model includes two independent encoders that process two different omics data respectively, a common embedding unit, and a classifier composed of fully connected layers.

[0037] According to some embodiments, a second aspect of the present invention provides a pairwise single-cell multi-omics data integration system, employing the following technical solution:

[0038] Pair-based single-cell multi-omics data integration systems include:

[0039] The data acquisition module is configured to acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices of different omics.

[0040] The data integration module is configured to integrate the expression matrices of different omics using a pre-trained pseudo-Twin neural network model, embedding the expression matrices of different omics into the same dimensional space to obtain integrated single-cell multi-omics data.

[0041] The training process of the pseudo-twin neural network model is as follows:

[0042] Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics;

[0043] Different cell expression matrices are generated using different pre-trained variational autoencoders for expression matrices of different omics.

[0044] By using a pseudo-twin neural network model, cell expression matrices and expression matrices from pairwise omics of different dimensions are embedded in the same dimensional space to obtain integrated single-cell multi-omics data;

[0045] The classification loss and triple loss are calculated in the embedding dimension to continuously optimize the pseudo-twin neural network model, resulting in a well-trained pseudo-twin neural network model.

[0046] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium.

[0047] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the pairwise single-cell multi-omics data integration method described in the first aspect above.

[0048] According to some embodiments, a fourth aspect of the present invention provides a computer device.

[0049] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the pairwise single-cell multi-omics data integration method described in the first aspect above.

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0051] This invention enables paired single-cell multi-omics data to be embedded in the same feature space and eliminates batch effects in the original data as much as possible. It preserves the one-to-one correspondence between cells in the original paired single-cell multi-omics data, so that data from different omics with the same cell label are as close as possible in Euclidean distance in low-dimensional space in the embedding dimension, thus protecting biological information. It also enables cells of the same cell type from different omics to be clustered together in the embedding dimension, overcoming technical and batch effects. Attached Figure Description

[0052] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0053] Figure 1 This is a flowchart of a pairing-based single-cell multi-omics data integration method in an embodiment of the present invention;

[0054] Figure 2 This is a structural diagram of the variational autoencoder in an embodiment of the present invention;

[0055] Figure 3 This is a flowchart of the pseudo-twin neural network training process in an embodiment of the present invention. Detailed Implementation

[0056] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0057] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0058] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0059] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0060] Example 1

[0061] like Figure 1As shown, this embodiment provides a pairing-based single-cell multi-omics data integration method. This embodiment uses the application of this method to a server as an example for illustration. It is understood that this method can also be applied to terminals, and can also be applied to systems including terminals, servers, and other components, and is implemented through interaction between the terminal and the server. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communication, middleware services, domain name services, CDN security services, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein. In this embodiment, the method includes the following steps:

[0062] Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics;

[0063] Based on the expression matrices of different omics, a pre-trained pseudo-Twin neural network model is used to embed the expression matrices of different omics into the same dimensional space for data integration, resulting in integrated single-cell multi-omics data.

[0064] The training process of the pseudo-twin neural network model is as follows:

[0065] Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics;

[0066] Different cell expression matrices are generated using different pre-trained variational autoencoders for expression matrices of different omics.

[0067] By using a pseudo-twin neural network model, cell expression matrices and expression matrices from pairwise omics of different dimensions are embedded in the same dimensional space to obtain integrated single-cell multi-omics data;

[0068] The classification loss and triple loss are calculated in the embedding dimension to continuously optimize the pseudo-twin neural network model, resulting in a well-trained pseudo-twin neural network model.

[0069] This embodiment provides a method for jointly embedding and integrating paired single-cell multi-omics data. This method can be applied to various paired single-cell multi-omics data. It can obtain more diverse and referable high-quality cell expression matrices through variational autoencoders. Then, by using these generated expression matrices and the original expression matrices to construct triples, a pre-trained pseudo-Siamese neural network is obtained. The pre-trained pseudo-Siamese neural network embeds data from different omics into the same dimension, thereby providing data support for subsequent downstream analysis.

[0070] like Figure 3 As shown, the training and testing process of the pseudo-Siamese neural network model includes the following steps:

[0071] Step 1: Perform quality control and preprocessing on the paired single-cell multi-omics datasets, including:

[0072] Cells with high mitochondrial gene content, shallow cell count depth, and low gene expression levels are filtered out to screen for genes with high expression levels.

[0073] This method employs different screening strategies for different omics data.

[0074] In this embodiment, bone marrow mononuclear cell data from 12 healthy human donors in the Cite-seq dataset were selected. This dataset consists of paired multi-omics data, comprising 90,261 cells and 13,953 genes, and 90,261 cells and 134 proteins. Initial data screening was performed, using data from s1d1, s1d3, s2d1, s2d4, s2d5, s3d1, and s3d6 as the training set. For the transcriptome data, cell quality control was performed based on mitochondrial gene content, cell count depth, and gene expression levels, ultimately resulting in 44,277 cells. High-expression genes were then selected, yielding an expression matrix X containing 2,175 gene features. For the proteome data, the same screening results as for the transcriptome data were used for cell selection, but protein features were not further filtered, resulting in the final expression matrix Y.

[0075] Step 2: Construct a variational autoencoder to generate a high-quality, referable representation matrix, including:

[0076] Different variational autoencoder networks are constructed for different omics to learn the latent variable distribution of the expression matrix of different omics datasets;

[0077] By randomly sampling latent variables, high-quality cellular omics reference data similar to the input can be obtained.

[0078] like Figure 2As shown, in this embodiment, different variational autoencoder models are designed for the two different omics, with the inputs being the preprocessed single-cell multi-omics data obtained in step 1. Assume the expression matrices of the two omics are X and Y, respectively. Taking expression matrix X as an example, its autoencoder network is constructed. First, an encoder consisting of fully connected layers is designed based on the number of genes in expression matrix X, mapping the input to the hidden layer H. x Then, two encoders consisting of fully connected layers are constructed to obtain μ. x and σ x Finally, the latent variable Z is obtained by sampling through reparameterization. x Then, the hidden layer H′ is mapped through a decoder consisting of a fully connected layer. x Finally, a decoder consisting of fully connected layers generates the final high-quality cell expression matrix X′.

[0079] The formula for the variational autoencoder is as follows:

[0080] H x =LeakyRelu(XW H )

[0081] μ x =H x W μ

[0082] σ x =H x W σ

[0083] Z x =μ x +εσ x , ε∈Norm(0,1)

[0084]

[0085] X′=LeakyRelu(H′ x W X ,)

[0086] Among them, W * Z represents the weight parameters of different fully connected layers. The network uses the LeakyReLU function as the activation function, which increases the convergence speed and prevents gradient vanishing. x The embedding dimension is selected as 32-dimensional.

[0087] Another omics approach also requires the design of an independent variational autoencoder network, with similar formulas and embedding dimensions as described above, but learning different parameters.

[0088] The loss function of a variational autoencoder can be expressed as:

[0089]

[0090] The first term is the reconstruction loss, which measures the difference between the generated representation matrix and the input representation matrix. The second term is the KL loss, which iteratively fits the posterior probability p(Z). x |X), thereby learning the embedding Z of input X. x .

[0091] Step 3: Generate a referenceable, high-quality representation matrix using the trained variational autoencoder. Specifically, for each omics representation matrix, train its unique variational autoencoder model, setting the training epochs to 100, the learning rate to 10⁻⁴, the dimension of the latent variables to 32, and the dimension of the hidden layers to 1024 for the X representation matrix and 64 for the Y representation matrix. After training, referenceable representation matrices for both omics are obtained.

[0092] Step 4: Construct the input triples for the pseudo-twin neural network.

[0093] The data comes from quality-controlled and preprocessed expression matrices and high-quality cell expression matrices generated by a variational autoencoder network.

[0094] When any row in the first omics is used as the anchor cell, the row expression matrix X is obtained. anchor In this case, the positive example should be selected as the row expression matrix Y in the second omics that corresponds one-to-one with the row label of the anchor cell. positive The negative example is chosen to be the same as X. anchor The cell row labels maintain a completely different row expression matrix Y in the second omics. negative Similarly, using any row in the second omics dataset as the anchor cell, we obtain the row expression matrix Y. anchor If the construction idea is consistent with the above, then the positive example cell should be selected as the first row expression matrix X in the omics that corresponds one-to-one with the row label of the anchor cell. positive The negative example is selected as the one with Y. anchor Cell row labels remain completely different in the first omics row expression matrix X negative .

[0095] It is understandable that in the first and second omics studies, each row of data represents the expression matrix of a cell, i.e., the row expression matrix. When selecting input triples, the selection is done on a row-by-row basis.

[0096] Step 5: Construct a pseudo-Siamese neural network to embed the two omics datasets into the same low-dimensional space for subsequent downstream analysis. This learns the joint embedding space of single-cell multi-omics, bridging the gap between different omics data of the same cell label in low-dimensional space. The construction method is as follows:

[0097] Two independent encoders are constructed, each processing two different omics datasets, and the input data is scaled by the encoders to the same hidden layer dimension.

[0098] Then, the encoded result is embedded into the required common low-dimensional space by using a fully connected layer with shared weights between the two omics as a common embedding unit.

[0099] The cell types of the input anchor and positive examples are obtained. Given that the cell types corresponding to the input expression matrix are known, a classifier consisting of fully connected layers is finally constructed to classify cell types in this dimension in order to learn the features of cell types.

[0100] The loss function used in the Siamese neural network is:

[0101]

[0102] loss2=γ×d(a,p)+max(d(a,p)-d(a,n)+margin,0)γ∈(0,1)

[0103] loss=loss1+β×loss2,β∈(0,1)

[0104] loss1 is the cross-entropy loss function, where M is the total number of cell types, and X... ic The function is a sign function; it is 1 if cell type i matches cell type c, otherwise it is 0. ic This represents the probability that cell i belongs to cell type c.

[0105] `loss2` represents the improved triple loss function. Here, `a` represents the anchor, `p` represents the positive example, and `n` represents the negative example, expressing the matrix representation in the common low-dimensional space. `d(a, p)` represents the Euclidean distance between the anchor and the positive example in the embedding low-dimensional space; the other distances are similarly represented. `margin` represents the boundary value, and `γ` and `β` represent scaling factors.

[0106] Compared to the original triple loss function, loss2 adds a loss term d(a, p), avoiding the hard truncation problem. This allows the distance between the anchor and the positive example to continue to narrow even when the original loss function is 0. This loss function ensures that the distance between the same cell from different omics in the low-dimensional space is narrowed, while distancing the distance between different cells from different omics in the low-dimensional space. This aligns with the assumption of multi-omics joint embedding, allowing cells with similar biological states to cluster together as much as possible.

[0107] Step 6: Based on the trained pseudo-Siamese neural network model, perform joint embedding on the test set consisting of s1d2 and s3d7, which were not involved in the model. The model input needs to undergo the same preprocessing as the test set, which yields the embedding matrices of the transcriptome and proteome in low-dimensional space. It's understandable that the embedding matrices of the transcriptome and proteome mentioned here mean that their dimensions become consistent after passing through the Siamese neural network, resulting in a common low-dimensional dataset. This is because the expression matrices of the transcriptome and proteome have different dimensions if they are not processed through the Siamese neural network. The low-dimensional embedding matrices of the transcriptome and proteome are obtained by jointly embedding data from two different omics.

[0108] By jointly embedding the transcriptome and proteome, the batch effect inherent in the original data was eliminated, and a high level of cell alignment was ensured. Simultaneously, it allowed cells of the same type to cluster in the reduced-dimensional space, resulting in a high average contour width. This high-quality embedding matrix provides a data foundation for subsequent downstream analysis.

[0109] This embodiment discloses a method for integrating paired single-cell multi-omics data, addressing the issues of eliminating batch effects in bioinformatics and performing multi-omics joint embedding. It includes two key steps: first, using a generative model to create more high-quality cell omics reference data to augment the training data; and second, training a Siamese neural network using an improved triple loss function for multi-omics joint embedding. This eliminates the batch effect problem in single-cell multi-omics data and encourages cells with similar biological states to move closer together. This embodiment eliminates the batch effect problem between different batches of data during paired cell joint embedding, preserves a large amount of biological information, makes the cell type distribution in low-dimensional space more obvious, and maintains a high level of cell alignment.

[0110] Example 2

[0111] This embodiment provides a pairwise single-cell multi-omics data integration system, including:

[0112] The data acquisition module is configured to acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices of different omics.

[0113] The data integration module is configured to integrate the expression matrices of different omics using a pre-trained pseudo-Twin neural network model, embedding the expression matrices of different omics into the same dimensional space to obtain integrated single-cell multi-omics data.

[0114] The training process of the pseudo-twin neural network model is as follows:

[0115] Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics;

[0116] Different cell expression matrices are generated using different pre-trained variational autoencoders for expression matrices of different omics.

[0117] By using a pseudo-twin neural network model, cell expression matrices and expression matrices from pairwise omics of different dimensions are embedded in the same dimensional space to obtain integrated single-cell multi-omics data;

[0118] The classification loss and triple loss are calculated in the embedding dimension to continuously optimize the pseudo-twin neural network model, resulting in a well-trained pseudo-twin neural network model.

[0119] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0120] The descriptions of each embodiment in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0121] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and the division of modules described above is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0122] Example 3

[0123] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the pairwise single-cell multi-omics data integration method described in Embodiment 1 above.

[0124] Example 4

[0125] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the pairwise single-cell multi-omics data integration method described in Embodiment 1 above.

[0126] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0127] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0131] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for integrating paired single-cell multi-omics data, characterized in that, include: Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics; Based on the expression matrices of different omics, a pre-trained pseudo-Twin neural network model is used to embed the expression matrices of different omics into the same dimensional space for data integration, resulting in integrated single-cell multi-omics data. The training process of the pseudo-twin neural network model is as follows: Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics; Different cell expression matrices are generated using different pre-trained variational autoencoders for expression matrices of different omics. Using a pseudo-twin neural network model, cell expression matrices and expression matrices from pairwise omics of different dimensions are embedded in the same dimensional space to obtain integrated single-cell multi-omics data, including: Based on the expression matrices and cell expression matrices of different omics, input triples of different omics are constructed; By using a pseudo-twin neural network model, the input triples of pairwise omics from different dimensions are scaled to the same dimensional space; By using common embedding units, the encoding results in the same dimensional space are embedded into the required common low-dimensional space to obtain integrated single-cell multi-omics data. Given that the cell type corresponding to the input triple is known, a classifier consisting of fully connected layers is used to classify the cell type in this dimension in order to learn the features of the cell type; The classification loss and triple loss are calculated in the embedding dimension to continuously optimize the pseudo-twin neural network model, resulting in a well-trained pseudo-twin neural network model.

2. The method for pairing-based single-cell multi-omics data integration as described in claim 1, characterized in that, The process of acquiring paired single-cell multi-omics data and preprocessing it to obtain expression matrices for different omics is as follows: Acquire paired single-cell multi-omics data; Cells with high mitochondrial gene content, shallow cell count depth, and low gene expression data in different omics groups were filtered out. The expression matrices of different omics were obtained.

3. The method for pairing-based single-cell multi-omics data integration as described in claim 1, characterized in that, The expression matrices for different omics are generated using different pre-trained variational autoencoders to produce different cell expression matrices, specifically as follows: The learned representation matrix is ​​encoded and mapped for the first time to obtain the encoded data; By simultaneously performing two different secondary encodings on the first-encoded data, two secondary-encoded data are obtained; Latent variables are obtained by sampling based on two secondary encoded data through reparameterization. The latent variables are decoded twice to obtain the cell expression matrix.

4. The method for pairing-based single-cell multi-omics data integration as described in claim 3, characterized in that, The expression matrices of the different omics correspond to the trained variational autoencoders with different parameters; The variational autoencoders have the same structure, including a first encoder consisting of a fully connected layer, a first hidden layer, a second encoder consisting of two fully connected layers, a first decoder consisting of a fully connected layer, a second hidden layer, and a second decoder consisting of a fully connected layer.

5. The method for pairing-based single-cell multi-omics data integration as described in claim 1, characterized in that, The construction of input triples based on expression matrices and cell expression matrices from different omics is as follows: Anchor cells are selected from the row expression matrix of any row in the first omics; Then, the positive example cells are selected from the row expression matrix in the second omics that corresponds one-to-one with the row labels of the anchor cells; The negative example cells are selected from the row expression matrix in the second omics, which has completely different row labels from the anchor cells. Based on the input anchor cells, positive example cells, and negative example cells, different omics input triplets are constructed.

6. The method for pairing-based single-cell multi-omics data integration as described in claim 1, characterized in that, The structure of the pseudo-twin neural network model includes two independent encoders that process two different omics data respectively, a common embedding unit, and a classifier composed of fully connected layers.

7. A pair-based single-cell multi-omics data integration system, characterized in that, include: The data acquisition module is configured to acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices of different omics. The data integration module is configured to integrate the expression matrices of different omics using a pre-trained pseudo-Twin neural network model, embedding the expression matrices of different omics into the same dimensional space to obtain integrated single-cell multi-omics data. The training process of the pseudo-twin neural network model is as follows: Acquire paired single-cell multi-omics data and preprocess them to obtain expression matrices for different omics; Different cell expression matrices are generated using different pre-trained variational autoencoders for expression matrices of different omics. Using a pseudo-twin neural network model, cell expression matrices and expression matrices from pairwise omics of different dimensions are embedded in the same dimensional space to obtain integrated single-cell multi-omics data, including: Based on the expression matrices and cell expression matrices of different omics, input triples of different omics are constructed; By using a pseudo-twin neural network model, the input triples of pairwise omics from different dimensions are scaled to the same dimensional space; By using common embedding units, the encoding results in the same dimensional space are embedded into the required common low-dimensional space to obtain integrated single-cell multi-omics data. Given that the cell type corresponding to the input triple is known, a classifier consisting of fully connected layers is used to classify the cell type in this dimension in order to learn the features of the cell type; The classification loss and triple loss are calculated in the embedding dimension to continuously optimize the pseudo-twin neural network model, resulting in a well-trained pseudo-twin neural network model.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps in the pair-based single-cell multi-omics data integration method as described in any one of claims 1-6.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the pairwise single-cell multi-omics data integration method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Domain generalization and domain self-adaption blood cell classification method based on feature decoupling

    CN115937566A

  • Method and system for deconvolution of bulk RNA-sequencing data

    WO2023025419A1