Single-cell RNA-seq and ATAC-seq data integration method

Through imbalanced optimal transmission algorithm and transfer learning strategy, aligning single-cell RNA-seq and ATAC-seq data in the shared embedding space, the integration distortion problem caused by distribution imbalance is solved, efficient data integration and cell type tag alignment are achieved, and the accuracy and reliability of integration are improved.

CN120564833APending Publication Date: 2025-08-29HARBIN INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510639096.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

When integrating single-cell RNA-seq and ATAC-seq data, there are biologically related signal distortion problems caused by distribution imbalance and technical noise, which affects the accuracy and reliability of data integration.

Method used

The unbalanced optimal transmission algorithm and transfer learning strategy are adopted to realize the joint representation of different annotation data in the shared embedding space by building an encoder network, and supervised learning and alignment are performed in the embedding space and label space, combining the unbalanced optimal transmission algorithm to align scRNA-seq and scATAC-seq data.

Benefits of technology

It effectively solves the integration distortion problem caused by cell type distribution imbalance and technical deviation across modal data, realizes cell granularity alignment, and improves the accuracy and reliability of data integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564833A_ABST
    Figure CN120564833A_ABST
Patent Text Reader

Abstract

The method disclosed by the invention comprises the following steps: S1, collecting a gene expression matrix of scRNA-seq, a gene activity matrix of scATAC-seq and cell type annotation of the scRNA-seq; s2, the scRNA-seq data and the scATAC-seq data are subjected to preprocessing, and a gene shared by the scRNA-seq data and the scATAC-seq data is obtained; s3, constructing an encoder network, and realizing joint representation of different omics data in the shared embedding space; s4, performing supervised learning guidance on the embedded space by adopting a cell type label in the scRNA-seq data; s5, in the embedding space and the label space, performing alignment on the scRNA-seq data and the scATAC-seq data by adopting an unbalanced optimal transmission algorithm; and S6, according to the matching probability matrix, endowing each cell in the scATAC-seq data with a cell type annotation, and realizing integration of the scRNA-seq data and the scATAC-seq data. The problem that the accuracy and reliability of data integration are affected due to distortion of biological related signals caused by application of an existing OT frame to single cell data integration is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biological cell data integration, and specifically relates to a method for integrating single-cell RNA-seq and ATAC-seq data. Background Art

[0002] In recent years, the rapid development of single-cell multi-omics technologies has provided unprecedented high-resolution data resources for analyzing cellular heterogeneity, revealing gene regulatory mechanisms, and exploring the molecular basis of disease development. However, how to efficiently integrate single-cell data across modalities, batches, and experimental conditions has become a key challenge in the field of bioinformatics. Traditional data integration methods, such as canonical correlation analysis (CCA), mutual nearest neighbor (MNN), or deep learning-based cross-modal embedding models, usually assume balanced sample distributions or similar latent feature spaces between different datasets or modalities. However, in practical applications, single-cell data often suffer from significant distribution shifts, technical noise, and sparsity, which limit the effectiveness of existing integration methods.

[0003] To address these issues, optimal transport theory (OT) has been increasingly applied to single-cell data integration due to its mathematical rigor in measuring distribution differences and implementing probability mass mapping. However, traditional OT frameworks require strict conservation of the total mass between the source and target domains, while single-cell multi-omics data often exhibit significant distribution imbalances, such as differences in cell type ratios between modalities or mismatches in cell numbers between batches. Direct application of OT can distort biologically relevant signals, thereby compromising the accuracy and reliability of data integration. Summary of the Invention

[0004] To address the problem that the application of existing OT frameworks to single-cell data integration leads to distortion of biologically relevant signals, thereby affecting the accuracy and reliability of data integration, the present invention provides a single-cell RNA-seq and ATAC-seq data integration method.

[0005] In order to achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0006] A method for integrating single-cell RNA-seq and ATAC-seq data, comprising the steps of:

[0007] S1. Collect gene expression matrix for scRNA-seq, gene activity matrix for scATAC-seq, and cell type annotation for scRNA-seq;

[0008] S2. Preprocess scRNA-seq and scATAC-seq data to obtain genes shared by scRNA-seq and scATAC-seq;

[0009] S3. Build an encoder network to achieve joint representation of different omics data in a shared embedding space;

[0010] S4. Supervised learning guides the embedding space using cell type labels from scRNA-seq data.

[0011] S5. Align scRNA-seq and scATAC-seq data using the unbalanced optimal transfer algorithm in both embedding and label spaces.

[0012] S6. Assign cell type annotations to each cell in the scATAC-seq data based on the matching probability matrix, and integrate scRNA-seq data with scATAC-seq data.

[0013] Furthermore, each row of the gene expression matrix of scRNA-seq and the gene activity matrix of scATAC-seq represents a cell, and each column represents a gene.

[0014] Furthermore, an encoder is jointly trained by scRNA-seq data and scATAC-seq data to learn a shared low-dimensional embedding space by weight sharing. The small batch B0 of training data is constructed by sampling cell subsets of equal size from each dataset, i.e. Each subset B (s) and B (t) All have B cells;

[0015] Low-dimensional orthogonal features are captured when projecting each data batch into the embedding space. The encoding loss measures the deviation of the embedded representation from the identity matrix, and the encoder is optimized by minimizing the encoding loss.

[0016] Furthermore, dimensionality reduction is performed on each RNA dataset, and its encoding loss is expressed as:

[0017]

[0018] Among them, B (s) represents the cell subset sampled from the scRNA-seq dataset, B is the number of cells in the current subset, D is the dimension of the embedding space, and θ is the parameter set of the encoder network. represents the embedding vector generated by the b-th scRNA-seq sample under the parameter θ, represents the value of the embedding vector in the jth dimension, Represents the mean of all scRNA-seq samples in this dimension, represents the covariance between the i-th and j-th embedding dimensions in the RNA modality, i,j∈{1,…,D} represents the dimension index, and |·| represents the absolute value, which is used to unify the measurement deviation or correlation size.

[0019] The encoding loss represented by ATAC embedding is the same as that of RNA. Let ATAC embedding be Its dimensionality reduction loss is expressed as:

[0020]

[0021] Among them, B (t) represents the cell subset sampled from the scATAC-seq dataset, B is the number of cells in the current subset, D is the dimension of the embedding space, θ is the parameter set of the encoder network, represents the embedding vector generated by the b-th scATAC-seq sample under parameter θ, represents the value of the embedding vector in the jth dimension, represents the mean value of all scATAC-seq samples on this dimension, represents the covariance between the i-th and j-th embedding dimensions in the ATAC mode, i,j∈{1,…,D} represents the dimension index, and |·| represents the absolute value, which is used to unify the measurement deviation or correlation size.

[0022] In this loss function, the first term (the inverse) encourages each coordinate dimension to have a larger intra-sample variability, the second term penalizes the correlation between different dimensions to enhance the orthogonality of the space, and the third term prevents the mean drift of the embedding space and improves the stability of the model.

[0023] Finally, the encoder loss is the sum of the RNA encoding loss and the ATAC encoding loss, which uniformly constrains the embedding space features of different modalities:

[0024]

[0025] Furthermore, in step S4, after softmax transformation, the classifier network outputs a probability vector for cell type prediction.

[0026] Each with cell type annotation B (s) , the classification loss is used to supervise the training of scRNA-seq dataset, defined as the cross entropy loss, and the formula is expressed as:

[0027]

[0028] Among them, 1(.) represents the indicator function, represents the predicted probability that the bth cell belongs to the cth class.

[0029] Furthermore, the imbalanced optimal transfer algorithm is used to align RNA and ATAC embeddings with predicted labels, and the cost matrix It represents the distance between the source distribution a and the target distribution b, and is defined as:

[0030] M=αM embed +λ t M sce

[0031] Among them, M embed =||f RNA -f ATAC || 2 Represents the Euclidean distance matrix between RNA and ATAC embeddings, which is used to measure the feature similarity of the two datasets in the embedding space. represents the cross-entropy-based label alignment matrix used to align RNA cell type labels with the probability distribution predicted by the ATAC model, and α and are hyperparameters used to tune M embed and M sce Weights in the cost matrix M, optimal transmission matrix It is used to describe the transmission between the source distribution and the target distribution. The calculation formula is as follows:

[0032]

[0033] in, and are the source and target distributions, is the regularized reference distribution, reg m1 and reg m2 Represents the regularization term, which is used to prevent the model from overfitting, <γ,M> F Used to measure transmission plan loss, reg m1 KL(γ1||a) and reg m2 KL(γ T 1||b) is the marginal constraint term for the transfer matrix γ, which measures the difference between the row marginal distribution and column marginal distribution of γ and a and b respectively. reg·KL(γ||c) is the regularization term, which represents the KL difference between the transfer matrix γ and the reference distribution c.

[0034] The constraints are:

[0035]

[0036] That is, all elements in the transfer matrix are non-negative;

[0037] scUOTL uses KL divergence to update the transfer matrix γ and calculate the error until convergence. The KL divergence update formula is:

[0038]

[0039] Among them, γ represents the current transfer matrix, which is used to describe the correspondence between source domain samples and target domain samples; γ i,1 represents the element in row i and column 1, γ j,0 represents the element in the jth row and the 0th column, which are used to measure the marginal distribution of the current transmission matrix in two directions; r1 and r2 are hyperparameters that control the weight of the KL divergence regularization term, which are used to control the importance of the distribution at both ends in the optimization; γ d is a normalization coefficient that controls the marginal distribution power and adds a small numerical stability term 1×10 -16 , to avoid division by zero in subsequent calculations.

[0040] In the update of the transfer matrix, K is the kernel matrix, which is usually calculated by the cost matrix C and the regularization coefficient ∈, and its form is K = exp(-C / ∈); Represents an element-by-element power operation on each element of the transfer matrix, with a power of r1+r2; the final updated γ is obtained by element-by-element multiplication with K and division by the normalization factor γ d This multiplicative update method has good numerical stability and convergence characteristics in KL divergence optimization.

[0041] The final loss function of the model is the weighted sum of the above losses, optimizing the overall performance of the model:

[0042]

[0043] Where λ1 and λ2 are weights of different loss terms. The final loss combines different loss terms to balance the embedding dimensionality reduction, classification task, and cross-dataset alignment.

[0044] Furthermore, after obtaining the transfer matrix in step S5, the best matching RNA cell is found for each ATAC cell:

[0045] RNA label(i) =argmax j T ij

[0046] Assign ATAC cell i to its best matching RNA cell j, and use the label of this RNA cell as the prediction result.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] This paper integrates single-cell RNA-seq and ATAC-seq data by combining an unbalanced optimal transfer algorithm and a transfer learning strategy. This innovative approach incorporates unbalanced optimal transfer into end-to-end neural network training, achieving alignment at the cell-level. A dual alignment mechanism aligns the embedding representations and cell type labels of scATAC-seq and scRNA-seq in both embedding and label spaces. By relaxing the strict mass conservation constraints of traditional over-the-air (OT) training, this approach effectively addresses cross-modal data integration artifacts caused by imbalanced cell type distributions and technical bias. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is an overall flow chart of an embodiment of the present invention;

[0050] Figure 2 This is a schematic diagram of UMAP visualization of joint dimensionality reduction in an embodiment of the present invention, colored by cell type;

[0051] Figure 3 This is a schematic diagram of UMAP visualization of joint dimensionality reduction in an embodiment of the present invention, colored by omics;

[0052] Figure 4 Schematic diagram of the comparison of experimental results of the method of the present application and five comparative algorithms in the embodiment of the present invention;

[0053] Figure 5 In the embodiment of the present invention Figure 4 Corresponding data schematic diagram;

[0054] Figure 6 In the embodiment of the present invention, the three cases of balanced matching, RNA imbalance matching and ATAC imbalance matching are shown. Figure 5 Schematic diagram of accuracy, F1 score, ARI and NMI total scores of the compared methods;

[0055] Figure 7 In the embodiment of the present invention, the three cases of balanced matching, RNA imbalance matching and ATAC imbalance matching are shown. Figure 5 Schematic diagram of the F1 silhouette coefficient scores of the compared methods. DETAILED DESCRIPTION

[0056] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and drawings. The contents mentioned in the embodiments are not intended to limit the present invention.

[0057] like Figure 1 As shown, this embodiment provides a method for integrating single-cell RNA-seq and ATAC-seq data, including the steps of:

[0058] S1. Collect gene expression matrix for scRNA-seq, gene activity matrix for scATAC-seq, and cell type annotation for scRNA-seq;

[0059] S2. Preprocess scRNA-seq and scATAC-seq data to obtain genes shared by scRNA-seq and scATAC-seq;

[0060] S3. Build an encoder network to achieve joint representation of different omics data in a shared embedding space;

[0061] S4. Supervised learning guides the embedding space using cell type labels from scRNA-seq data.

[0062] S5. Align scRNA-seq and scATAC-seq data using the unbalanced optimal transfer algorithm in both embedding and label spaces.

[0063] S6. Assign cell type annotations to each cell in the scATAC-seq data based on the matching probability matrix, and integrate scRNA-seq data with scATAC-seq data.

[0064] The gene expression matrix for scRNA-seq and the gene activity matrix for scATAC-seq each row represents a cell and each column represents a gene.

[0065] An encoder is trained jointly by scRNA-seq data and scATAC-seq data, and a shared low-dimensional embedding space is learned by weight sharing. The small batch B0 of training data is constructed by sampling cell subsets of equal size from each dataset, that is, Each subset B (s) and B (t) All have B cells;

[0066] Low-dimensional orthogonal features are captured when projecting each data batch into the embedding space. The encoding loss measures the deviation of the embedded representation from the identity matrix, and the encoder is optimized by minimizing the encoding loss.

[0067] The dimension reduction is performed on each RNA dataset, and its encoding loss is expressed as:

[0068]

[0069] Among them, B (s) represents the cell subset sampled from the scRNA-seq dataset, B is the number of cells in the current subset, D is the dimension of the embedding space, and θ is the parameter set of the encoder network. represents the embedding vector generated by the b-th scRNA-seq sample under the parameter θ, represents the value of the embedding vector in the jth dimension, Represents the mean of all scRNA-seq samples in this dimension, represents the covariance between the i-th and j-th embedding dimensions in the RNA modality, i,j∈{1,…,D} represents the dimension index, and |·| represents the absolute value, which is used to unify the measurement deviation or correlation size.

[0070] The encoding loss represented by ATAC embedding is the same as that of RNA. Let ATAC embedding be Its dimensionality reduction loss is expressed as:

[0071]

[0072] Among them, B (t) represents the cell subset sampled from the scATAC-seq dataset, B is the number of cells in the current subset, D is the dimension of the embedding space, θ is the parameter set of the encoder network, represents the embedding vector generated by the b-th scATAC-seq sample under parameter θ, represents the value of the embedding vector in the jth dimension, represents the mean value of all scATAC-seq samples on this dimension, represents the covariance between the i-th and j-th embedding dimensions in the ATAC mode, i,j∈{1,…,D} represents the dimension index, and |·| represents the absolute value, which is used to unify the measurement deviation or correlation size.

[0073] In this loss function, the first term (the inverse) encourages each coordinate dimension to have a larger intra-sample variability, the second term penalizes the correlation between different dimensions to enhance the orthogonality of the space, and the third term prevents the mean drift of the embedding space and improves the stability of the model.

[0074] Finally, the encoder loss is the sum of the RNA encoding loss and the ATAC encoding loss, which uniformly constrains the embedding space features of different modalities:

[0075]

[0076] In step S4, after softmax transformation, the classifier network outputs a probability vector for cell type prediction.

[0077] Each with cell type annotation B (s) , the classification loss is used to supervise the training of scRNA-seq dataset, defined as the cross entropy loss, and the formula is expressed as:

[0078]

[0079] Among them, 1(.) represents the indicator function, represents the predicted probability that the bth cell belongs to the cth class.

[0080] Alignment between RNA and ATAC embeddings and predicted labels using the imbalanced optimal transfer algorithm, cost matrix It represents the distance between the source distribution a and the target distribution b, and is defined as:

[0081] M=αM embed +λ t M sce

[0082] Among them, M embed =||f RNA -f ATAC || 2 Represents the Euclidean distance matrix between RNA and ATAC embeddings, which is used to measure the feature similarity of the two datasets in the embedding space. represents the cross-entropy-based label alignment matrix used to align RNA cell type labels with the probability distribution predicted by the ATAC model, and α and are hyperparameters used to tune M embed and M sce Weights in the cost matrix M, optimal transmission matrix It is used to describe the transmission between the source distribution and the target distribution. The calculation formula is as follows:

[0083]

[0084] in, and are the source and target distributions, is the regularized reference distribution, reg m1 and reg m2 Represents the regularization term, which is used to prevent the model from overfitting, <γ,M> F Used to measure transmission plan loss, reg m1 KL(γ1||a) and reg m2 KL(γ T 1||b) is the marginal constraint term for the transfer matrix γ, which measures the difference between the row marginal distribution and column marginal distribution of γ and a and b respectively. reg·KL(γ||c) is the regularization term, which represents the KL difference between the transfer matrix γ and the reference distribution c.

[0085] The constraints are:

[0086]

[0087] That is, all elements in the transfer matrix are non-negative;

[0088] scUOTL uses KL divergence to update the transfer matrix γ and calculate the error until convergence. The KL divergence update formula is:

[0089]

[0090] Among them, γ represents the current transfer matrix, which is used to describe the correspondence between source domain samples and target domain samples; γ i,1 represents the element in row i and column 1, γ j,0 represents the element in the jth row and the 0th column, which are used to measure the marginal distribution of the current transmission matrix in two directions; r1 and r2 are hyperparameters that control the weight of the KL divergence regularization term, which are used to control the importance of the distribution at both ends in the optimization; γ d is a normalization coefficient that controls the marginal distribution power and adds a small numerical stability term 1×10 -16 , to avoid division by zero in subsequent calculations.

[0091] In the update of the transfer matrix, K is the kernel matrix, which is usually calculated by the cost matrix C and the regularization coefficient ∈, and its form is Represents an element-by-element power operation on each element of the transfer matrix, with a power of r1+r2; the final updated γ is obtained by element-by-element multiplication with K and division by the normalization factor γ d This multiplicative update method has good numerical stability and convergence characteristics in KL divergence optimization.

[0092] The final loss function of the model is the weighted sum of the above losses, optimizing the overall performance of the model:

[0093]

[0094] Where λ1 and λ2 are weights of different loss terms. The final loss combines different loss terms to balance the embedding dimensionality reduction, classification task, and cross-dataset alignment.

[0095] After obtaining the transfer matrix in step S5, find the best matching RNA cell for each ATAC cell:

[0096] RNA label(i) =argmax j T ij

[0097] Assign ATAC cell i to its best matching RNA cell j, and use the label of this RNA cell as the prediction result.

[0098] This example uses a human myocardial infarction dataset downloaded from https: / / cellxgene.cziscience.com / collections / 8191c283-0816-424b-9b61c3e1d6258a77. After preprocessing, 53,206 scRNA-seq cells and 30,739 scATAC-seq cells were obtained. 17,878 genes were shared between the two omics methods. scRNA-seq covered 10 cell types, while scATAC-seq covered 8, with 8 of these shared cell types.

[0099] This example uses the true and predicted labels of scATAC-seq to calculate the overall accuracy, weighted F1 score of cell type classification, Adjusted Rand Index (ARI), and Normalized Mutual Information (NMI). modality , cell type silhouette coefficient s cellTypes , F1 silhouette coefficient, to evaluate the clustering effect of the model.

[0100] This example compares it with five single-cell multi-omics data integration methods (scJoint, Seurat, Portal, SCOT, and UnionCom). To ensure experimental fairness, we used default parameters when running each algorithm model to ensure that the experimental results of the compared algorithms can reach or approach the optimal results reported in the paper. Figure 2 For the UMAP visualization of joint dimensionality reduction, such as Figure 2 and Figure 3 Color by cell type and omics, respectively. Figure 4 and Figure 5 It shows that this application is significantly better than other methods in terms of label transfer accuracy, F1 score, ARI, and NMI. Figure 6 and Figure 7To remove some cell types from scRNA-seq and scATAC-seq of human myocardial infarction data, two imbalanced matching tasks were performed. First, this example removed the "neural receptor cell", "smooth muscle cell" and "pericyte" types from the scRNA-seq data while keeping the scATAC-seq data unchanged, and expressed the integration task as RNA imbalance matching (UBM-RNA). Second, this example removed the same cell types from the scATAC-seq data while keeping the scRNA-seq data unchanged, and expressed the integration task as ATAC imbalance matching (UBM-ATAC). For comparison, this example also defined the integration of the complete human myocardial infarction data as balanced matching (BM). The results show that the model achieved stable performance in all three cases and performed optimally, indicating that the model is suitable for the integration of imbalanced data.

[0101] Compared with the prior art, the present invention has the following beneficial effects:

[0102] This paper integrates single-cell RNA-seq and ATAC-seq data by combining an unbalanced optimal transfer algorithm and a transfer learning strategy. This innovative approach incorporates unbalanced optimal transfer into end-to-end neural network training, achieving alignment at the cell-level. A dual alignment mechanism aligns the embedding representations and cell type labels of scATAC-seq and scRNA-seq in both embedding and label spaces. By relaxing the strict mass conservation constraints of traditional over-the-air (OT) training, this approach effectively addresses cross-modal data integration artifacts caused by imbalanced cell type distributions and technical bias.

[0103] The above describes in detail a method for integrating single-cell RNA-seq and ATAC-seq data provided by this application. The description of the specific embodiments is only intended to help understand the method and core ideas of this application. It should be noted that for those skilled in the art, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A method for integrating single-cell RNA-seq and ATAC-seq data, characterized in that: Including steps: S1. Collect gene expression matrix for scRNA-seq, gene activity matrix for scATAC-seq, and cell type annotation for scRNA-seq; S2. Preprocess scRNA-seq and scATAC-seq data to obtain genes shared by scRNA-seq and scATAC-seq; S3. Build an encoder network to achieve joint representation of different omics data in a shared embedding space; S4. Supervised learning guides the embedding space using cell type labels from scRNA-seq data. S5. Align scRNA-seq and scATAC-seq data using the unbalanced optimal transfer algorithm in both embedding and label spaces. S6. Assign cell type annotations to each cell in the scATAC-seq data based on the matching probability matrix and integrate scRNA-seq data with scATAC-seq data.

2. A method for integrating single-cell RNA-seq and ATAC-seq data according to claim 1, characterized in that: The gene expression matrix for scRNA-seq and the gene activity matrix for scATAC-seq each row represents a cell and each column represents a gene.

3. A method for integrating single-cell RNA-seq and ATAC-seq data according to claim 2, characterized in that: An encoder is trained jointly by scRNA-seq data and scATAC-seq data, and a shared low-dimensional embedding space is learned by weight sharing. The small batch B0 of training data is constructed by sampling cell subsets of equal size from each dataset, that is, Each subset B (s) and B (t) All have B cells; Low-dimensional orthogonal features are captured when projecting each data batch into the embedding space. The encoding loss is measured as the deviation of the embedding representation from the identity matrix, and the encoder is optimized by minimizing the encoding loss.

4. A method for integrating single-cell RNA-seq and ATAC-seq data according to claim 3, characterized in that: The dimension reduction is performed on each RNA dataset, and its encoding loss is expressed as: Among them, B (s) represents the cell subset sampled from the scRNA-seq dataset, B is the number of cells in the current subset, D is the dimension of the embedding space, and θ is the parameter set of the encoder network. represents the embedding vector generated by the b-th scRNA-seq sample under the parameter θ, represents the value of the embedding vector in the jth dimension, Represents the mean of all scRNA-seq samples in this dimension, represents the covariance between the i-th and j-th embedding dimensions in the RNA mode, i,j∈{1,…,D} represents the dimension index, and |·| represents the absolute value, which is used to unify the measurement deviation or correlation size; The encoding loss represented by ATAC embedding is the same as that of RNA. Let ATAC embedding be Its dimensionality reduction loss is expressed as: Among them, B (t) represents the cell subset sampled from the scATAC-seq dataset, B is the number of cells in the current subset, D is the dimension of the embedding space, θ is the parameter set of the encoder network, represents the embedding vector generated by the b-th scATAC-seq sample under parameter θ, represents the value of the embedding vector in the jth dimension, represents the mean value of all scATAC-seq samples on this dimension, represents the covariance between the i-th and j-th embedding dimensions in the ATAC mode, i,j∈{1,…,D} represents the dimension index, |·| represents the absolute value, which is used to unify the measurement deviation or correlation size; In this loss function, the inverse of the first term encourages each coordinate dimension to have a large intra-sample variability, and the second term penalizes the correlation between different dimensions; Finally, the encoder loss is the sum of the RNA encoding loss and the ATAC encoding loss, which uniformly constrains the embedding space features of different modalities:

5. A method for integrating single-cell RNA-seq and ATAC-seq data according to claim 4, characterized in that: In step S4, after softmax transformation, the classifier network outputs a probability vector for cell type prediction; Each with cell type annotation B (s) , the classification loss is used to supervise the training of scRNA-seq dataset, defined as the cross entropy loss, and the formula is expressed as: Among them, 1(.) represents the indicator function, represents the predicted probability that the bth cell belongs to the cth class.

6. A method for integrating single-cell RNA-seq and ATAC-seq data according to claim 5, characterized in that: Alignment between RNA and ATAC embeddings and predicted labels using the imbalanced optimal transfer algorithm, cost matrix It represents the distance between the source distribution a and the target distribution b, and is defined as: M=αM embed +λ t M sce Among them, M embed =||f RNA -f ATAC || 2 Represents the Euclidean distance matrix between RNA and ATAC embeddings, which is used to measure the feature similarity of the two datasets in the embedding space. represents the cross-entropy-based label alignment matrix used to align RNA cell type labels with the probability distribution predicted by the ATAC model, and α and are hyperparameters used to tune M embed and M sce Weights in the cost matrix M, optimal transmission matrix It is used to describe the transmission between the source distribution and the target distribution. The calculation formula is as follows: in, and are the source and target distributions, is the regularized reference distribution, reg m1 and reg m2 Represents the regularization term, used to prevent the model from overfitting, <γ,M> F Used to measure transmission plan loss, reg m1 KL(γ1||a) and reg m2 KL(γ T 1||b) is the marginal constraint term for the transfer matrix γ, which measures the difference between the row marginal distribution and column marginal distribution of γ and a and b respectively. reg·KL(γ||c) is the regularization term, which represents the KL difference between the transfer matrix γ and the reference distribution c. The constraints are: That is, all elements in the transfer matrix are non-negative; scUOTL uses KL divergence to update the transfer matrix γ and calculate the error until convergence. The KL divergence update formula is: Among them, γ represents the current transfer matrix, which is used to describe the correspondence between source domain samples and target domain samples; γ i,1 represents the element in row i and column 1, γ j,0 represents the element in the jth row and the 0th column, which is used to measure the marginal distribution of the current transmission matrix in two directions; r1 and r2 are the hyperparameters that control the weight of the KL divergence regularization term, which are used to control the importance of the distribution at both ends in the optimization; γ d is a normalization coefficient that controls the marginal distribution power and adds a small numerical stability term 1×10 -16 ; In the update of the transfer matrix, K is the kernel matrix, which is usually calculated by the cost matrix C and the regularization coefficient ∈, and its form is K = exp(-C / ∈); Represents an element-by-element power operation on each element of the transfer matrix, with a power of r1+r2; the final updated γ is obtained by element-by-element multiplication with K and division by the normalization factor γ d obtained; The final loss function of the model is the weighted sum of the above losses, optimizing the overall performance of the model: Where λ1 and λ2 are weights of different loss terms. The final loss combines different loss terms to balance the embedding dimensionality reduction, classification task, and cross-dataset alignment.

7. A method for integrating single-cell RNA-seq and ATAC-seq data according to claim 6, characterized in that: After obtaining the transfer matrix in step S5, find the best matching RNA cell for each ATAC cell: RNA label(i) =argmax j T ij Assign ATAC cell i to its best matching RNA cell j, and use the label of this RNA cell as the prediction result.

Citation Information

Cited By

  • Single cell identification method and storage medium

    CN121354689A