A single-cell data denoising method

By extracting and fusion of single-cell data and modifying the counting matrix using comparison learning, the problem of ‘false zero’ noise in single-cell data is solved, and the data noise reduction effect and processing efficiency are improved.

CN119479808BActive Publication Date: 2025-06-24SICHUAN COMPUTER RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411301445.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-06-24
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

The existence of ‘false zero’ noise in the current single-cell data limits the further work of biomedical workers and requires denoising single-cell data to promote the development of corresponding disciplines.

Method used

A single-cell data denoising method is adopted to extract DNA and RNA sequences from scRNA-seq and scATAC-seq data, encode them into a single-hot encoding matrix, and feature extraction and fusion are performed through stacked convolutional neural networks and gated networks. Finally, the counting matrix is ​​corrected through comparative learning to complete single-cell data denoising.

Benefits of technology

It improves the noise reduction effect of single-cell data, can more accurately characterize the single-cell characteristics in multi-omics data, reduces the overfitting phenomenon of deep learning models in single-cell data processing, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479808B_ABST
    Figure CN119479808B_ABST
Patent Text Reader

Abstract

The present invention discloses a single-cell data denoising method, which relates to the field of cell data processing. This method is based on DNA sequences and RNA sequences for modeling, simultaneously predicts chromatin accessibility and gene expression status, and integrates paired single-cell multi-omics data. The weights of the classifier are used as the low-dimensional features of the cells in the two omics, and contrastive learning is used to learn the cell heterogeneity therein, characterize the true profile and biological information of the single-cell data, and effectively denoise the single-cell data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cell data processing, and particularly to a method for denoising single-cell data. Background Art

[0002] As an important branch of gene technology, single-cell omics technology can reveal genetic and expression differences between different cells through high-throughput sequencing and analysis of single cells, deeply understand information on cell diversity, development processes, disease mechanisms, etc., and thus promote the development of precision medicine. So far, numerous large international cooperation projects have accumulated a vast amount of data on human cells and genomes, providing valuable resources for the development of single-cell omics technology.

[0003] Single-cell omics data analysis, as an important topic in modern bioinformatics, plays an important role in modern life sciences. Traditionally, biological research mainly focused on homogeneous samples of a large number of cells, ignoring the heterogeneity between cells. However, the emergence of single-cell omics technology has changed this situation. Parallel detection of cell heterogeneity in the genome and proteome enables biomedical workers to understand diseases and cells more deeply. However, the existence of "false zeros" noise in current single-cell data will limit the further work of biomedical workers. Therefore, denoising single-cell data will effectively promote the development of corresponding disciplines. Summary of the Invention

[0004] Aiming at the above deficiencies in the prior art, a method for denoising single-cell data provided by the present invention realizes the correction of the count matrix and completes the denoising of single-cell data.

[0005] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0006] A method for denoising single-cell data is provided, which includes the following steps:

[0007] S1. Extract DNA sequences from scRNA-seq data and transcribe them into RNA sequences; extract DNA sequences from scATAC-seq data;

[0008] S2. Encode the transcribed RNA sequences and the DNA sequences extracted from scATAC-seq data into one-hot encoding matrices respectively, and correspondingly obtain a DNA sequence encoding matrix and an RNA sequence encoding matrix;

[0009] S3. Concatenate the DNA sequence encoding matrix and the RNA sequence encoding matrix to obtain a concatenated matrix;

[0010] S4. Obtain the features of the DNA sequence encoding matrix, the features of the RNA sequence encoding matrix, and the features of the splicing matrix respectively through the stacked convolutional neural network;

[0011] S5. Concatenate the DNA sequence encoding matrix with its features to obtain the DNA fusion feature; concatenate the RNA sequence encoding matrix with its features to obtain the RNA fusion feature;

[0012] S6. Fuse the DNA fusion feature with the features of the splicing matrix through the trained gating network to obtain the DNA hybrid feature; fuse the RNA fusion feature with the features of the splicing matrix through the trained gating network to obtain the RNA hybrid feature;

[0013] S7. Perform residual connection on the DNA hybrid feature and the DNA sequence encoding matrix to obtain the sequence embedding of the scATAC-seq data; perform residual connection on the RNA hybrid feature and the RNA sequence encoding matrix to obtain the sequence embedding of the scRNA-seq data;

[0014] S8. Use the sequence embedding of the scATAC-seq data as the input of the classifier in the trained gating network to obtain the predicted probability value of chromatin accessibility; use the sequence embedding of the scRNA-seq data as the input of the classifier in the trained gating network to obtain the predicted probability value of gene expression status;

[0015] S9. Correct the count matrices of the scATAC-seq data and the scRNA-seq data according to the predicted probability value of chromatin accessibility and the predicted probability value of gene expression status to complete the denoising of single-cell data.

[0016] The beneficial effects of the present invention are as follows:

[0017] 1. This method is based on DNA sequences and RNA sequences for modeling, simultaneously predicts chromatin accessibility and gene expression status, and integrates paired single-cell multi-omics data. The weights of the classifier are used as the low-dimensional features of cells for the two omics, and contrast learning is used to learn the cell heterogeneity therein, improving the denoising effect of single-cell data.

[0018] 2. This method adopts a lightweight deep learning model, avoiding the embarrassing situation of difficult training of current single-cell data fusion models based on the encoder-decoder structure, and can also improve the data processing efficiency.

[0019] 3. This method takes into account the influence of cell heterogeneity on single-cell multi-omics data fusion, and uses contrast learning to be able to more accurately characterize single-cell features in multi-omics data, laying a foundation for improving the denoising effect of single-cell data. Description of the Drawings

[0020] Figure 1 It is a flow schematic diagram of this method;

[0021] Figure 2 It is a comparison of the performance of this method and existing methods in chromatin accessibility prediction in the embodiment;

[0022] Figure 3 It is a comparison of the performance of this method and existing methods in gene expression state prediction in the embodiment;

[0023] Figure 4 It is the inference result of this method in the 10X multiome PBMC dataset in the embodiment;

[0024] Figure 5 It is the inference result of this method in the SNARE-seq dataset in the embodiment;

[0025] Figure 6 It is the inference result of this method in the SHARE-seq dataset in the embodiment;

[0026] Figure 7 It is the performance of this method in multi-omics data integration in the embodiment;

[0027] Figure 8 It is the performance of GLUE in multi-omics data integration in the embodiment;

[0028] Figure 9 It is the performance of scBridge in multi-omics data integration in the embodiment;

[0029] Figure 10 It is the performance of scJoint in multi-omics data integration in the embodiment;

[0030] Figure 11 It is the cosine similarity between the same type of cells in two omics in the 10X multiome PBMC dataset of this method and the comparative method in the embodiment;

[0031] Figure 12 It is the cosine similarity between the same type of cells in two omics in the SNARE-seq dataset of this method and the comparative method in the embodiment;

[0032] Figure 13 It is the cosine similarity between the same type of cells in two omics in the SHARE-seq dataset of this method and the comparative method in the embodiment;

[0033] Figure 14 It is the performance of this method in multi-omics data integration on the SNARE-seq dataset in the embodiment;

[0034] Figure 15Performance of GLUE in multi-omics data integration on the SNARE-seq dataset in the examples;

[0035] Figure 16 Performance of scBridge in multi-omics data integration on the SNARE-seq dataset in the examples;

[0036] Figure 17 Performance of scJoint in multi-omics data integration on the SNARE-seq dataset in the examples;

[0037] Figure 18 Performance of the present method and the comparative method in scATAC-seq data denoising and scRNA-seq data denoising on the 10X multiome PBMC dataset in the examples.

[0038] Figure 19 Effect of the denoising of the present method on the correlation between the same type of cells in scATAC-seq data and scRNA-seq data in the examples;

[0039] Figure 20 Performance of the present method and the comparative method in terms of inference efficiency in the examples;

[0040] Figure 21 Data profile of scATAC-seq data before denoising in the examples;

[0041] Figure 22 Data profile of scATAC-seq data after being denoised by the present method in the examples;

[0042] Figure 23 Data profile of scRNA-seq data before denoising in the examples;

[0043] Figure 24 Data profile of scRNA-seq data after being denoised by the present method in the examples. Detailed implementation manners

[0044] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0045] As Figure 1 shown, the single-cell data denoising method includes the following steps:

[0046] S1. Extract DNA sequences from scRNA-seq data and transcribe them into RNA sequences; extract DNA sequences from scATAC-seq data;

[0047] S2. Encode the transcribed RNA sequences and the DNA sequences extracted from scATAC-seq data into one-hot encoding matrices respectively, obtaining a DNA sequence encoding matrix and an RNA sequence encoding matrix;

[0048] S3. Concatenate the DNA sequence encoding matrix and the RNA sequence encoding matrix to obtain a concatenated matrix;

[0049] S4. Obtain the features of the DNA sequence encoding matrix, the features of the RNA sequence encoding matrix, and the features of the concatenated matrix respectively through a stacked convolutional neural network;

[0050] S5. Concatenate the DNA sequence encoding matrix with its features to obtain a DNA fusion feature; concatenate the RNA sequence encoding matrix with its features to obtain an RNA fusion feature;

[0051] S6. Fuse the DNA fusion feature with the features of the concatenated matrix through a trained gating network to obtain a DNA hybrid feature; fuse the RNA fusion feature with the features of the concatenated matrix through a trained gating network to obtain an RNA hybrid feature;

[0052] S7. Perform residual connection between the DNA hybrid feature and the DNA sequence encoding matrix to obtain the sequence embedding of scATAC-seq data; perform residual connection between the RNA hybrid feature and the RNA sequence encoding matrix to obtain the sequence embedding of scRNA-seq data;

[0053] S8. Use the sequence embedding of scATAC-seq data as the input of the classifier in the trained gating network to obtain the predicted probability value of chromatin accessibility; use the sequence embedding of scRNA-seq data as the input of the classifier in the trained gating network to obtain the predicted probability value of gene expression status;

[0054] S9. Correct the count matrices of scATAC-seq data and scRNA-seq data according to the predicted probability value of chromatin accessibility and the predicted probability value of gene expression status to complete single-cell data denoising.

[0055] The specific method of encoding into a one-hot encoding matrix is as follows: both the transcribed RNA sequence and the DNA sequence extracted from scATAC-seq data are encoded into a sparse one-hot encoding matrix with a dimension of 5; among them, the 5 dimensions correspond to adenine A, guanine G, cytosine C, thymine T, and uracil U respectively; the lengths of the transcribed RNA sequence and the DNA sequence extracted from scATAC-seq data are 1344-bp.

[0056] The stacked convolutional neural network includes 7 layers of convolutional structures, and the calculation expression of the first layer of convolutional structure is:

[0057]

[0058] where is the output of the first layer of convolutional structure; , represents the DNA sequence encoding matrix, represents the RNA sequence encoding matrix, represents the splicing matrix; and are the weights and biases of the first layer of convolutional structure respectively; is the GLUE activation function; represents a 1×1 convolutional operation; represents batch normalization;

[0059] The calculation expressions of the second to seventh layer of convolutional structures are:

[0060]

[0061] where is the output of the layer of convolutional structure. When takes the value of 7, is the feature corresponding to the input. ; is the output of the layer of convolutional structure; and are the weights and biases of the layer of convolutional structure respectively; represents a max pooling operation.

[0062] The expressions for obtaining the DNA mixed feature and the RNA mixed feature are:

[0063]

[0064] where , is the DNA mixed feature, is the RNA mixed feature; , is a DNA fusion feature, is an RNA fusion feature; and are both weights; and are both biases; is a feature of the splicing matrix; is a softmax activation function; is a GLUE activation function.

[0065] The expression for obtaining the sequence embeddings of scATAC-seq data and scRNA-seq data is:

[0066]

[0067] where , is the sequence embedding of scATAC-seq data, is the sequence embedding of scRNA-seq data; and are both weights; and are both biases; , is a DNA sequence encoding matrix, represents an RNA sequence encoding matrix.

[0068] The expression for the classifier in the trained gated network is:

[0069]

[0070] where , is the predicted probability value of chromatin accessibility; is the predicted probability value of gene expression status; , is the weight of the classifier for processing the sequence embedding , is the weight of the classifier for processing the sequence embedding ; , is the bias of the classifier for processing the sequence embedding , is the bias of the classifier for processing the sequence embedding ; is a sigmoid activation function.

[0071] The specific method for training the gated network includes the following sub-steps:

[0072] A1. Randomly initialize the weights and biases in the gating network. Take as the cell expression matrix of scATAC-seq data, such that each vector in represents the feature vector of a cell; take as the cell expression matrix of scRNA-seq data, such that each vector in represents the feature vector of a cell;

[0073] A2. Average all the feature vectors corresponding to the cells of the th class in the current along the vector dimension, and take the averaged result as the DNA cell prototype of the th class of cells; average all the feature vectors corresponding to the cells of the th class in the current along the vector dimension, and take the averaged result as the RNA cell prototype of the th class of cells; thus obtaining two cell prototypes of the cells used in the training process of the gating network for scATAC-seq data and scRNA-seq data;

[0074] A3. Calculate the distance vector between the prototypes of the non- class cells used in the training process of the gating network for scATAC-seq data and scRNA-seq data and ; the corresponding expression is:

[0075]

[0076] where is the total number of cell types used in the training process of the gating network for scATAC-seq data and scRNA-seq data; is the temperature coefficient; represents the th non- class cell; represents the exponential function with the natural constant e as the base;

[0077] A4. Calculate the distance vector between the current corresponding cell prototypes of the th class of cells; the corresponding expression is:

[0078] ;

[0079] A5. Construct a loss function to obtain the total loss value of the gating network, and update the weights and biases of the gating network using the total loss value until the loss function converges. The expression of the loss function is as follows:

[0080]

[0081]

[0082]

[0083]

[0084] where is the total loss value of the gating network; is the contrastive learning loss value; represents the -th element in the feature vector of a single cell; represents the total number of elements in the feature vector of a single cell; represents the natural logarithm function with base \(e\); is the binary cross-entropy loss value corresponding to chromatin accessibility of cells; represents the true probability value of chromatin accessibility of the -th class of cells; represents the predicted probability value of chromatin accessibility corresponding to the -th class of cells; represents the true probability value of gene expression status of the -th class of cells; represents the predicted probability value of gene expression status corresponding to the -th class of cells; represents the common logarithm function with base 10; is the binary cross-entropy loss value corresponding to gene expression status of cells.

[0085] The specific method for correcting the count matrices of scATAC-seq data and scRNA-seq data is as follows: Use the predicted probability values of chromatin accessibility and gene expression status as the data after noise reduction to complete the correction of "false zeros" in the count matrices of scATAC-seq data and scRNA-seq data, that is, complete the noise reduction of single-cell data.

[0086] In the specific implementation process, scRNA-seq data and scATAC-seq data contain a large amount of single-cell data. Therefore, during the training process, only part of the single-cell data is needed to train the gating network, and then the remaining single-cell data is denoised using this method.

[0087] In the specific implementation process, The initial distance between the two matrices in may be very large, and training makes the distance between the two matrices in

[0088] become smaller and smaller. From the cellular level, the gap between the same type of cells in the two omics will also be very small, and this process is the fusion process. In this embodiment, the 10X multiome PBMC dataset, SNARE-seq dataset, and SHARE-seq dataset were collected from the datasets processed by the Gao Ge team at Peking University. The 10X multiome PBMC dataset contains 19 single-cell populations, 107,162 ATAC peak data, and 29,083 gene expression data; the SNARE-seq dataset provides 22 single-cell populations, 241,757 ATAC peak data, and 28,918 gene expression data; while the SHARE-seq dataset contains 19 single-cell populations, 340,314 ATAC peak data, and 21,478 gene expression data.

[0089] In this embodiment, the area under the receiver operating characteristic curve (Area Under the Receiver Operating Characteristic Curve, auROC) is used to evaluate the performance of this method (CONTRACT) in the tasks of chromatin accessibility prediction and gene state prediction.

[0090] In this embodiment, the learning rate is set to 0.001, the number of training epochs is set to 40, the data batch size is set to 128, and the adaptive moment estimation optimizer (Adam) is used to train CONTRCT. The training set, validation set, and test set are divided according to the ratio of 8:1:1.

[0091] For the DNA sequences retained in the above three standard datasets, this embodiment compares the auROC values of CONTRACT with scBasset, Basenji, and DeepSEA. For fair comparison, DNA and RNA sequences with a length of 1344-bp are used as the input for all comparison methods.

[0092] As Figure 2 shown, in terms of chromatin accessibility prediction, CONTRACT performs better than the existing prediction methods in all three datasets. Especially in the 10X multiome PBMC dataset, the advantage of CONTRACT is more obvious, and its auROC is much higher than the other comparison methods.

[0093] As Figure 3As shown, in terms of gene expression status prediction, CONTRACT still shows better prediction indicators than other methods, except for being second only to Basenji in the SHARE dataset.

[0094] In this example, almost all benchmark methods perform much better on the SHARE-seq dataset than the other two datasets, probably because there is a lot of noise in the single-cell dataset, making it difficult for the model to fit the true distribution of the data, and the number of genes in the SHARE-seq dataset is much higher than the other two datasets, giving the model a better chance to learn the true data labels. In the case of limited data, the performance of CONTRACT shows that the use of multi-task learning and the mutual learning of chromatin accessibility prediction and gene state expression can help reduce the overfitting phenomenon of the model in scRNA-seq data processing and improve the performance of the model in the task of predicting gene expression state.

[0095] The high sparsity and "false zero" noise of the count matrices of scATAC-seq and scRNA-seq make it difficult for current methods to identify open chromatin regions to make further progress. In order to evaluate the robustness of CONTRACT in datasets with a lot of noise, this example randomly samples the non-zero counts in the count matrices of the three datasets to 80%, 60%, 40%, 20% and 10% of the original and retrains CONTRACT, and then completes the inference in the test set divided from the original dataset. Figure 4 , Figure 5 and Figure 6 The reasoning results of CONTRACT are presented, and the results show that even if the noise in the dataset is increased by sampling, the performance of CONTRACT is still slightly affected, which indicates that CONTRACT is very robust to noisy datasets.

[0096] like Figure 7 , Figure 8 , Figure 9 and Figure 10As shown, this implementation qualitatively analyzed the performance of several comparison methods in multi-omics data integration. Both CONTRACT and GLUE well integrated scATAC-seq omics data and scRNA-seq omics data, eliminating the differences between the two omics data. While scBridge and scJoint missed part of the integration of scRNA-seq and scATAC-seq omics data. The reason may be that, on the one hand, the operation of the calculation methods of scBridge and scJoint depends on the alignment of the number of data in the non-cell dimension. The experimental data adopted in this example did not follow this characteristic. The number of characteristic peaks of scATAC-seq was much higher than the number of genes in scRNA-seq data. Therefore, some characteristics of scATAC-seq data may be missed during the sampling process. On the other hand, as Figure 11 、 Figure 12 、 Figure 13 、 Figure 14 、 Figure 15 、 Figure 16 and Figure 17 shown, scBridge and scJoint respectively adopted reliability modeling and the "voting" method to select cells for integration. However, high-score cells are not necessarily correct, which may lead to the model integrating incorrect cells. This phenomenon is more obvious in the integration of the SNARE-seq dataset. After integration, CONTRACT obtained the highest inter-omics cosine similarity, and GLUE also achieved relatively high performance among the omics. While scBridge and scJoint had some cell data that were partially integrated and not integrated, and there were significant differences in the integration performance between cells.

[0097] In this example, the predicted probability value of CONTRACT was used as the denoised data. This example evaluated the denoising performance of CONTRACT on the 10Xmultiome PBMC dataset. Specifically, the denoised count matrix should make the distances between the same cells, the same characteristic peaks, and the same genes closer. Therefore, in this example, based on the denoised cell-characteristic peak count matrix and cell-gene count matrix, the cell type ASW was calculated. This is an index used to measure cell similarity. In addition to CONTRACT, this example used the count matrices generated by SCALE, MAGIC, and scBasset and the original count matrix to calculate the cell type ASW respectively. As Figure 18 shown, in the tasks of scATAC-seq data denoising and scRNA-seq data denoising, CONTRACT obtained the best index, followed by SCALE; MAGIC also achieved relatively good performance in the denoising of scATAC-seq data.

[0098] In this example, we further observed whether there is a higher correlation in the distances between the same cells after noise reduction by CONTRACT, and the cosine similarity was used to measure such a correlation. Specifically, the cosine similarities between the same cell vectors were calculated when only the scATAC-seq data was denoised, only the scRNA-seq data was denoised, neither was denoised, and both were denoised, as Figure 19 shown. It was found that after denoising both omics data, the correlation between the same cell vectors was higher. If only one type of omics data was denoised, the cosine similarity between the same cell vectors was the second highest, but higher than that of the data without denoising, indicating the effectiveness of CONTRACT in denoising both scATAC-seq data and scRNA-seq data.

[0099] Due to the continuous growth of single-cell data volume, the efficiency of the model has become an important reference in the process of single-cell data analysis. Since the denoising process runs through the inference processes of CONTRACT and the above comparison methods, as Figure 20 shown, the running times of CONTRACT and the above several methods for denoising scATAC-seq data were compared and recorded in detail in Table 1. Among them, the inference time of scBasset was the fastest, at 190.01 seconds, and the running times of CONTRACT, SCALE, and MAGIC were 246.05, 883.00, and 344.57 seconds respectively, indicating that CONTRACT can be used in the practical application of single-cell data denoising tasks like other comparison methods. However, all comparison methods can only denoise one type of omics data, while using multi-task learning, CONTRACT can denoise both scATAC-seq data and scRNA-seq data simultaneously. Deep learning methods such as CONTRACT, SCALE, and scBasset are expected to accelerate the running efficiency with the improvement of GPU hardware conditions, but the disadvantage is that the number of cells they can analyze is restricted by the GPU memory.

[0100] Table 1

[0101]

[0102] As Figure 21 , Figure 22 , Figure 23 and Figure 24 shown, a visual comparison of the denoising and the original count matrices of scATAC-seq data and scRNA-seq data was carried out. After denoising, the data can more accurately display the biological significance of single-cell data, making the contours of the same cells closer. This indicates that CONTRACT can show a more realistic distribution of single-cell data, and this performance is particularly obvious in the denoising of scRNA-seq data.

[0103] In summary, in a limited single-cell dataset, most existing single-cell chromatin accessibility prediction or gene state prediction models have serious overfitting problems. By using multi-task learning and contrastive learning methods, the present invention can mitigate such overfitting phenomena when the single-cell dataset is limited, improve the prediction performance of the model, effectively utilize cell heterogeneity, depict the true profile and biological information of single-cell data, and effectively denoise the single-cell dataset.

Claims

1. A single cell data denoising method, characterized in that: The following steps are involved: S1. Extract DNA sequences from scRNA-seq data and transcribe them into RNA sequences; Extract DNA sequences from scATAC-seq data; S2. Encode the transcribed RNA sequence and the DNA sequence extracted from the scATAC-seq data into one-hot encoding matrices, respectively, to obtain a DNA sequence encoding matrix and an RNA sequence encoding matrix; S3, splicing the DNA sequence coding matrix and the RNA sequence coding matrix to obtain a splicing matrix; S4, respectively obtaining the features of the DNA sequence encoding matrix, the features of the RNA sequence encoding matrix, and the features of the splicing matrix through stacked convolutional neural networks; S5, splicing the DNA sequence encoding matrix with its features to obtain a DNA fusion feature; splicing the RNA sequence encoding matrix with its features to obtain an RNA fusion feature; S6. The DNA fusion feature is fused with the feature of the splicing matrix through the trained gating network to obtain a DNA mixed feature; the RNA fusion feature is fused with the feature of the splicing matrix through the trained gating network to obtain an RNA mixed feature; S7, residually connect the DNA mixed features with the DNA sequence encoding matrix to obtain the sequence embedding of the scATAC-seq data; residually connect the RNA mixed features with the RNA sequence encoding matrix to obtain the sequence embedding of the scRNA-seq data; S8. Use the sequence embedding of scATAC-seq data as the input of the classifier in the trained gating network to obtain the predicted probability value of chromatin accessibility; The sequence embedding of scRNA-seq data is used as the input of the classifier in the trained gating network to obtain the predicted probability value of gene expression status; S9. According to the predicted probability values ​​of chromatin accessibility and gene expression status, the counting matrix of scATAC-seq data and scRNA-seq data is modified to complete single-cell data noise reduction.

2. The single cell data denoising method according to claim 1, characterized in that: The specific method of encoding into a one-hot encoding matrix is: The transcribed RNA sequences and the DNA sequences extracted from the scATAC-seq data were encoded as sparse one-hot encoding matrices with a dimension of 5; the 5 dimensions corresponded to adenine A, guanine G, cytosine C, thymine T, and uracil U, respectively; the transcribed RNA sequences and the DNA sequences extracted from the scATAC-seq data were 1344-bp in length.

3. The single cell data denoising method according to claim 1, characterized in that: The stacked convolutional neural network consists of 7 layers of convolutional structures. The calculation expression of the first layer of convolutional structure is: in is the output of the first layer of convolutional structure; , represents the DNA sequence encoding matrix, represents the RNA sequence encoding matrix, represents the concatenation matrix; and They are the weights and biases of the first layer of convolutional structure respectively; is the GLUE activation function; Represents a 1×1 convolution operation; represents batch normalization; The calculation expression of the second to seventh convolution structures is: in For the The output of the layer convolution structure, when When the value is 7, That is, the feature corresponding to the input, ; For the The output of the layer convolutional structure; and Respectively The weights and biases of the layer convolutional structure; Represents a max pooling operation.

4. The single cell data denoising method according to claim 1, characterized in that: The expressions for obtaining DNA mixed features and RNA mixed features are: in , DNA mixing characteristics, It is RNA mixed feature; , DNA fusion characteristics, It is the RNA fusion feature; and All are weights; and All are biased; is the feature of the concatenated matrix; is the softmax activation function; is the GLUE activation function.

5. The single cell data denoising method according to claim 4, characterized in that: The expressions for obtaining the sequence embeddings for scATAC-seq data and scRNA-seq data are: in , For sequence embedding of scATAC-seq data, Sequence embedding for scRNA-seq data; and All are weights; and All are biased; , is the DNA sequence encoding matrix, Represents the RNA sequence encoding matrix.

6. The single cell data denoising method according to claim 5, characterized in that: The expression of the classifier in the trained gating network is: in , is the predicted probability value of chromatin accessibility; is the predicted probability value of gene expression status; , Embedding for processing sequences The weight of the classifier, Embedding for processing sequences The weight of the classifier; , Embedding for processing sequences The bias of the classifier, Embedding for processing sequences The bias of the classifier; is the sigmoid activation function.

7. The single cell data denoising method according to claim 6, characterized in that: The specific method for training the gated network includes the following sub-steps: A1. Randomly initialize the weights and biases in the gating network. As a cell expression matrix for scATAC-seq data, Each vector in represents the feature vector of a cell; As the cell expression matrix of scRNA-seq data, Each vector in represents the feature vector of a cell; A2. Set the current Middle All feature vectors corresponding to the cell-like cells are averaged in the vector dimension, and the average result is used as the first Cell-like DNA cell prototype ; Set the current Middle All feature vectors corresponding to the cell-like cells are averaged in the vector dimension, and the average result is used as the first Cell-like RNA cell prototype ; Then we obtain two cell prototypes of cells used in the gating network training process of scATAC-seq data and scRNA-seq data; A3. Calculate the non-significant amount of gating network used in the training process of scATAC-seq data and scRNA-seq data. Cell-like prototypes and The distance vector ; The corresponding expression is: in is the total number of cell types used in the gating network training process for scATAC-seq data and scRNA-seq data; is the temperature coefficient; Indicates Non Cell-like; It represents the exponential function with the natural constant e as the base; A4. Calculate the The distance vector between the cell prototypes currently corresponding to the class cell ; The corresponding expression is: ; A5. Construct a loss function to obtain the total loss value of the gating network, and update the weights and biases of the gating network by the total loss value until the loss function converges; the expression of the loss function is: in is the total loss value of the gating network; is the comparison learning loss value; The feature vector representing a single cell elements; The total number of elements in the feature vector representing a single cell; It represents the logarithmic function with the natural constant e as the base; is the binary cross entropy loss value corresponding to cell chromatin accessibility; Indicates True probability value of cell-like chromatin accessibility; Indicates The predicted probability value of chromatin accessibility corresponding to the cell class; Indicates The true probability value of the cell-like gene expression state; Indicates The predicted probability value of the cell-like gene expression state; represents the logarithmic function with base 10; is the binary cross entropy loss value corresponding to the cell gene expression state.

8. The single cell data denoising method according to claim 1, characterized in that: The specific method for correcting the count matrix of scATAC-seq data and scRNA-seq data is as follows: The predicted probability values ​​of chromatin accessibility and gene expression status were used as denoised data to correct the "false zero" counts in the count matrix of scATAC-seq data and the count matrix of scRNA-seq data, thus completing the denoising of single-cell data.

Citation Information

Patent Citations

  • Single cell data integration method based on domain confrontation and variation inference

    CN114819056A

  • Method for analyzing active pathway in single cell multi-omics based on graph neural network

    CN115240772A