Deconvolution method for RNA-seq sequencing data of tumor Bulk samples based on deep learning model

By constructing a deep learning model, integrating the scRNA-seq data set and generating pseudoBulk RNA-seq data, optimizing model parameters, solving the shortcomings of existing tools in deconvolution of tumor samples, and achieving efficient and accurate deconvolution of tumor Bulk RNA-seq data, revealing tumor heterogeneity and cell interactions, providing tools for precision medicine.

CN119864089BActive Publication Date: 2025-07-29HANGZHOU INSTITUTE OF MEDICAL SCIENCES CHINESE ACADEMY OF SCIENCES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510343511.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-29
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing Bulk RNA-seq deconvolution tools are insufficiently accurate for tumor samples, cannot effectively handle heterogeneity and dynamic changes in cell types, and rely on known reference genomic data, neglecting technical errors and bionoise.

Method used

A deconvolution method for tumor Bulk sample RNA-seq sequencing data based on deep learning models is constructed, multiple scRNA-seq data sets are integrated for cell annotation, and pseudoBulk RNA-seq data sets are generated. Model parameters are optimized through error backpropagation and gradient descent, and deep feature learning and decoding are combined with residual connection module to achieve efficient deconvolution.

Benefits of technology

It improves the generalization ability and prediction accuracy of the model, enhances the adaptability to diverse tumor samples, and can be widely used in research on multiple tumor types, reveals tumor heterogeneity and cell-interactions, and provides support for precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864089B_ABST
    Figure CN119864089B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for deconvoluting RNA-seq sequencing data of tumor Bulk samples based on a deep learning model, which integrates multiple scRNA-seq data sets containing tumor cells, performs cell annotation, constructs a scRNA-seq reference data set and generates a pseudo-Bulk RNA-seq data set; subsequently, the two data sets are normalized and gene screened, and then input into two encoders to extract deep latent features, and the extracted features are deconvoluted into cell proportions and cell-specific gene expression matrices; the model parameters are iterated through the error backpropagation algorithm and the gradient descent method, and finally the model weights with the best validation effect are saved; through the combination of deep learning and the construction and annotation of the scRNA-seq reference data set, the present invention realizes the efficient deconvolution of data, and uses random sampling to generate pseudo-Bulk data for training, improving the generalization ability and prediction accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of bioinformatics and computer applications, and particularly relates to a method for deconvoluting tumor Bulk sample RNA-seq sequencing data based on a deep learning model. Background Art

[0002] In the fields of bioinformatics and cancer research, Bulk RNA-seq and single-cell RNA-seq (scRNA-seq) are two important transcriptomic research techniques, each with unique advantages and application scenarios. Bulk RNA-seq provides a transcriptomic snapshot of a tissue or cell population by measuring the average gene expression levels of the cell population. This method has the advantages of low cost and high throughput, and can reveal the average gene expression profile at the macroscopic level, and is suitable for studying complex biological processes. In contrast, scRNA-seq analyzes gene expression characteristics at the single-cell level, can identify rare cell types, reveal the heterogeneity in the cell population, and discover cell states or transitional states that may be missed in Bulk analysis. However, scRNA-seq has a high cost, large experimental and computational workload, and requires specialized computational analysis methods to handle its data complexity.

[0003] Existing Bulk RNA-seq deconvolution tools have several significant drawbacks, including dependence on known reference genome datasets, neglect of the heterogeneity and dynamic changes of cell types in samples, and deficiencies in handling technical errors and biological noise. In addition, there is currently no tool that can accurately deconvolute tumor Bulk samples. For example, in the Chinese patent with the application number US20230049525A1 and the title "Method for Identifying Cell Type-Specific Gene Expression Levels by Deconvoluting Bulk Gene Expression", a method for identifying gene expression levels in specific cell types based on the measurement of bulk gene expression levels in tissue samples containing multiple cell types is disclosed. In the Chinese patent with the application number WO2021108556A1 and the title "Method for Identifying Cell Type-Specific Gene Expression Levels by Deconvoluting Bulk Gene Expression", a method for identifying gene expression levels in specific cell types based on the measurement of bulk gene expression levels in tissue samples containing multiple cell types is disclosed. The identification methods involved in the above two patents are representative and neither can provide an accurate deconvolution tool specifically for tumor Bulk samples. Summary of the Invention

[0004] To address the deficiencies of current methods for deconvoluting tumor Bulk sample RNA-seq data, the present invention discloses a method for deconvoluting tumor Bulk sample RNA-seq sequencing data based on a deep learning model, including the following steps:

[0005] S1. Integrate multiple scRNA-seq datasets containing tumor cells, perform cell annotation, and construct a large-scale scRNA-seq reference dataset;

[0006] S2. Simulate the cell composition of real Bulk RNA-seq to provide training and validation data for the model: randomly sample from the annotated scRNA-seq reference dataset, combine according to the specified cell ratios, and generate a pseudo-Bulk RNA-seq dataset;

[0007] S3. Data preprocessing: normalize and perform gene screening on the scRNA-seq reference dataset and the pseudo-Bulk RNA-seq dataset to ensure the consistency of the input data in terms of feature expression and gene sets;

[0008] S4. Input the preprocessed scRNA-seq reference dataset and pseudo-Bulk RNA-seq dataset into the reference set encoder and the Bulk encoder respectively to extract the deep latent features representing the two datasets, and then use the decoder to deconvolve the extracted features into cell ratios and cell-specific gene expression matrices;

[0009] S5. Continuously update and optimize the model parameters through the error backpropagation algorithm and the gradient descent method: in each training iteration, evaluate the error between the model prediction and the real pseudo-Bulk data through the mean absolute error and mean squared error loss functions, gradually reduce the prediction error, and finally save the model weights with the best validation effect.

[0010] Preferably, in S1, perform principal component analysis and UMAP visualization on the scRNA-seq dataset, and use the scVI model to remove batch effects from the entire scRNA-seq dataset;

[0011] Adopt the Leiden algorithm to perform clustering analysis on the scRNA-seq reference dataset, divide it into multiple clusters, display the highly expressed marker genes in each cluster through a dot plot, and thus determine the specific cell types of each cluster. Finally, annotate the scRNA-seq reference dataset as 10 different cell types, namely T cells, NK cells, B cells, plasma cells, myeloid cells, mast cells, fibroblasts, endothelial cells, epithelial cells, and unrecognized cells. Among them, the unrecognized cells will be discarded in subsequent analyses.

[0012] Preferably, further subgroup annotation is performed on some of the cell types: all T cells and NK cells are extracted from the scRNA-seq reference dataset, and the previous steps are repeated to further annotate T cells and NK cells as resting T cells, helper T cells, cytotoxic T cells, exhausted T cells, regulatory T cells, and NK cells; similarly, subgroup annotation is performed on B cells, myeloid cells, fibroblasts, and endothelial cells.

[0013] Preferably, the inferCNV algorithm is used to analyze the scRNA-seq reference dataset to identify malignant tumor cells; for epithelial cells, cells with higher CNV scores are defined as malignant tumor cells, thereby obtaining the finally annotated scRNA-seq reference dataset for subsequent deep learning network model training.

[0014] Preferably, in S2, the annotated scRNA-seq reference dataset is randomly sampled to construct a pseudo BulkRNA-seq dataset, and the specific steps are as follows:

[0015] S2.1. Randomly draw an integer between 1000 and 5000 as the total number of cells M in the Bulk sample.

[0016] S2.2. Randomly draw N numbers between 0 and 1, where N is the number of cell categories finally annotated in the scRNA-seq reference dataset, and normalize these numbers so that their sum is 1 as the proportion P of each cell category in the Bulk sample.

[0017] S2.3. According to the total number of cells M and the cell category proportion P, calculate the number of cells T in each cell category; then, randomly select Ti (i = 1, 2, 3 …, N) cells for each cell category from the scRNA-seq reference dataset.

[0018] S2.4. Calculate the cell category-specific gene expression for the randomly selected cells and obtain the pseudo Bulk RNA-seq data through addition calculation.

[0019] Through the above steps, pseudo Bulk RNA-seq datasets with sample sizes of a and b are constructed for the training set and the validation set respectively, where a > b.

[0020] Preferably, in S3, the scRNA-seq reference dataset is standardized: the total reads of each cell are standardized to c and logarithmically transformed; then, the Seurat_v3 algorithm in scanpy is used to screen for highly variable genes, and the top d highly variable genes are selected. Then, cell type classification addition operations are performed on the screened dataset to generate a feature matrix of cell categories and highly variable genes.

[0021] Next, according to these d highly variable genes, gene screening is performed on the pseudo-Bulk RNA-seq data and the pseudo-cell-specific gene expression matrix to ensure that the genes in the pseudo-Bulk samples are consistent with the highly variable genes in the scRNA-seq reference dataset; finally, min-max normalization is performed on the pseudo-Bulk RNA-seq data, and log1p transformation is performed on the pseudo-cell-specific gene expression matrix; where a, b, c, and d are generally integer multiples of 10,000.

[0022] Preferably, in S4, the preprocessed scRNA-seq feature matrix and the pseudo-Bulk RNA-seq data are respectively input into two fully connected layers with residual connection modules: the reference set encoder and the Bulk encoder, to perform deep low-dimensional feature learning to obtain the latent features Zref and Zbulk; subsequently, matrix multiplication and dot product operations are performed on these two latent features to calculate the cell proportion feature Zprop and the cell-specific gene expression feature Zexpr; finally, Zprop and Zexpr are respectively input into the cell proportion decoder and the cell-specific gene expression decoder for decoding, so as to obtain the predicted cell proportion prop and the cell-specific gene expression matrix Expr.

[0023] Preferably, in S5, the entire deep network model is optimized by minimizing the error loss, and the optimization objective is:

[0024]

[0025] Where:

[0026] is the known true cell proportion in the pseudo-Bulk data;

[0027] is the known true cell-specific gene expression matrix in the pseudo-Bulk data;

[0028] is the cell proportion predicted by the Prop Decoder;

[0029] is the cell-specific gene expression matrix predicted by the Expr Decoder;

[0030] is a hyperparameter used to adjust the weight between the two parts of the loss;

[0031] represents the Mean Absolute Error (MAE), which is used to evaluate the difference between the predicted cell proportion and the true proportion;

[0032] Denotes the Mean Squared Error, which is used to evaluate the difference between the predicted cell-specific gene expression matrix and the true value.

[0033] Preferably, when the training reaches the preset stop condition or the error converges to a lower level, the model weights with the best validation effect are saved for subsequent deconvolution analysis of real Bulk RNA-seq data.

[0034] Preferably, the application of the model weights: Input the preprocessed Bulk RNA-seq data into the pre-trained deep learning model, and the model will predict the proportions of each cell subset in the Bulk sample and the corresponding cell type-specific gene expression matrix.

[0035] The deconvolution method for tumor Bulk sample RNA-seq sequencing data based on the deep learning model of the present invention has at least the following advantages:

[0036] 1. The present invention realizes the efficient deconvolution of tumor Bulk RNA-seq data through deep learning combined with the construction and annotation of a large-scale scRNA-seq reference dataset.

[0037] 2. Random sampling is used to generate pseudo-Bulk data for training, combined with the residual connection module and the optimized loss function, which improves the generalization ability and prediction accuracy of the model.

[0038] 3. The adaptability to diverse tumor samples is enhanced, and it can be widely applied to the research of various tumor types.

[0039] 4. It provides a powerful and convenient tool for the in-depth analysis of the tumor microenvironment, which helps to reveal the heterogeneity of tumors, cell-cell interactions, and the key mechanisms of tumor occurrence and development.

[0040] 5. It can better analyze the response of the tumor microenvironment to treatment regimens such as immunotherapy and targeted therapy, and provides support for the development of precision medicine and individualized treatment strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The present invention is further described with reference to the accompanying drawings, but the content in the drawings does not constitute any limitation to the present invention.

[0042] Figure 1 It is the UMAP visualization result of the scRNA-seq reference dataset before and after batch correction after integration.

[0043] Figure 2It is the cell clustering map of the scRNA-seq reference dataset after Leiden clustering analysis and its corresponding cell type annotation map.

[0044] Figure 3 It is a dot plot of the expression of marker genes in different cell clusters of the scRNA-seq reference dataset.

[0045] Figure 4 It is the final cell type annotation result of the scRNA-seq reference dataset.

[0046] Figure 5 It is a process for deconvoluting the RNA-seq sequencing data of Bulk samples based on a deep learning model. Detailed implementation mode

[0047] The present invention will be further described in detail below through embodiments, so that those skilled in the art can implement it according to the text of the specification.

[0048] It should be understood that terms such as "having", "comprising" and "including" used herein do not exclude the presence or addition of one or more other elements or their combinations.

[0049] The present invention will be further described in conjunction with the accompanying drawings, including the following steps:

[0050] 1) Construct a large-scale scRNA-seq reference dataset

[0051] We obtained 6 scRNA-seq datasets related to lung cancer from the Curated Cancer Cell Atlas (3CA) database. These datasets are all in the raw count format and are derived from Bischoff etal. 2021, Chan et al. 2021, Kim et al. 2020, Laughney et al. 2020, Qian et al.2020 and Xing et al. 2021. After quality control of these datasets, a scRNA-seq dataset containing 120 samples, 33410 genes and 378115 cells was finally obtained. We performed principal component analysis (PCA) and UMAP visualization on this dataset, and the results are as Figure 1 shown on the left. Subsequently, the scVI model was used to remove batch effects from the entire scRNA-seq dataset, and the obtained results are as Figure 1 shown on the right.

[0052] 2) Perform cell type annotation on the large-scale scRNA-seq reference dataset

[0053] First, we used the Leiden algorithm to perform clustering analysis on the scRNA-seq reference dataset, dividing it into multiple clusters, as shown in Figure 2 the left figure below. Subsequently, we used the dot plot shown in Figure 3 to display the highly expressed marker genes in each cluster, so as to determine the specific cell types of each cluster. Finally, we annotated the scRNA-seq reference dataset into 10 different cell types, namely T cells, Natural Killer cells, B cells, Plasma cells, Myeloid cells, Mast cells, Fibroblast cells, Endothelial cells, Epithelial cells, and Unidentified cells. Unidentified cells may be doublet cells or low-quality cells, so they will be discarded in subsequent analyses. The annotation results of each cell type are shown in Figure 2 the right figure below.

[0054] Next, we performed further subpopulation annotation on some cell types. For example, all T cells and NK cells were extracted from the scRNA-seq reference dataset, and the previous steps were repeated to further annotate T cells and NK cells as Quiescent T, Helper T, Cytotoxic T, Exhausted T, Regulatory T, and NK cells. Similarly, subpopulation annotation was performed on B cells, myeloid cells, fibroblasts, and endothelial cells; specifically, B cells were further annotated as Naive B, Memory B, GC B, fibroblasts were further annotated as iCAF, mCAF, pCAF, meCAF, pnCAF, and apCAF, endothelial cells were further annotated as Lymphatic EC and Vascular EC, and myeloid cells were further annotated as Neutrophil, CD14+Mono, CD16+Mono, CD14+CD16+Mono, pDC, cDC1, cDC2, INF TAM, Inflam TAM, Angio TAM, Prolif TAM, LA TAM, and TRM TAM.

[0055] Finally, we used the inferCNV algorithm to analyze the scRNA-seq reference dataset to identify malignant tumor cells. For Epithelial cells, cells with higher CNV scores were defined as malignant tumor cells, thus obtaining the finally annotated scRNA-seq reference dataset for subsequent deep learning network model training. The annotation results of the scRNA-seq reference set are as Figure 4 shown.

[0056] 3) Construction of pseudo-Bulk RNA-seq dataset

[0057] Randomly sample the annotated scRNA-seq reference dataset to construct a pseudo-Bulk RNA-seq dataset. The specific steps are as follows:

[0058] i. Randomly draw an integer between 1000 and 5000 as the total number of cells M in the Bulk sample.

[0059] ii. Randomly draw N numbers (N is the number of cell types finally annotated in the scRNA-seq reference dataset) between 0 and 1, and normalize these numbers so that their sum is 1, as the proportion P of each cell type in the Bulk sample.

[0060] iii. According to the total number of cells M and the cell type proportion P, calculate the number of cells T of each cell type. Then, randomly select Ti (i = 1, 2, 3 …, N) cells for each cell type from the scRNA-seq reference dataset.

[0061] iv. Calculate the cell type-specific gene expression for the randomly selected cells and obtain the pseudo-Bulk RNA-seq data through addition calculation.

[0062] Through the above steps, we constructed pseudo-Bulk RNA-seq datasets with sample sizes of 80,000 and 20,000 (or 100,000 and 10,000) for the training set and validation set respectively.

[0063] 4) Data preprocessing

[0064] First, standardize the scRNA-seq reference dataset. Specifically, we standardize the total reads of each cell to 10,000 and perform a logarithmic transformation (log1p). Then, use the Seurat_v3 algorithm in scanpy to screen for highly variable genes and select the top 20,000 highly variable genes. Next, perform a cell type classification addition operation on the screened dataset to generate a feature matrix of cell types and highly variable genes.

[0065] Next, according to these 20,000 highly variable genes, screen the pseudo-Bulk RNA-seq data and the pseudo cell-specific gene expression matrix to ensure that the genes in the pseudo-Bulk sample are consistent with the highly variable genes in the scRNA-seq reference dataset. Finally, perform min-max normalization on the pseudo-Bulk RNA-seq data and perform a log1p transformation on the pseudo cell-specific gene expression matrix.

[0066] 5) Training of the deep network model:

[0067] As Figure 5As shown, the preprocessed scRNA-seq feature matrix and pseudo-Bulk RNA-seq data are respectively input into two fully connected layers with residual connection modules: the Reference Encoder and the Bulk Encoder, for deep low-dimensional feature learning to obtain the latent feature Z ref and Z bulk . Subsequently, matrix multiplication and dot product operations are performed on these two latent features to calculate the cell proportion feature Z prop and the cell-specific gene expression feature Z expr . Finally, Z prop and Z expr are respectively input into the Prop Decoder and the Expr Decoder for decoding, so as to obtain the predicted cell proportion prop and the cell-specific gene expression matrix Expr.

[0068] The entire deep network model is optimized by minimizing the error loss, and the optimization objective is:

[0069]

[0070] where,

[0071] is the known true cell proportion in the pseudo-Bulk data;

[0072] is the known true cell-specific gene expression matrix in the pseudo-Bulk data;

[0073] is the cell proportion predicted by the Prop Decoder;

[0074] is the cell-specific gene expression matrix predicted by the Expr Decoder;

[0075] is a hyperparameter used to adjust the weight between the two parts of the loss;

[0076] represents the Mean Absolute Error, which is used to evaluate the difference between the predicted cell proportion and the true proportion;

[0077] represents the Mean Squared Error, which is used to evaluate the difference between the predicted cell-specific gene expression matrix and the true value.

[0078] Through the Backpropagation algorithm and the Gradient Descent method, the model parameters are continuously updated and optimized. In each training iteration, the model gradually reduces the gap between the predicted value and the true value, thereby improving the accuracy of the model in the deconvolution task. Finally, when the training reaches the preset stop condition or the error converges to a low level, the model weights with the best validation effect are saved for subsequent deconvolution analysis of real Bulk RNA-seq data.

[0079] In this embodiment, taking the lung cancer sample GSM3827136 in the GEO database as an example, a method for deconvoluting Bulk RNA-seq sequencing data of tumors based on deep learning includes the following steps:

[0080] Load the deep network model:

[0081] First, load the deep learning model pre-trained on the lung cancer scRNA-seq reference dataset and import the corresponding model weights. At the same time, obtain 20,000 highly variable genes screened during the pre-training process.

[0082] Preprocess the Bulk RNA-seq data:

[0083] Next, use the 20,000 highly variable genes screened during pre-training to screen the Rawcount data of Bulk RNA-seq. For genes that do not appear, fill them with zero values. Subsequently, perform min-max normalization on the screened Bulk RNA-seq data to ensure data standardization and consistency.

[0084] Prediction of the network model:

[0085] Input the preprocessed Bulk RNA-seq data into the pre-trained deep learning model, and the model will predict the proportions of each cell subpopulation in the Bulk sample and the corresponding cell type-specific gene expression matrix.

[0086] Annotation:

[0087] Bulk RNA-seq: A technique for measuring the average gene expression level of all cells in a tissue or cell population, providing macroscopic gene expression information but unable to distinguish the expression characteristics of individual cells.

[0088] scRNA-seq: Single-cell RNA sequencing technology, which analyzes gene expression at the single-cell level, reveals the characteristics and heterogeneity of different cell types, and identifies rare cell types.

[0089] Deep learning: Deep learning is a big data analysis method based on artificial neural networks, capable of automatically extracting features, learning complex patterns, and making predictions through the training of a large amount of data. In bioinformatics, deep learning can be used to process and analyze a large amount of genomic data to achieve tasks such as deconvolution, classification, and feature extraction.

[0090] Tumor microenvironment: The tumor microenvironment refers to the complex environment surrounding tumor cells, including cancer cells, immune cells, fibroblasts, vascular cells, matrix components, and secreted molecules, etc. These components interact with each other, affecting tumor growth, metastasis, immune escape, and response to treatment.

[0091] Tumor biology: Tumor biology is the science that studies tumor occurrence, development, metastasis, and the interaction with the host. It covers a wide range of fields from molecular mechanisms, gene regulation to cell behavior, helping to understand the basic mechanisms of tumors and develop new treatment methods.

[0092] Tumor heterogeneity: Tumor heterogeneity refers to the existence of various different cell types, gene mutations, expression patterns, and microenvironmental characteristics inside and outside the tumor. Heterogeneity makes the same type of tumor show significant differences in different patients, and also makes the same tumor have different evolutionary trajectories and treatment responses during the progression process.

[0093] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily achieved. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the embodiments shown and described here.

Claims

1. A deconvolution method for RNA-seq sequencing data of tumor Bulk samples based on a deep learning model, characterized in that, It includes the following steps: S1. Integrate multiple scRNA-seq datasets containing tumor cells, perform cell annotation, and construct a large-scale scRNA-seq reference dataset; S2. Simulate the cell composition of real Bulk RNA-seq to provide training and validation data for the model: randomly sample from the annotated scRNA-seq reference dataset, combine according to the specified cell ratios, and generate a pseudo-Bulk RNA-seq dataset; S3. Data preprocessing: normalize and perform gene screening on the scRNA-seq reference dataset and the pseudo-Bulk RNA-seq dataset to ensure the consistency of the input data in terms of feature expression and gene sets; S4. Input the preprocessed scRNA-seq reference dataset and pseudo-Bulk RNA-seq dataset into the reference set encoder and Bulk encoder respectively to extract the deep latent features representing the two datasets, and then use the decoder to deconvolve the extracted features into cell ratios and cell-specific gene expression matrices; S5. Continuously update and optimize the model parameters through the error backpropagation algorithm and gradient descent method: in each training iteration, evaluate the error between the model prediction and the real pseudo-Bulk data through the mean absolute error and mean squared error loss functions, gradually reduce the prediction error, and finally save the model weights with the best validation effect; In S4, input the preprocessed scRNA-seq feature matrix and pseudo-Bulk RNA-seq data into two fully connected layers with residual connection modules respectively: the reference set encoder and the Bulk encoder, to perform deep low-dimensional feature learning and obtain the latent features Zref and Zbulk; subsequently, perform matrix multiplication and dot product operations on these two latent features to calculate the cell ratio feature Zprop and the cell-specific gene expression feature Zexpr; finally, input Zprop and Zexpr into the cell ratio decoder and the cell-specific gene expression decoder respectively for decoding, so as to obtain the predicted cell ratio prop and the cell-specific gene expression matrix Expr; In S5, optimize the entire deep network model by minimizing the error loss, and its optimization objective is: ; Where: is the proportion of true cells known in the pseudo-Bulk data; is the known true cell-specific gene expression matrix in the pseudo-Bulk data; is the proportion of cells predicted by the Prop Decoder; is the cell-specific gene expression matrix predicted by the Expr Decoder; is a hyperparameter used to adjust the weight between two parts of the loss; Indicates the Mean Absolute Error, which is used to evaluate the difference between the predicted cell proportion and the true proportion; Denotes the Mean Squared Error, which is used to evaluate the difference between the predicted cell-specific gene expression matrix and the true values.

2. The deconvolution method for tumor Bulk sample RNA-seq sequencing data based on the deep learning model according to claim 1, characterized in that: In S1, perform principal component analysis and UMAP visualization on the scRNA-seq dataset, and use the scVI model to remove batch effects from the entire scRNA-seq dataset; Adopt the Leiden algorithm to perform clustering analysis on the scRNA-seq reference dataset, divide it into multiple clusters, display the highly expressed marker genes in each cluster through a dot plot, and thus determine the specific cell types of each cluster. Finally, annotate the scRNA-seq reference dataset as 10 different cell types, namely T cells, NK cells, B cells, plasma cells, myeloid cells, mast cells, fibroblasts, endothelial cells, epithelial cells, and unrecognized cells. Among them, the unrecognized cells will be discarded in subsequent analyses.

3. The deconvolution method for RNA-seq sequencing data of tumor Bulk samples based on a deep learning model according to claim 2, wherein Perform further subset annotation for some of the cell types: Extract all T cells and NK cells from the scRNA-seq reference dataset, repeat the previous steps, and further annotate T cells and NK cells as resting T cells, helper T cells, cytotoxic T cells, exhausted T cells, regulatory T cells, and NK cells; Similarly, perform subset annotation for B cells, myeloid cells, fibroblasts, and endothelial cells.

4. The deconvolution method for tumor Bulk sample RNA-seq sequencing data based on a deep learning model according to claim 3, characterized in that, Use the inferCNV algorithm to analyze the scRNA-seq reference dataset to identify malignant tumor cells; For epithelial cells, define cells with higher CNV scores as malignant tumor cells, thereby obtaining the final annotated scRNA-seq reference dataset for subsequent deep learning network model training.

5. The deconvolution method for tumor Bulk sample RNA-seq sequencing data based on a deep learning model according to claim 1, wherein In S2, randomly sample the annotated scRNA-seq reference dataset to construct a pseudo-BulkRNA-seq dataset. The specific steps are as follows: S2.

1. Randomly draw an integer between 1000 and 5000 as the total number of cells M in the Bulk sample. S2.

2. Randomly draw N numbers between 0 and 1, where N is the number of cell types finally annotated in the scRNA-seq reference dataset, and normalize these numbers so that their sum is 1, as the proportion P of each cell type in the Bulk sample. S2.

3. According to the total number of cells M and the cell type proportion P, calculate the number of cells T for each cell type; Then, randomly select Ti (i = 1, 2, 3 …, N) cells for each cell type from the scRNA-seq reference dataset. S2.

4. Calculate the cell type-specific gene expression for the randomly selected cells and obtain the pseudo-Bulk RNA-seq data through addition calculation. Through the above steps, construct pseudo-Bulk RNA-seq datasets with sample sizes of a and b for the training set and validation set respectively, where a > b.

6. The deconvolution method for tumor Bulk sample RNA-seq sequencing data based on a deep learning model according to claim 5, characterized in that In S3, perform standardization processing on the scRNA-seq reference dataset: Normalize the total reads of each cell to c and perform logarithmic transformation; Then, use the Seurat_v3 algorithm in scanpy to screen for highly variable genes, select the top d highly variable genes, and then perform cell type classification addition operation on the screened dataset to generate a feature matrix of cell types and highly variable genes. Next, according to these d highly variable genes, perform gene screening on the pseudo-Bulk RNA-seq data and the pseudo-cell-specific gene expression matrix to ensure that the genes in the pseudo-Bulk sample are consistent with the highly variable genes in the scRNA-seq reference dataset; Finally, perform min-max normalization on the pseudo-Bulk RNA-seq data and perform log1p transformation on the pseudo-cell-specific gene expression matrix; where a, b, c, and d are integer multiples of 10000.

7. The deconvolution method for tumor Bulk sample RNA-seq sequencing data based on a deep learning model according to claim 1, characterized in that When the training reaches the preset stopping condition or the error converges to a lower level, save the model weights with the best validation effect for subsequent deconvolution analysis of real Bulk RNA-seq data.

8. The deconvolution method for tumor Bulk sample RNA-seq sequencing data based on a deep learning model according to claim 1, characterized in that For the model weights, input the preprocessed Bulk RNA-seq data into the pre-trained deep learning model, and the model will predict the proportions of each cell subset in the Bulk sample, as well as the corresponding cell type-specific gene expression matrix.

Citation Information

Patent Citations

  • Methods of identifying cell-type-specific gene expression levels by deconvolving bulk gene expression

    US20230049525A1

  • Methods of identifying cell-type-specific gene expression levels by deconvolving bulk gene expression

    WO2021108556A1

  • Accurate robust information deconvolution from large numbers of tissue transcriptomes

    CN115136242A

  • Cancer drug response prediction method based on deep transfer learning

    CN118888007A