Method for determining signal channels of cell subpopulation and cell model, method for predicting efficacy of prescription and related device
By screening disease-related cell subpopulations and signaling pathways using low-depth single-cell transcriptome data and combining this with cell models of traditional Chinese medicine (TCM) effects, the problem of inaccurate cell subpopulation segmentation in TCM efficacy prediction was solved, achieving accurate prescription efficacy prediction and cost reduction.
Patent Information
- Application Number
- CN202511749406.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies lack the ability to finely classify cell subpopulations in predicting the efficacy of traditional Chinese medicine, resulting in insufficient accuracy and targeting of prescriptions. Furthermore, single-cell transcription sequencing is costly and difficult to apply on a large scale.
We used low-depth single-cell transcriptome data and employed methods such as pseudo-batch analysis to screen disease-related cell subpopulations and signaling pathways. We then combined these with cell models of the effects of traditional Chinese medicine to predict the efficacy of herbal formulas.
It improves the accuracy of cell models of drug action and disease-related signaling pathways, reduces costs, and enables precise prediction of prescription efficacy.
Smart Images

Figure CN121583467A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of traditional Chinese medicine, and in particular to a method for determining a signal pathway of a cell subpopulation and a cell model, a method for predicting the efficacy of a prescription, and related devices. BACKGROUND
[0002] Traditional Chinese medicine has the characteristics of multiple components, multiple targets, and multiple pathways. In order to reverse all the disorder signal pathways of a patient's disease state, the efficacy of a prescription is quantitatively predicted, so that the effect of the prescription can be seen and clearly explained, thereby providing a new modern method for evaluating the efficacy of a prescription in the field of traditional Chinese medicine. At present, when the efficacy of traditional Chinese medicine is predicted, the principle of reversing the disease signal pathway based on the different cell models of the drug is used to calculate the reversal rate and the efficacy of the prescription. Therefore, an important step is to determine the cell model of the drug and the fine disorder signal pathway of the patient's disease state.
[0003] In the past, the cell model was mainly determined based on expert experience or the cognition of the existing disease pathological mechanism. This method of selecting a cell model has the limitation of experience dependence, strong subjectivity, and lack of objective omics evidence support. In addition, it is common to analyze the omics data to obtain the disorder signal pathway. The common method is to analyze the conventional bulk transcriptome sequencing data to screen all the disorder signal pathways of the patient's disease state. However, since the disease signal pathway is screened based on the bulk transcriptome sequencing data, the fine division ability of the cell subpopulation is lacking, it is difficult to analyze the heterogeneity characteristics of the disease-related signal pathway at the single-cell resolution, the corresponding relationship between the cell subpopulation and the molecular signal pathway cannot be identified, and the accuracy and targeting of the prescription compatibility are insufficient. In addition, the current single-cell transcription sequencing has a relatively high data volume requirement. The recommended data volume of each sample is at least 90 Gb (double-end sequencing data volume), which leads to high cost and is not conducive to large-scale use. It is difficult to popularize and apply in the fields of drug screening and traditional Chinese medicine compound evaluation. For the above reasons, the accuracy of determining the cell model of the drug and the fine disorder signal pathway related to the disease is low, and the cost is high. SUMMARY
[0004] In view of the above problems, the present application provides a method for determining a signal pathway of a cell subpopulation and a cell model, a method for predicting the efficacy of a prescription, and related devices, to achieve the purpose of improving the accuracy of determining the cell model of the drug and the fine disorder signal pathway related to the disease. The specific scheme is as follows:
[0005] The first aspect of the present application provides a method for determining a signal pathway of a cell subpopulation and a cell model, comprising:
[0006] obtaining low-depth single-cell transcriptome data of a disease sample and low-depth single-cell transcriptome data of a normal control sample;
[0007] perform cell clustering and cell subpopulation identification operations on the low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples to obtain cell subpopulation identification results;
[0008] For the same cell subpopulation, the cell numbers of the cell subpopulation in the disease sample and the cell subpopulation in the normal control sample in the cell subpopulation identification results are compared to screen out a disease-related cell subpopulation with abnormal cell number change in the disease sample.
[0009] The gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample is summed, and the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample is summed to obtain summary data. The genes in the disease sample and the genes in the normal control sample in the summary data are subjected to differential gene analysis and signal pathway enrichment analysis to obtain a significance analysis result of a signal pathway in the disease sample.
[0010] The signal pathway that meets the requirement of a dysregulated signal pathway in the significance analysis result is screened out as a dysregulated signal pathway of the disease-related cell subpopulation.
[0011] Based on the correspondence between the disease-related cell subpopulation and a cell model of traditional Chinese medicine action, a cell model for prescription efficacy prediction is determined, and the dysregulated signal pathway of the disease-related cell subpopulation is used as a dysregulated signal pathway of the cell model for prescription efficacy prediction.
[0012] In one possible implementation, the cell clustering and cell subpopulation identification operations are performed on the low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples to obtain cell subpopulation identification results, including:
[0013] The low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples are respectively subjected to standardization, normalization, dimensionality reduction, clustering, and cell type annotation operations to obtain cell subpopulation identification results.
[0014] In one possible implementation, for the same cell subpopulation, the cell numbers of the cell subpopulation in the disease sample and the cell subpopulation in the normal control sample in the cell subpopulation identification results are compared to screen out a disease-related cell subpopulation with abnormal cell number change in the disease sample, including:
[0015] The cell numbers of each cell subpopulation of the disease samples and the cell numbers of each cell subpopulation of the normal control samples in the cell subpopulation identification results are counted.
[0016] calculate a quantity change ratio of the number of cells of the same cell subpopulation of the disease sample to the number of cells of the corresponding cell subpopulation of the normal control sample;
[0017] screen out the cell subpopulation with the quantity change ratio greater than the first ratio threshold or less than the second ratio threshold, and take it as a disease-related cell subpopulation.
[0018] In a possible implementation, the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample is summed, and the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample is summed to obtain summary data, and the genes of the disease sample and the genes of the normal control sample in the summary data are subjected to differential gene analysis and signal pathway enrichment analysis to obtain a significance analysis result of the signal pathway in the disease sample, including:
[0019] The gene expression data of all cells located in the same disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample is summed, and the gene expression data of all cells located in the same disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample is summed to obtain summary data by using a pseudo-batch analysis method.
[0020] The summary data is processed by using a differential gene analysis method to obtain an expression change fold of the gene.
[0021] The expression change fold of the gene is subjected to signal pathway enrichment analysis by using a gene set enrichment analysis method to obtain a signal pathway related to each disease-related cell subpopulation and a significance analysis result of the signal pathway.
[0022] The second aspect of the application provides a prescription efficacy prediction method, including:
[0023] obtaining the cell model for prescription efficacy prediction and the disorder signal pathway of the cell model for prescription efficacy prediction obtained by the determination method of the cell subpopulation and the signal pathway of the cell model.
[0024] Based on the signal pathway results of the cell model affected by traditional Chinese medicine and the disorder signal pathway of the cell model for prescription efficacy prediction, the efficacy is quantitatively predicted by calculating the reversal rate and the efficacy score of the candidate prescription.
[0025] In a possible implementation, based on the signal pathway results of the cell model affected by traditional Chinese medicine and the disorder signal pathway of the cell model for prescription efficacy prediction, the efficacy is quantitatively predicted by calculating the reversal rate and the efficacy score of the candidate prescription, including:
[0026] obtaining a candidate prescription;
[0027] acquire a signal pathway result of the single traditional Chinese medicine acting on the cell model for prescription efficacy prediction, and extract a standardized enrichment score of the single traditional Chinese medicine in the disordered signal pathway of the cell model for prescription efficacy prediction;
[0028] perform summation operation on the standardized enrichment scores of each traditional Chinese medicine in the candidate prescription in the disordered signal pathway of the cell model for prescription efficacy prediction, to obtain a standardized enrichment score of the candidate prescription;
[0029] Based on the standardized enrichment score of the candidate prescription, the efficacy is quantitatively predicted by calculating the reversal rate and the efficacy score of the candidate prescription.
[0030] The third aspect of the present application provides a device for determining a signal pathway of a cell subpopulation and a cell model, comprising:
[0031] A data acquisition module is configured to acquire low-depth single-cell transcriptome data of a disease sample and low-depth single-cell transcriptome data of a normal control sample.
[0032] A recognition module is configured to perform cell clustering and cell subpopulation recognition operation on the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample, to obtain a cell subpopulation recognition result.
[0033] A comparison module is configured to compare the number of cells of the cell subpopulation of the disease sample and the cell subpopulation of the normal control sample in the cell subpopulation recognition result for the same cell subpopulation, to screen out a disease-related cell subpopulation with abnormal cell number change in the disease sample.
[0034] An analysis module is configured to perform summation operation on gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample, and perform summation operation on gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample, to obtain summary data, and perform differential gene analysis and signal pathway enrichment analysis on genes of the disease sample and genes of the normal control sample in the summary data, to obtain a significance analysis result of the signal pathway in the disease sample.
[0035] A screening module is configured to screen out a signal pathway satisfying the requirement of disordered signal pathway from the significance analysis result, and take the signal pathway as a disordered signal pathway of the disease-related cell subpopulation.
[0036] The determining module is configured to determine a cell model for prescription efficacy prediction based on the correspondence between the disease-related cell subpopulation and the cell model affected by the traditional Chinese medicine, and to determine a disorder signal pathway of the disease-related cell subpopulation as a disorder signal pathway of the cell model for prescription efficacy prediction.
[0037] The fourth aspect of the present application provides a prescription efficacy prediction device, comprising:
[0038] The information acquisition module is configured to acquire the cell model for prescription efficacy prediction and the disorder signal pathway of the cell model for prescription efficacy prediction obtained by the above-mentioned determination method of the signal pathway of the cell subpopulation and the cell model.
[0039] The prescription efficacy prediction module is configured to quantitatively predict the efficacy of a candidate prescription based on the signal pathway results of the cell model affected by the traditional Chinese medicine and the disorder signal pathway of the cell model for prescription efficacy prediction by calculating the reversal rate and the efficacy score of the candidate prescription.
[0040] The fifth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0041] The memory is configured to store a computer program.
[0042] The processor is configured to execute the computer program to enable the electronic device to implement the above-mentioned determination method of the signal pathway of the cell subpopulation and the cell model, or implement the above-mentioned prescription efficacy prediction method.
[0043] The sixth aspect of the present application provides a computer storage medium, which carries one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can implement the above-mentioned determination method of the signal pathway of the cell subpopulation and the cell model, or implement the above-mentioned prescription efficacy prediction method.
[0044] By means of the technical solutions, the application provides a method for determining a signal pathway of a cell subpopulation and a cell model, a method for predicting the efficacy of a prescription, and related devices. In the application, low-depth single-cell transcriptome data of a disease sample and low-depth single-cell transcriptome data of a normal control sample are obtained, cell clustering and cell subpopulation identification operations are performed on the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample, and a cell subpopulation identification result is obtained. For the same cell subpopulation, the number of cells of the cell subpopulation of the disease sample and the cell subpopulation of the normal control sample in the cell subpopulation identification result are compared to screen out a disease-related cell subpopulation with abnormal cell quantity change in the disease sample. The gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample is summed, and the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample is summed to obtain summary data. The genes of the disease sample and the genes of the normal control sample in the summary data are subjected to differential gene analysis and signal pathway enrichment analysis to obtain a significance analysis result of a signal pathway in the disease sample. The signal pathway that meets the requirement of a disorder signal pathway in the significance analysis result is screened out as a disorder signal pathway of the disease-related cell subpopulation. Based on the correspondence between the disease-related cell subpopulation and a cell model of the action of traditional Chinese medicine, a cell model for predicting the efficacy of a prescription is determined. The disorder signal pathway of the disease-related cell subpopulation is used as a disorder signal pathway of the cell model for predicting the efficacy of a prescription. The entire process is automatically implemented. The cell model and the disorder signal pathway for predicting the efficacy of a prescription are determined based on single-cell transcriptome data analysis, the accuracy of the determination of the cell model and the disorder signal pathway for predicting the efficacy of a prescription is improved, and the efficacy of a prescription is predicted based on a cross-cell line multi-target synergistic intervention strategy. In addition, the data depth of the low-depth single-cell transcriptome data in the application is low. The low-depth data and the subsequent gene expression data are summed to accurately determine the disease-related cell subpopulation and the disorder signal pathway at a low data depth, reduce the cost, and realize fine disease target screening. BRIEF DESCRIPTION OF DRAWINGS
[0045] The above and other features, advantages, and aspects of the present disclosure will become more apparent by describing in detail the following specific embodiments thereof with reference to the attached drawings. Throughout the drawings, the same or similar reference numerals refer to the same or similar elements. It should be understood that the drawings are schematic and elements are not necessarily drawn to scale.
[0046] Figure 1 A flowchart of a method for determining a signal pathway of a cell subpopulation and a cell model provided by the application;
[0047] Figure 2aA schematic diagram of the correlation of NES values of signal pathways between different cell types under different sequencing amounts provided by the present application;
[0048] Figure 2b A schematic diagram of the correlation of NES values of signal pathways between different cell types under different sequencing amounts provided by the present application;
[0049] Figure 3 A schematic diagram of four disease-related cell subgroups and their disordered signal pathways of colorectal adenoma provided by the present application;
[0050] Figure 4 A flowchart of a prescription efficacy prediction method provided by the present application;
[0051] Figure 5 A flowchart of a method for quantitatively predicting efficacy by calculating the reversal rate and efficacy score of a candidate prescription provided by the present application;
[0052] Figure 6 A structural schematic diagram of a signal pathway determination device for a cell subpopulation and a cell model provided by the present application;
[0053] Figure 7 A structural schematic diagram of a prescription efficacy prediction device provided by the present application;
[0054] Figure 8 A structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0055] The embodiments of the present application are described below in conjunction with the accompanying drawings. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0056] The embodiments of the present application are described below in conjunction with the accompanying drawings. It is known to those of ordinary skill in the art that as technology develops and new scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0057] The terms “first”, “second”, and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or devices containing a series of units do not have to be limited to those units, but can include other units that are not clearly listed or inherent to these processes, methods, products or devices.
[0058] Traditional Chinese medicine has the characteristics of multi-component, multi-target and multi-pathway, and the principle of reversing all disorder signal pathways of the patient's disease state, and the drug efficacy of the prescription can be quantitatively predicted by calculating the reversal rate and drug efficacy of the candidate prescription on the disease-related signal pathway.
[0059] When evaluating the efficacy of the prescription, bulk mRNA sequencing (also known as bulk transcriptome sequencing) can be used. However, this method cannot identify disease-related cell subpopulations and has the problem of not being able to accurately depict cell subpopulation-specific molecular signal pathway targets. Specifically, bulk mRNA measures the mixed expression signals of all cells in the sample, lacks the ability to finely divide cell subpopulations, and is difficult to analyze the heterogeneity characteristics of disease-related signal pathways at single-cell resolution. Bulk transcriptome sequencing cannot distinguish the contributions of different cell types, and disease-related signal pathways may only be active in specific cell subpopulations, which can be easily masked by the expression noise of other high-abundance cells (such as stromal cells). When facing complex diseases, using a single cell model for evaluation makes it difficult to comprehensively and systematically evaluate the drug effect of traditional Chinese medicine combinations. In addition, bulk transcriptome sequencing cannot identify the correspondence between cell subpopulations and signal pathways, resulting in insufficient precision and targeting of prescription compatibility. Therefore, in this application, in order to distinguish different cell types, low-depth single-cell transcriptome data obtained by single-cell transcriptome sequencing (scRNA-seq) is used to analyze disease-related cell subpopulations and signal pathways, which is used to guide subsequent prescription efficacy evaluation. Single-cell transcriptome sequencing can perform gene expression analysis at single-cell resolution, more accurately screen the disorder signal pathways of different types of cells, avoid the interference of bulk cells, and thus achieve more accurate prescription efficacy prediction.
[0060] In addition, due to the complexity of the disease, when evaluating the efficacy of the prescription, a multi-target synergistic intervention strategy across cell models is used. However, the cell models currently used are mainly determined based on the understanding and experience of experts on the pathological mechanism of the disease, and this method of selecting cell models has the limitation of experience dependence and strong subjectivity. When the experience is insufficient, the selected cell model may not accurately reflect the pathological mechanism of the disease, resulting in low accuracy of the determination of the cell model, which is not conducive to accurate quantitative evaluation of the efficacy of the prescription.
[0061] To this end, the application introduces low-depth single-cell transcriptome sequencing data, compares single-cell transcriptome sequencing data of clinical patients with low-depth single-cell transcriptome data of normal healthy people, screens a cell subpopulation with a significantly changed number as a disease-related cell subpopulation, further adopts pseudobulk analysis, etc., obtains a standardized enrichment score of a disease-related cell subpopulation and a corresponding dysregulated molecular signaling pathway, and determines a cell model for prescription efficacy prediction and a dysregulated signaling pathway of the cell model for prescription efficacy prediction based on the related cell subpopulation and the dysregulated molecular signaling pathway, and carries out prescription efficacy evaluation.
[0062] On the basis of the above, an embodiment of the application provides a method for determining a signaling pathway of a cell subpopulation and a cell model, and the execution subject can be a processor, a controller, a server, or other devices with logical processing capability.
[0063] Reference Figure 1 A method for determining a signaling pathway of a cell subpopulation and a cell model can include:
[0064] S11, obtaining low-depth single-cell transcriptome data of a disease sample and low-depth single-cell transcriptome data of a normal control sample.
[0065] The disease sample refers to a diseased tissue, peripheral blood, etc. of a certain disease, and the normal control sample refers to a normal person or a normal control tissue adjacent to a diseased tissue or peripheral blood. When determining the low-depth single-cell transcriptome data, the data can be determined based on peripheral blood mononuclear cells (PBMCs), tissues, etc.
[0066] In an example, the disease sample is a colorectal adenoma tissue, and the normal control sample is an adjacent normal tissue. The colorectal adenoma tissue and the normal control tissue are obtained, single-cell library construction is performed, and sequencing is performed on the machine to obtain corresponding low-depth single-cell transcriptome data, including low-depth single-cell transcriptome data of the disease tissue and low-depth single-cell transcriptome data of the normal control tissue.
[0067] In this embodiment, the use of low-depth single-cell transcriptome data can break through the bulk transcriptome cell resolution limit.
[0068] The data depth of the low-depth single-cell transcriptome data is within a preset data depth range. The preset data depth range refers to a low data amount depth range, which is also a low sequencing depth range. In actual scenarios, the data amount of single-end sequencing bases in single-cell transcriptome data is usually about 50 Gb, and the preset data depth range in the application can be as low as 20 Gb, such as 20 Gb or more.
[0069] In actual implementation, in the embodiments of the present application, in order to avoid the problem of high cost caused by large amount of data samples, which is not conducive to large-scale use and is difficult to popularize and apply in drug screening and traditional Chinese medicine compound evaluation, low-sequencing-depth single-cell transcriptome data, i.e., low-depth single-cell transcriptome data, is used to simplify the amount of sample data.
[0070] In order to verify that the use of low-depth single-cell transcriptome data has high accuracy, in the embodiments of the present application, the high-depth cell transcriptome data of the patient's disease sample can be intercepted for subsequent analysis, and the amount of single-end data of the intercepted partial sequencing data can be 10 Gb or 20 Gb. The single-cell transcriptome data of the normal control sample of the patient is processed in the same way.
[0071] In order to verify that the low-depth single-cell transcriptome data can also meet the accuracy requirement, refer to Figure 2a and Figure 2b , Figure 2a and Figure 2b Pearson correlation coefficients of NES (Normalized Enrichment Score) values of disease signaling pathways identified in B lymphocytes and T lymphocytes are calculated under different data amounts, Figure 2a and Figure 2b In the diagonal line, 10 Gb, 20 Gb, and 50 Gb are used for analysis, and each point in the lower left part represents the NES value of the pathway in the corresponding data amount sample, and the right upper part represents the correlation coefficient. When the data amount is 20 Gb, in the B lymphocytes, the correlation of the disease pathway NES values obtained by analyzing the 20 Gb data and the 50 Gb data reaches 0.991, and in the T lymphocytes, the correlation of the disease pathway NES values obtained by analyzing the 20 Gb data and the 50 Gb data reaches 0.943, that is, when using low-depth single-cell transcriptome data, the disease signaling pathways screened by different cell types have high correlation. It shows that at low sequencing depth, the signaling pathways screened by each cell type have consistency and can be used for subsequent prescription evaluation and optimization.
[0072] S12, cell clustering and cell subpopulation identification operations are performed on the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample to obtain cell subpopulation identification results.
[0073] In the embodiments, in order to distinguish which cell types of the disease sample and the normal control sample have changed, cell clustering operation needs to be performed on the low-depth single-cell transcriptome data to identify the changed cell types.
[0074] In an implementation manner, step S12 can include:
[0075] The low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample are respectively standardized, normalized, dimensionally reduced, clustered, and cell type annotated to obtain a cell subpopulation recognition result.
[0076] In this embodiment, after obtaining the low-depth single-cell transcriptome data, the low-depth single-cell transcriptome data can be analyzed based on preset statistical indicators such as the number of genes, UMI (Unique Molecular Identifier) count, and the like, to obtain indicator values of these indicators, and the high-quality single cells are screened by using the indicator values.
[0077] Then, the gene expression data in the low-depth single-cell transcriptome data is standardized and normalized, and the values of indicators such as the coefficient of variation are calculated to screen high-coefficient-of-variation genes.
[0078] Then, the high-coefficient-of-variation genes are dimensionally reduced. In this embodiment, when the dimensional reduction operation is specifically performed, the high-coefficient-of-variation genes can be dimensionally reduced by using principal component analysis (PCA) or other dimensional reduction techniques such as t-SNE (t-Distributed Stochastic Neighbor Embedding) or UMAP (Uniform Manifold Approximation and Projection).
[0079] Finally, the dimensionally reduced data is clustered to obtain a clustering result.
[0080] In this embodiment, the Louvain clustering, Leiden clustering, and the like in the Scanpy analysis package are used for the clustering operation to identify cell subpopulations and obtain a clustering result.
[0081] In an actual scenario, the types and / or quantities of at least one cell subpopulation of the disease sample and at least one cell subpopulation of the normal control sample in the clustering result can be the same or different.
[0082] It should be noted that after obtaining the clustering result, cell subpopulations with a cell quantity greater than 500 can be screened for subsequent operations, and cell subpopulations with a small cell quantity have a small impact on the disease and can be ignored.
[0083] After obtaining the clustering result, the cell subpopulations in the clustering result can be subjected to cell type annotation operation to obtain a cell subpopulation recognition result.
[0084] In a specific implementation, SingleR, CellTypist, or other analysis software can be used to annotate the cell types in the cell clustering result to obtain the cell subpopulation identification result.
[0085] In step S13, the number of cells in the cell subpopulation of the disease sample and the number of cells in the cell subpopulation of the normal control sample in the cell subpopulation identification result are compared for the same cell subpopulation to screen out a disease-related cell subpopulation with an abnormal change in the number of cells in the disease sample.
[0086] In the embodiments of the present application, the number of cells of different cell types in the cell subpopulation identification result can be counted to determine a cell subpopulation with a large change in the number of cells in the cell subpopulation identification result of the disease sample and the normal control sample, which is the identified main cell subpopulation related to the disease, and can be referred to as a disease-related cell subpopulation.
[0087] In an implementation, step S13 can include:
[0088] 1) Count the number of cells of each cell subpopulation in the disease sample and the number of cells of each cell subpopulation in the normal control sample in the cell subpopulation identification result.
[0089] In the embodiments, after obtaining the cell subpopulation identification result, the number of cells of each cell subpopulation in the disease sample and the number of cells of each cell subpopulation in the normal control sample in the cell subpopulation identification result can be counted. In a specific counting process, since the cell subpopulations have been clustered together through the clustering operation, the number of cells in the cluster where the cell subpopulation is located in the cell subpopulation identification result can be directly counted.
[0090] 2) Calculate the number change ratio of the number of cells in the cell subpopulation of the disease sample and the number of cells in the corresponding same cell subpopulation of the normal control sample.
[0091] In the embodiments, it is assumed that the disease sample and the normal control sample each include A, B, C, and D cell subpopulations, and the number change ratio of the number of A cells in the disease sample and the number of A cells in the normal control sample is calculated. The number change ratio can be the number of A cells in the disease sample / the number of A cells in the normal control sample. The calculation process of the number change ratios of B, C, and D cell subpopulations is similar.
[0092] 3) Screen out a cell subpopulation with a number change ratio greater than a first ratio threshold or less than a second ratio threshold, and use it as a disease-related cell subpopulation.
[0093] In the embodiments of the present application, the first proportion threshold and the second proportion threshold are pre-configured, the first proportion threshold is greater than the second proportion threshold, that is, the first proportion threshold is a threshold with a larger value, and the second proportion threshold is a threshold with a smaller value. If the quantity change proportion is greater than the first proportion threshold, it indicates that the cell quantity of the cell subpopulation in the disease sample is significantly increased compared with the cell quantity of the cell subpopulation in the normal control sample. In a special example, if the cell quantity of a cell subpopulation is large in the disease sample and is zero in the normal control sample, the quantity change proportion is infinite, and in this case, it is also considered that the quantity change proportion is greater than the first proportion threshold.
[0094] If the quantity change proportion is less than the second proportion threshold, it indicates that the cell quantity of the cell subpopulation in the disease sample is significantly reduced compared with the cell quantity of the cell subpopulation in the normal control sample.
[0095] Whether the cell quantity is significantly increased or significantly reduced, it indicates that the quantity change of the cell subpopulation is abnormal, and it also indicates that the cell subpopulation of the patient is different from that of a normal person, that is, the cell subpopulation is a disease-related cell subpopulation, that is, the cell subpopulation can be determined as the disease-related cell subpopulation in the embodiments of the present application.
[0096] In an example, single-cell transcriptome data of colorectal adenoma lesion tissue with data number GSE161277 and single-cell transcriptome data of paired normal tissue are obtained from a GEO (Gene Expression Omnibus, Gene Expression Omnibus) database, the cell detection numbers of different cell types between the three colorectal adenoma tissues and the normal control tissues are analyzed, and the quantity change proportions of the detection numbers of each cell type under the disease state are compared, and the results are as shown in Table 1.
[0097] Table 1 Cell detection of different cell types between colorectal adenoma tissue and normal control tissue
[0098]
[0099] Among them, the Adenoma sample is the colorectal adenoma tissue, and the Normal sample is the normal control tissue.
[0100] Further filter out the cell subpopulation with cell number less than 500 in the colorectal adenoma and normal control samples, so as to eliminate the small amount of cell subpopulation, then compare the colorectal adenoma tissue with the normal control tissue, and screen out the cell subpopulation with the number change ratio ≥1.5 times or ≤0.75 times, which is set as the disease-related cell subpopulation with obvious change. As can be seen from Table 1, the epithelial cell subpopulation and B lymphocyte subpopulation in the colorectal adenoma tissue are obviously increased, while the T lymphocyte subpopulation and natural killer cell subpopulation are obviously decreased, and the four cell subpopulations can be used as the disease-related cell subpopulation in the embodiment.
[0101] In the embodiment, the differences in cell composition between the disease tissue and the normal control tissue are compared, and the cell subpopulation with significant change in number or ratio is screened out, which is used to guide the subsequent quantitative prediction of the efficacy of the prescription based on the reversal of the disease signal pathway as the core.
[0102] S14, summing the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample and summing the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample to obtain summary data, and performing differential gene analysis and signal pathway enrichment analysis on the genes of the disease sample and the genes of the normal control sample in the summary data to obtain a significant analysis result of the signal pathway in the disease sample.
[0103] In the embodiment, after analyzing the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample to obtain the disease-related cell subpopulation related to the disease, in one implementation, the gene expression data of all cells located in the same disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample can be summed, and the gene expression data of all cells located in the same disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample can be summed to obtain summary data. Specifically, the expression matrix of each cell in the disease-related cell subpopulation is extracted, and the gene expression data of all cells of each sample in the same cell subpopulation is summed by using the pseudobulk method to generate a “pseudobulk” expression profile. When summing, the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample and the normal control sample can be summed respectively to obtain summary data.
[0104] Then, the summary data is processed by using a differential gene analysis method to obtain the expression change fold of the gene.
[0105] Specifically, a differential gene analysis method such as DEseq, DEseq2, etc. is used to perform differential gene analysis on genes of the disease samples and the normal control samples in the aggregated data from the molecular feature perspective, to obtain the expression change fold of the genes.
[0106] Finally, a gene set enrichment analysis method is used to perform enrichment analysis of the signal pathways on the expression change fold of the genes, to obtain the signal pathways related to each disease-related cell subpopulation, and the significance analysis result of the signal pathways.
[0107] In this embodiment, a gene set enrichment analysis (GSEA) is used to perform enrichment analysis of the KEGG (Kyoto Encyclopedia of Genes and Genomes) signal pathways on the expression change fold of the genes of the disease-related cell subpopulation, to obtain the disease signal pathways related to each disease-related cell subpopulation, and the standardized enrichment score NES and the significance analysis result of the disease signal pathways.
[0108] Wherein, the NES is positive, indicating that the disease-related gene set is enriched at the top of the genes (high expression value). The NES is negative, indicating that the gene set is enriched at the bottom of the genes (low expression value).
[0109] Wherein, the significance analysis result can be filtered and selected by FDR (False Discovery Rate).
[0110] S15, a signal pathway whose significance analysis result meets the requirement of the disorder signal pathway is selected as the disorder signal pathway of the disease-related cell subpopulation.
[0111] Wherein, the requirement of the disorder signal pathway can be:
[0112] FDR<=0.25.
[0113] That is, the signal pathway with FDR<=0.25 can be selected as the disorder signal pathway of the disease-related cell subpopulation. The signal pathway with FDR>0.25 indicates that the pathway is not significantly enriched, and cannot be used as the disorder signal pathway of the disease-related cell subpopulation, and the NES value thereof is set to 0.
[0114] It should be noted that different disease-related cell subpopulations have their own corresponding disorder signal pathways, and the disorder signal pathways of different disease-related cell subpopulations can be the same or different.
[0115] In an example, the two cell subpopulations (epithelial cells and B lymphocytes) that are significantly increased and the two cell subpopulations (T lymphocytes and natural killer cells) that are significantly decreased in the colorectal adenoma tissue analyzed in the above embodiment are subjected to pathway enrichment analysis, and a pathway enrichment cluster diagram is as shown in Figure 3 The signal pathways enriched by epithelial cells mainly include oxidative phosphorylation, pyruvate metabolism, mRNA surveillance pathway, RNA degradation, DNA replication, spliceosome, proteasome, nucleotide excision repair, cell cycle, ubiquitin-mediated proteolysis, protein processing in endoplasmic reticulum, cell adhesion molecules, apoptosis. The pathways enriched by B lymphocytes mainly include oxidative phosphorylation, RNA degradation, DNA replication, spliceosome, proteasome, cell cycle, p53 signaling pathway, protein processing in endoplasmic reticulum, endocytosis, antigen processing and presentation, hematopoietic cell lineage, Th1 and Th2 cell differentiation, Th17 cell differentiation, IgA-producing intestinal immune network. The pathways enriched by T lymphocytes mainly include oxidative phosphorylation, mRNA surveillance pathway, spliceosome, NF-κB signaling pathway, FoxO signaling pathway, cell cycle, endocytosis, apoptosis, antigen processing and presentation, NOD-like receptor signaling pathway, IL-17 signaling pathway, Th1 and Th2 cell differentiation, Th17 cell differentiation, T cell receptor signaling pathway, TNF signaling pathway, PD-L1 expression in cancer and PD-1 checkpoint pathway. The pathways enriched by natural killer cells mainly include pentose phosphate pathway, RNA degradation, DNA replication, spliceosome, proteasome, NF-κB signaling pathway, cell cycle, ubiquitin-mediated proteolysis, protein processing in endoplasmic reticulum, endocytosis, apoptosis, ferroptosis, antigen processing and presentation, C-type lectin receptor signaling pathway, TNF signaling pathway.
[0116] S16, based on the correspondence between the disease-related cell subpopulation and the cell model of traditional Chinese medicine action, determine the cell model for prescription efficacy prediction, and use the disorder signal pathway of the disease-related cell subpopulation as the disorder signal pathway of the cell model for prescription efficacy prediction.
[0117] In actual scenarios, when the efficacy of a prescription is predicted, the signal pathways affected by the drug acting on different cell models can be obtained. A single cell model refers to obtaining the effect pathway spectrum of a drug acting on a cell model, and a multi-cell model refers to obtaining the effect pathway spectrum of a drug acting on multiple cell models. In the embodiments of the present application, a multi-cell model is used to predict the efficacy of a prescription, and therefore, a plurality of single cell models need to be pre-configured, each disease-related cell subpopulation corresponds to a single cell model, and the correspondence between the disease-related cell subpopulation and the cell model of the action of traditional Chinese medicine can be established based on biological knowledge and the source of the cell line. Therefore, in an implementation manner, the cell model corresponding to the disease-related cell subpopulation for prescription efficacy prediction can be obtained based on the correspondence between the disease-related cell subpopulation and the cell model of the action of traditional Chinese medicine, and the disorder signal pathway of the disease-related cell subpopulation is used as the disorder signal pathway of the cell model for prescription efficacy prediction. In addition, the influence (NES value) of the drug acting on the corresponding cell model on the disorder signal pathway can also be obtained based on the disorder signal pathway of the disease-related cell subpopulation.
[0118] In an example, as shown in Table 1, the Jurkat cell model corresponds to the cell model of the T lymphocyte subpopulation, the Raji cell model corresponds to the cell model of the B lymphocyte subpopulation, the NK92Mi cell model corresponds to the cell model of the natural killer cell subpopulation, and the Caco2 cell model corresponds to the cell model of the epithelial cell subpopulation.
[0119] The disorder signal pathway to be extracted when the drug acts on the Jurkat cell model is the disorder signal pathway of the T lymphocyte subpopulation, the disorder signal pathway to be extracted when the drug acts on the Raji cell model is the disorder signal pathway of the B lymphocyte subpopulation, the disorder signal pathway to be extracted when the drug acts on the NK92Mi cell model is the disorder signal pathway of the natural killer cell subpopulation, and the disorder signal pathway to be extracted when the drug acts on the Caco2 cell model is the disorder signal pathway of the epithelial cell subpopulation.
[0120] In actual scenarios, the disease-related cell subpopulation can be mined by analyzing low-depth single-cell transcriptome data, and then when the drug combination is evaluated, the cell model corresponding to the disease-related cell subpopulation for prescription efficacy prediction can be used. The cell model has an effect pathway spectrum of different drugs acting on the cell model, the signal pathway used for prescription efficacy prediction is determined based on the disorder signal pathway of the disease-related cell subpopulation, and the cell model and the signal pathway are used for subsequent prescription efficacy prediction.
[0121] In this embodiment, low-depth single-cell transcriptome data of disease samples and low-depth single-cell transcriptome data of normal control samples are obtained, cell clustering and cell subpopulation identification operations are performed on the low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples, cell subpopulation identification results are obtained, for the same cell subpopulation, the cell numbers of the cell subpopulation of the disease samples and the cell subpopulation of the normal control samples in the cell subpopulation identification results are compared to screen out disease-related cell subpopulations with abnormal cell number changes in the disease samples, gene expression data in the disease-related cell subpopulations in the low-depth single-cell transcriptome data of the disease samples are summed, and gene expression data in the disease-related cell subpopulations in the low-depth single-cell transcriptome data of the normal control samples are summed to obtain summary data, difference gene analysis and signal pathway enrichment analysis are performed on genes of the disease samples and genes of the normal control samples in the summary data, significant analysis results of signal pathways in the disease samples are obtained, signal pathways meeting the requirements of the dysregulated signal pathways in the significant analysis results are screened out as dysregulated signal pathways of the disease-related cell subpopulations, based on the correspondence between the disease-related cell subpopulations and the cell models of the action of traditional Chinese medicine, a cell model used for prescription efficacy prediction is determined, and the dysregulated signal pathways of the disease-related cell subpopulations are used as the dysregulated signal pathways of the cell model for prescription efficacy prediction. The whole process is automatically implemented, the cell model and the dysregulated signal pathway used for prescription efficacy prediction are determined based on single-cell transcriptome data analysis, the accuracy of the determination of the cell model and the dysregulated signal pathway used for prescription efficacy prediction is improved, and then the prescription efficacy is predicted based on the cross-cell line multi-target synergistic intervention strategy. In addition, the low-depth single-cell transcriptome data in this application has low data depth, and the summation operation of the low-depth data and the subsequent gene expression data can accurately determine the disease-related cell subpopulation and the dysregulated signal pathway under low data depth, has low cost, and can realize fine disease target screening.
[0122] In addition, based on the low-depth single-cell transcriptome data analysis of clinical patients and normal control populations, the application screens cell subpopulations and dysregulated signal pathways with significant changes in cell number under disease conditions, which are used to guide the selection of cell models and signal pathways when performing prescription efficacy prediction.
[0123] Based on the above-mentioned embodiment of the method for determining the cell subpopulation and the signal pathway of the cell model, another embodiment of the application provides a prescription efficacy prediction method, which refers to Figure 4 may include:
[0124] S21, obtaining low-depth single-cell transcriptome data of disease samples and low-depth single-cell transcriptome data of normal control samples.
[0125] S22, cell clustering and cell subpopulation identification operations are performed on the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample to obtain cell subpopulation identification results.
[0126] S23, for the same cell subpopulation, the cell numbers of the cell subpopulation of the disease sample and the cell subpopulation of the normal control sample in the cell subpopulation identification results are compared to screen out disease-related cell subpopulations with abnormal cell number changes in the disease sample.
[0127] S24, the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample is summed, and the gene expression data in the disease-related cell subpopulation in the low-depth single-cell transcriptome data of the normal control sample is summed to obtain summary data. The genes of the disease sample and the genes of the normal control sample in the summary data are subjected to differential gene analysis and signal pathway enrichment analysis to obtain a significant analysis result of the signal pathway in the disease sample.
[0128] S25, screening out the signal pathway that meets the requirement of the disorder signal pathway in the significant analysis result, and taking it as the disorder signal pathway of the disease-related cell subpopulation.
[0129] S26, based on the corresponding relationship between the disease-related cell subpopulation and the cell model of the action of traditional Chinese medicine, determining the cell model for prescription efficacy prediction, and taking the disorder signal pathway of the disease-related cell subpopulation as the disorder signal pathway of the cell model for prescription efficacy prediction.
[0130] Wherein, the specific implementation of steps S21-S26, please refer to the above description. In this embodiment, the cell model for prescription efficacy prediction and the disorder signal pathway of the cell model for prescription efficacy prediction need to be obtained.
[0131] In this embodiment, still taking the above-mentioned embodiment as an example, according to the above-mentioned embodiment, 4 disease-related cell subpopulations related to colorectal adenoma are obtained, which are epithelial cells, T lymphocytes, B lymphocytes and natural killer cells. The Jurkat cell model is the cell model for prescription efficacy prediction corresponding to the T lymphocyte subpopulation, the Raji cell model is the cell model for prescription efficacy prediction corresponding to the B lymphocyte subpopulation, the NK92Mi cell model is the cell model for prescription efficacy prediction corresponding to the natural killer cell subpopulation, and the Caco2 cell model is the cell model for prescription efficacy prediction corresponding to the epithelial cell subpopulation. In addition, 4 disease-related cell subpopulations corresponding disorder signal pathways are obtained, which are the disorder signal pathways of the cell model for prescription efficacy prediction.
[0132] S27, the signal pathway results based on the traditional Chinese medicine action cell model and the disorder signal pathway of the cell model for prescription efficacy prediction, the efficacy is quantitatively predicted by calculating the reversal rate and the efficacy of the candidate prescription.
[0133] In this embodiment, after obtaining the cell model for prescription efficacy prediction and the disorder signal pathway of the cell model for prescription efficacy prediction, the disorder signal pathway for prescription efficacy prediction and the cell model are used for prescription efficacy prediction.
[0134] In one implementation manner, with reference to Figure 5 , step S27 can include:
[0135] S31, obtaining a candidate prescription.
[0136] The candidate prescription refers to a compound or single prescription composed of at least one traditional Chinese medicine, and the efficacy thereof has not been evaluated by the cell model described in the method, and needs to be predicted by subsequent steps. The number of drugs in the candidate prescription is greater than or equal to 1.
[0137] The source of the candidate prescription includes but is not limited to the following ways: (1) clinical commonly used empirical formula; (2) classical prescription recorded in ancient literature, etc.
[0138] In this embodiment, the empirical formula of Shengbai granules for treating intestinal adenoma is taken as an example of a candidate prescription for subsequent evaluation.
[0139] S32, using the single herb in the candidate prescription to process the cell model for prescription efficacy prediction, obtaining the signal pathway results of the single herb acting on the cell model, and using the disorder signal pathway of the cell model for prescription efficacy prediction to extract the standardized enrichment score of the single herb in the disorder signal pathway of the cell model for prescription efficacy prediction.
[0140] Specifically, the traditional Chinese medicine extract in the candidate prescription, i.e. each single herb, is used to process Jurkat cell model (T lymphocyte subpopulation), Raji cell model (corresponding to B lymphocyte subpopulation), NK92Mi cell model (corresponding to natural killer cell subpopulation), and Caco2 cell model (corresponding to epithelial cell subpopulation), and the transcriptional sequencing is performed and the transcriptional data is analyzed to obtain the signal pathway results, thereby obtaining the standardized enrichment score NES and FDR of each traditional Chinese medicine in the candidate prescription in the disorder signal pathway for prescription efficacy prediction on the cell model corresponding to the disease-related cell subpopulation.
[0141] Wherein, FDR<=0.25 is considered to be significantly enriched, NES is positive, indicating that the gene set of the traditional Chinese medicine is enriched at the top of the gene (high expression value), and NES is negative, indicating that the gene set is enriched at the bottom of the gene (low expression value), thereby obtaining the change direction of the traditional Chinese medicine in the disorder signal pathway.
[0142] S33, summing the normalized enrichment scores of each traditional Chinese medicine in the candidate prescription on the disorder signal pathway of the cell model for prescription efficacy prediction to obtain the normalized enrichment score of the candidate prescription.
[0143] Specifically, in order to quantitatively evaluate the effect of the candidate prescription on different cell models, two indexes of reversal rate and efficacy score are used for evaluation. First, for a disorder signal pathway in a cell model, the NES value of each drug in each candidate prescription on the disorder signal pathway of the cell model is added to represent the strength of the prescription in regulating the disorder signal pathway.
[0144] For example, for a cell model, there are ten drugs in the candidate prescription, and the NES value of each drug on the disorder signal pathway is obtained by treating the cell model, and then the NES values of the same disorder signal pathway obtained by treating the cell model with the ten drugs are added to obtain the NES value of each disorder signal pathway of the candidate prescription on the cell model.
[0145] For example, a cell model corresponds to five disorder signal pathways, namely disorder signal pathways 1-5, and the NES values of the ten drugs on the disorder signal pathway 1 are added to obtain the NES value of the candidate prescription on the disorder signal pathway 1, and the processing logic of the disorder signal pathways 2-5 is the same.
[0146] The processing procedures of the remaining models are the same. Through the above steps, the NES value of the candidate prescription on each disorder signal pathway corresponding to each cell model can be obtained.
[0147] S34, based on the normalized enrichment score of the candidate prescription, the efficacy is quantitatively predicted by calculating the reversal rate and the efficacy score of the candidate prescription.
[0148] Specifically, for each disorder signal pathway, if the NES value of the disease sample on the disorder signal pathway is opposite to the NES value of the candidate prescription on the disorder signal pathway, it indicates that the signal pathway is reversed. According to whether it is reversed, the reversal rate of each cell model can be calculated, and the reversal rate is equal to the number of reversed signal pathways divided by the total number of disorder signal pathways.
[0149] For each cell model, the efficacy score = the sum of the absolute values of the NES of the reversed disorder signal pathway - the sum of the absolute values of the NES of the unreversed disorder signal pathway.
[0150] Through the above steps, the reversal rate and the efficacy score corresponding to each cell model can be calculated.
[0151] Subsequently, the efficacy of the prescription can be evaluated based on the reversal rate and the efficacy score of the candidate prescription corresponding to each cell model.
[0152] In this embodiment, the corresponding reversal rate and efficacy score of the candidate prescription in each cell model are obtained, and then the average reversal rate and average efficacy score of the candidate prescription are calculated. The specific calculation method is: average reversal rate = sum of reversal rates corresponding to each cell model / total number of cell models; average efficacy score = sum of efficacy scores corresponding to each cell model / total number of cell models.
[0153] In this embodiment, the molecular efficacy of the prescription is evaluated based on the reversal of the disordered signal pathway, and the clinical efficacy is predicted.
[0154] In one example, the experienced prescription Shengbai Granules for treating intestinal adenoma is evaluated based on disease-related cell subgroups and disordered signal pathways of intestinal adenoma patients. According to the average reversal rate of the four disease-related cell subgroups, the average reversal rate of Shengbai Granules is 57%.
[0155] Table 2 Evaluation of Shengbai Granules Prescription
[0156]
[0157] Currently, Shengbai Granules is used in clinical practice to prevent metachronous colorectal tumors after polypectomy. A multi-center, randomized, double-blind, placebo-controlled clinical trial has been conducted in multiple hospitals to strictly evaluate the efficacy and safety of Shengbai Granules. Patients who have recently undergone complete polypectomy and have been diagnosed with adenoma are randomly assigned (1:1) to receive Shengbai Granules or placebo treatment twice a day for 6 months. During the 2-year follow-up period, a colonoscopy is performed once a year. The primary outcome is the proportion of patients with at least one adenoma detected in the adjusted intention-to-treat population at 2 years of follow-up. Using logistic regression, the data of 336 adjusted intention-to-treat populations were analyzed, and there was a significant difference in the proportion of patients with at least one recurrent adenoma between the treatment group and the placebo group (42.5% vs. 58.6%; OR, 0.47; 95% CI, 0.29-0.74; p=0.001) (Phytomedicine. 2024 May; 127: 155496.). From this clinical trial, it can be seen that 57.5% of colorectal adenoma patients in the Shengbai Granules treatment group did not relapse 2 years after surgery, and the clinical efficacy value is consistent with the molecular efficacy evaluation value of Shengbai Granules (average reversal rate: 57%).
[0158] In this embodiment, the molecular efficacy of the prescription is evaluated based on the reversal of the disordered signal pathway, and the clinical efficacy is predicted, and the predicted value is consistent with the clinical research efficacy value.
[0159] In summary, the embodiment of the present application analyzes the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample, aggregates the same cell subpopulation through cell clustering, screens the cell subpopulation with a significant change in quantity as a disease-related cell subpopulation, further performs differential gene expression analysis, obtains the change of the disease-related cell subpopulation and the corresponding dysregulated signal pathway related to the disease, and develops prescription efficacy evaluation based on the disease-related cell subpopulation and the dysregulated signal pathway, which can reduce the cost while stably identifying the signal pathway changes with biological significance.
[0160] In addition, the low-depth single-cell transcriptome sequencing data is integrated into the prescription efficacy prediction process, and the disease target is accurately screened based on the low-depth sequencing single-cell transcriptome data to establish a research and development path of "low-depth single-cell sequencing-cell subpopulation analysis-feature cell pathway target identification-prescription efficacy prediction".
[0161] In addition, the resolution limit of bulk transcriptome is broken through, the molecular characteristics of different cell subpopulations are analyzed, the basis for multi-target synergistic intervention is provided, and the prescription efficacy prediction is performed based on the principle of reversing the disease-related signal pathway.
[0162] On the basis of the embodiment of the method for determining the signal pathway of the cell subpopulation and the cell model, another embodiment of the present application provides a device for determining the signal pathway of the cell subpopulation and the cell model, which refers to Figure 6 may include:
[0163] The data acquisition module 11 is configured to acquire the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample.
[0164] The identification module 12 is configured to perform cell clustering and cell subpopulation identification operations on the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample to obtain a cell subpopulation identification result.
[0165] The comparison module 13 is configured to compare the cell quantities of the cell subpopulation of the disease sample and the cell subpopulation of the normal control sample in the cell subpopulation identification result for the same cell subpopulation, so as to screen out a disease-related cell subpopulation with abnormal change in cell quantity in the disease sample.
[0166] The analysis module 14 is configured to perform summation operation on gene expression data in a disease-related cell subpopulation in low-depth single-cell transcriptome data of a disease sample, and perform summation operation on gene expression data in a disease-related cell subpopulation in low-depth single-cell transcriptome data of a normal control sample, to obtain summary data, perform differential gene analysis and signal pathway enrichment analysis on genes in the disease sample and genes in the normal control sample in the summary data, and obtain a significant analysis result of a signal pathway in the disease sample;
[0167] The screening module 15 is configured to screen a signal pathway that meets a disorder signal pathway requirement in the significant analysis result, and take the signal pathway as a disorder signal pathway of the disease-related cell subpopulation.
[0168] The determination module 16 is configured to determine a cell model for prescription efficacy prediction based on a corresponding relationship between the disease-related cell subpopulation and a cell model of a traditional Chinese medicine action, and take the disorder signal pathway of the disease-related cell subpopulation as a disorder signal pathway of the cell model for prescription efficacy prediction.
[0169] In an implementation manner, the identification module 12 is specifically configured to:
[0170] The low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample are respectively subjected to standardization, normalization, dimension reduction, clustering and cell type annotation operation, to obtain a cell subpopulation identification result.
[0171] In an implementation manner, the comparison module 13 includes:
[0172] The statistical sub-module is configured to count the number of cells of each cell subpopulation of the disease sample and the number of cells of each cell subpopulation of the normal control sample in the cell subpopulation identification result.
[0173] The calculation sub-module is configured to calculate a number change ratio of the number of cells of a cell subpopulation of the disease sample and a corresponding same cell subpopulation of the normal control sample.
[0174] The screening sub-module is configured to screen a cell subpopulation with a number change ratio greater than a first ratio threshold or less than a second ratio threshold, and take the cell subpopulation as a disease-related cell subpopulation.
[0175] In an implementation manner, the analysis module 14 includes:
[0176] The summation sub-module is configured to perform summation operation on gene expression data of all cells in a same disease-related cell subpopulation in the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample by using a pseudo-batch analysis method, to obtain summary data.
[0177] a processing submodule, configured to process the aggregated data by using a differential gene analysis method to obtain expression change folds of genes;
[0178] an analysis submodule, configured to perform enrichment analysis of signal pathways on the expression change folds of genes by using a gene set enrichment analysis method to obtain signal pathways related to each disease-related cell subpopulation and significant analysis results of the signal pathways.
[0179] In this embodiment, low-depth single-cell transcriptome data of disease samples and low-depth single-cell transcriptome data of normal control samples are obtained, and cell clustering and cell subpopulation identification operations are performed on the low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples to obtain cell subpopulation identification results. For the same cell subpopulation, the number of cells in the cell subpopulation of the disease samples and the number of cells in the cell subpopulation of the normal control samples in the cell subpopulation identification results are compared to screen out disease-related cell subpopulations with abnormal changes in the number of cells in the disease samples. The gene expression data in the disease-related cell subpopulations in the low-depth single-cell transcriptome data of the disease samples is summed, and the gene expression data in the disease-related cell subpopulations in the low-depth single-cell transcriptome data of the normal control samples is summed to obtain aggregated data. Differential gene analysis and signal pathway enrichment analysis are performed on the genes of the disease samples and the genes of the normal control samples in the aggregated data to obtain significant analysis results of signal pathways in the disease samples. Signal pathways that meet the requirements of dysregulated signal pathways in the significant analysis results are screened out as dysregulated signal pathways of the disease-related cell subpopulations. Based on the correspondence between the disease-related cell subpopulations and cell models of the action of traditional Chinese medicines, a cell model for prescription efficacy prediction is determined, and the dysregulated signal pathways of the disease-related cell subpopulations are used as dysregulated signal pathways of the cell model for prescription efficacy prediction. The entire process is automatically implemented. The cell model and the dysregulated signal pathways for prescription efficacy prediction are determined based on single-cell transcriptome data analysis, which improves the accuracy of the determination of the cell model and the dysregulated signal pathways for prescription efficacy prediction, and further performs prescription efficacy prediction based on a cross-cell line multi-target synergistic intervention strategy. In addition, the low-depth single-cell transcriptome data in this application has a low data depth. Summation operation is performed on the low-depth data and subsequent gene expression data, which can accurately determine the disease-related cell subpopulations and the dysregulated signal pathways under low data depth, has low cost, and can realize fine disease target screening.
[0180] It should be noted that the working processes of each module and submodule in this embodiment are described above, and will not be repeated here.
[0181] On the basis of the above-mentioned embodiment of the prescription efficacy prediction method, another embodiment of the present application provides a prescription efficacy prediction device, which refers to Figure 7 may include:
[0182] The information acquisition module 21 is configured to acquire the cell model for prescription efficacy prediction and the disorder signal pathway of the cell model for prescription efficacy prediction obtained by the above-mentioned determination method of the signal pathway of the cell subpopulation and the cell model.
[0183] The prescription efficacy prediction module 22 is configured to quantitatively predict the efficacy of the candidate prescription by calculating the reversal rate and the efficacy score of the candidate prescription based on the signal pathway result of the cell model of the Chinese medicine and the disorder signal pathway of the cell model for prescription efficacy prediction.
[0184] In an implementation manner, the prescription efficacy prediction module 22 includes:
[0185] The prescription acquisition submodule is configured to acquire the candidate prescription.
[0186] The first score determination submodule is configured to acquire the signal pathway result of the cell model of the single Chinese medicine in the candidate prescription by using the single Chinese medicine to process the cell model for prescription efficacy prediction, and extract the standardized enrichment score of the single Chinese medicine in the disorder signal pathway of the cell model for prescription efficacy prediction by using the disorder signal pathway of the cell model for prescription efficacy prediction.
[0187] The second score determination submodule is configured to sum the standardized enrichment scores of each Chinese medicine in the candidate prescription in the disorder signal pathway of the cell model for prescription efficacy prediction to obtain the standardized enrichment score of the candidate prescription.
[0188] The efficacy prediction submodule is configured to quantitatively predict the efficacy of the candidate prescription by calculating the reversal rate and the efficacy score of the candidate prescription based on the standardized enrichment score of the candidate prescription.
[0189] In the embodiment, the precise identification of the key signal pathway of the disease can be realized at the single cell scale. Specifically, the disease-related cell subpopulation is subjected to differential gene analysis and pathway enrichment analysis, the disorder signal pathway driving the disease is identified, and the prescription efficacy prediction is performed by using the multi-cell multi-target collaborative intervention strategy across cell lines.
[0190] It should be noted that the working processes of each module and submodule in the embodiment are described above with reference to the corresponding description in the above-mentioned embodiment, and will not be described herein.
[0191] The electronic device provided in the embodiment of the present application includes at least one processor and a memory connected with the processor, wherein:
[0192] The memory is configured to store the computer program.
[0193] The processor is configured to execute the computer program to enable the electronic device to implement the method for determining a cell subpopulation and a signal pathway of a cell model, or implement the method for predicting the efficacy of a prescription.
[0194] Reference Figure 8 As shown in FIG. 1, a structural schematic diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present application is shown. The electronic device in the embodiments of the present application can include, but is not limited to, a fixed terminal such as a mobile phone, a notebook computer, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a desktop computer, and the like. Figure 8 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0195] As shown in FIG. 1, a structural schematic diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present application is shown. The electronic device in the embodiments of the present application can include, but is not limited to, a fixed terminal such as a mobile phone, a notebook computer, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a desktop computer, and the like. Figure 8 As shown in FIG. 1, the electronic device can include a processing device (for example, a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. In the state that the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0196] Generally, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 608 including, for example, a memory card, a hard disk, and the like; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 The electronic device with various devices is shown, but it should be understood that it is not required to implement or have all the devices shown. More or fewer devices can be alternatively implemented or provided.
[0197] The embodiments of the present application also provide a computer program product including computer readable instructions, which, when executed on an electronic device, enable the electronic device to implement any one of the methods for determining a cell subpopulation and a signal pathway of a cell model or the method for predicting the efficacy of a prescription provided by the embodiments of the present application.
[0198] The embodiment of the present application further provides a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement any of the cell subpopulation and cell model signal pathway determination method or formula drug efficacy prediction method provided by the embodiment of the present application.
[0199] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, and units described as separate components can or can not be physically separated, and components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiment provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and necessary general hardware, and of course can also be realized by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily realized by corresponding hardware, and specific hardware structures for realizing the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, training device, or network device, etc.) execute the method described in each embodiment of the present application.
[0201] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, it can be realized in the form of a computer program product in whole or in part.
[0202] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. A method for determining signaling pathways in cell subpopulations and cell models, characterized in that, include: Obtain low-depth single-cell transcriptome data from disease samples and normal control samples. Cell clustering and cell subpopulation identification operations were performed on the low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples to obtain cell subpopulation identification results. For the same cell subpopulation, the cell count of the cell subpopulation in the disease sample in the cell subpopulation identification results is compared with that of the cell subpopulation in the normal control sample to screen out disease-related cell subpopulations with abnormal changes in cell count in the disease sample. The gene expression data of disease-related cell subpopulations in the low-depth single-cell transcriptome data of disease samples and the gene expression data of disease-related cell subpopulations in the low-depth single-cell transcriptome data of normal control samples were summed to obtain summary data. Differential gene analysis and signaling pathway enrichment analysis were performed on the genes of disease samples and normal control samples in the summary data to obtain the significance analysis results of signaling pathways in disease samples. Signal pathways whose significance analysis results meet the requirements for dysregulated signaling pathways are selected and used as dysregulated signaling pathways for the disease-related cell subpopulations. Based on the correspondence between the disease-related cell subpopulations and the cell models of the effects of traditional Chinese medicine, a cell model for predicting the efficacy of the prescription is determined, and the dysregulated signaling pathways of the disease-related cell subpopulations are used as the dysregulated signaling pathways of the cell model for predicting the efficacy of the prescription.
2. The method for determining cell subpopulations and signaling pathways in cell models according to claim 1, characterized in that, Cell clustering and subpopulation identification operations were performed on the low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples to obtain cell subpopulation identification results, including: The low-depth single-cell transcriptome data of the disease samples and the low-depth single-cell transcriptome data of the normal control samples were standardized, normalized, reduced in dimensionality, clustered, and annotated with cell type to obtain cell subpopulation identification results.
3. The method for determining cell subpopulations and signaling pathways in cell models according to claim 1, characterized in that, For the same cell subpopulation, the cell counts of the cell subpopulation in the disease samples and the cell subpopulation in the normal control samples are compared in the cell subpopulation identification results to screen out disease-related cell subpopulations with abnormal changes in cell counts in the disease samples, including: The cell counts of each cell subpopulation in the disease samples and the cell counts of each cell subpopulation in the normal control samples were statistically analyzed. Calculate the percentage change in the number of cells in the same cell subpopulation of the disease sample compared to that of the normal control sample. Cell subpopulations with a change in number greater than a first threshold or less than a second threshold were selected and identified as disease-related cell subpopulations.
4. The method for determining cell subpopulations and signaling pathways in cell models according to claim 1, characterized in that, Gene expression data from disease-related cell subpopulations in low-depth single-cell transcriptome data of disease samples and gene expression data from disease-related cell subpopulations in low-depth single-cell transcriptome data of normal control samples were summed to obtain aggregated data. Differential gene analysis and signaling pathway enrichment analysis were then performed on the genes from disease samples and normal control samples in the aggregated data to obtain the significance analysis results of signaling pathways in the disease samples, including: The pseudo-batch analysis method was used to sum the gene expression data of all cells in the same disease-related cell subpopulation in the low-depth single-cell transcriptome data of disease samples, and to sum the gene expression data of all cells in the same disease-related cell subpopulation in the low-depth single-cell transcriptome data of normal control samples, so as to obtain the summary data. The aggregated data were processed using differential gene analysis to obtain the fold change in gene expression. Gene set enrichment analysis was used to enrich signaling pathways based on fold changes in gene expression, resulting in signaling pathways associated with each disease-related cell subpopulation, as well as the significance analysis results of these signaling pathways.
5. A method for predicting the efficacy of a prescription, characterized in that, include: Obtain a cell model for predicting the efficacy of a prescription drug, obtained by the method for determining the signaling pathways of cell subpopulations and cell models as described in any one of claims 1-4, and the dysregulated signaling pathways of the cell model for predicting the efficacy of the prescription drug. Based on the signaling pathway results of the cell model of the effects of traditional Chinese medicine and the dysregulated signaling pathways of the cell model used for predicting the efficacy of prescriptions, the efficacy of candidate prescriptions is quantitatively predicted by calculating the reversal rate and efficacy score.
6. The method for predicting the efficacy of a prescription according to claim 5, characterized in that, Based on the signaling pathway results of the cell model described in the traditional Chinese medicine action theory and the dysregulated signaling pathways of the cell model used for predicting the efficacy of prescriptions, the efficacy of candidate prescriptions is quantitatively predicted by calculating the reversal rate and efficacy fraction, including: Obtain candidate prescriptions; The cell model used for predicting the efficacy of the single Chinese herbal medicine in the candidate prescription is processed to obtain the signal pathway results of the single Chinese herbal medicine acting on the cell model. The standardized enrichment score of the single Chinese herbal medicine in the cell model used for predicting the efficacy of the prescription is extracted using the disordered signal pathways of the cell model used for predicting the efficacy of the prescription. The standardized enrichment scores of each herb in the candidate prescription on the dysregulated signaling pathways used in the cell model for predicting prescription efficacy are summed to obtain the standardized enrichment score of the candidate prescription. Based on the standardized enrichment score of the candidate prescriptions, the efficacy is quantitatively predicted by calculating the reversal rate and efficacy score of the candidate prescriptions.
7. A device for determining signaling pathways in cell subpopulations and cell models, characterized in that, include: The data acquisition module is used to acquire low-depth single-cell transcriptome data of disease samples and low-depth single-cell transcriptome data of normal control samples. The identification module is used to perform cell clustering and cell subpopulation identification operations on the low-depth single-cell transcriptome data of the disease sample and the low-depth single-cell transcriptome data of the normal control sample to obtain cell subpopulation identification results. The comparison module is used to compare the cell count of the cell subpopulation in the disease sample and the cell subpopulation in the normal control sample in the cell subpopulation identification results for the same cell subpopulation, so as to screen out disease-related cell subpopulations with abnormal changes in cell count in the disease sample. The analysis module is used to sum the gene expression data of disease-related cell subpopulations in the low-depth single-cell transcriptome data of disease samples and the gene expression data of disease-related cell subpopulations in the low-depth single-cell transcriptome data of normal control samples to obtain summary data. Differential gene analysis and signaling pathway enrichment analysis are performed on the genes of disease samples and normal control samples in the summary data to obtain the significance analysis results of signaling pathways in disease samples. The screening module is used to screen out signaling pathways whose significance analysis results meet the requirements of the dysregulated signaling pathways, and to use them as the dysregulated signaling pathways of the disease-related cell subpopulations. The determination module is used to determine the cell model for predicting the efficacy of traditional Chinese medicine based on the correspondence between the disease-related cell subpopulation and the cell model of the effect of traditional Chinese medicine, and to use the disordered signaling pathways of the disease-related cell subpopulation as the disordered signaling pathways of the cell model for predicting the efficacy of traditional Chinese medicine.
8. A device for predicting the efficacy of a prescription, characterized in that, include: The information acquisition module is used to acquire the cell model for predicting the efficacy of a prescription obtained by the method for determining the signal pathways of cell subpopulations and cell models as described in any one of claims 1-4, and the disordered signal pathways of the cell model for predicting the efficacy of the prescription. The formula efficacy prediction module is used to quantitatively predict the efficacy of a formula based on the signaling pathway results of the cell model of the effects of traditional Chinese medicine and the disordered signaling pathways of the cell model for formula efficacy prediction, by calculating the reversal rate and efficacy score of the candidate formula.
9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the method for determining the signaling pathways of cell subpopulations and cell models as described in any one of claims 1 to 4, or the method for predicting the efficacy of prescriptions as described in any one of claims 5 to 6.
10. A computer storage medium, characterized in that, The computer storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the method for determining the signaling pathways of cell subpopulations and cell models as described in any one of claims 1 to 4, or to implement the method for predicting the efficacy of prescriptions as described in any one of claims 5 to 6.