Tumor target spot and drug quantification system and method based on multi-gene expression profile
By establishing a tumor target and drug quantification system based on multigene expression profile and combining machine learning technology, the problem of lack of personalized drug selection in tumor targeted therapy is solved, and the scientific basis for personalized drug selection and treatment plans is realized, which improves the treatment effect and patient quality of life.
Patent Information
- Application Number
- CN202510495691.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, tumor targeted treatment lacks personalized scientific guidance, which makes drug selection vary from person to person and makes it difficult to improve the treatment effect.
Establish a tumor target and drug quantification system based on multigene expression profile, and use a comprehensive database module, assay and standardized processing module, target quantification module, gene fusion information identification module, gene mutation information identification module and drug quantification module, combined with machine learning technology, multi-dimensional analysis is carried out to identify individual tumor characteristics and drug response-related molecular targets, and provide personalized drug selection.
It improves the targeted treatment of tumors, reduces side effects, improves patients' quality of life, prolongs survival, and provides scientifically-based treatment plans and drug choices.
Smart Images

Figure CN120431992A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and in particular to a tumor target and drug quantification system and method based on multi-gene expression profiling. Background Art
[0002] In recent years, with the continuous advancement of cancer research, people have gradually realized that tumor heterogeneity is not only reflected between different cancer types, but also between different patients with the same cancer. Tumor development is the result of the combined effects of genetics and the environment. According to the central dogma of "gene->transcription->protein", transcription is not only the key bridge between genes and proteins, but also a comprehensive manifestation of the interaction between genetic information and environmental information in tumor patients.
[0003] In addition, targeted therapy and immunotherapy have become important treatments for various tumors, including lung cancer. However, the lack of personalized scientific guidance for drug selection is a major challenge currently faced. Although some targeted drugs have shown significant efficacy, due to the complexity of the patient's gene expression profile, clinical effects often vary from person to person. Therefore, in the field of tumor treatment, how to accurately select appropriate targets and drugs based on the patient's molecular characteristics has become the key to improving treatment efficacy. As more and more tumor targeted drugs and immunotherapy drugs are approved for marketing, how to scientifically select personalized drugs from a large number of drugs has become a core issue in personalized tumor treatment. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a tumor target and drug quantification system and method based on multi-gene expression profiles.
[0005] The technical solution is as follows: a tumor target and drug quantification system based on multi-gene expression profiles, characterized in that the system includes:
[0006] A comprehensive database module is used to establish a comprehensive tumor information database, which includes tumor transcriptome data, tumor clinical information data, and anti-tumor drug and target benchmark data;
[0007] The measurement and standardization processing module measures the transcriptome of the tissue sample to obtain the original expression counts of the tissue sample, and standardizes the obtained original expression counts to obtain stable, reliable, and comparable gene expression levels;
[0008] Target quantification module quantifies and identifies the target of each gene and obtains the quantitative value of each target;
[0009] Gene fusion information identification module, which uses the transcriptome of tissue samples to identify the gene fusion information of each target and output the results;
[0010] Gene mutation information identification module, which uses the high-depth transcriptome of tissue samples to identify gene mutation information of each target and output the results;
[0011] The drug quantification module uses the results of the target quantification module, the gene fusion information identification module and the gene mutation information identification module to conduct a comprehensive quantitative analysis of the drugs, and outputs the drug quantification results in descending order.
[0012] The tumor target and drug quantification method based on multi-gene expression profiling mainly includes the above-mentioned tumor target and drug quantification system based on multi-gene expression profiling, and the steps are as follows:
[0013] S1: Comprehensive database module, which collects tumor benchmark data and establishes a comprehensive tumor information database, which includes tumor transcriptome data, tumor clinical information dataset, anti-tumor drug and target benchmark data;
[0014] The tumor types include 15 solid tumors: lung cancer (including non-small cell lung cancer and small cell lung cancer), breast cancer, liver cancer, colorectal cancer, gastric cancer, pancreatic cancer, endometrial cancer, prostate cancer, ovarian cancer, bladder cancer, kidney cancer (including clear cell renal carcinoma and papillary renal cell carcinoma), esophageal cancer, brain tumor, head and neck squamous cell carcinoma, and malignant sarcoma; 1 non-solid tumor: acute leukemia;
[0015] S2: Measurement and standardization processing module, which measures the transcriptome of the tissue sample to obtain the original expression counts of the sample, and standardizes the original expression counts to obtain stable, reliable, and comparable gene expression levels;
[0016] S3: Target quantification module, which quantifies and identifies targets for each gene and calculates the comprehensive score corresponding to each target through the target comprehensive quantification function;
[0017] Factors considered in target quantification and identification include: a) the expression difference (log2FC) of each target in the tissue sample and the corresponding normal tissue in the tumor comprehensive information database; b) the expression level ranking percentage of each target among the same tumor patients worldwide (perc_expr%); c) the absolute value of the expression level (log2CPM and TPM); d) the overall expression positive level (Positive); e) the relationship between the target and overall survival (Harzard ratio, HR);
[0018] S4: Gene fusion information identification module, which uses the transcriptome of tissue samples to identify the gene fusion information of each target;
[0019] S5: Gene mutation information identification module, which uses the high-depth transcriptome of tissue samples to identify gene mutation information of each target;
[0020] S6: Drug quantification module, which comprehensively quantifies the potential therapeutic effect of each drug and calculates the comprehensive efficacy value of the drug through the drug comprehensive quantification function;
[0021] Factors considered in drug quantification include: a. mutation information of each target corresponding to the drug; b. gene fusion information of each target corresponding to the drug; c. classification information of each target, whether it is a core target; d. comprehensive scoring score corresponding to the target; e. number of targets corresponding to the drug; f. degree of indication matching of the drug with the current national or world standard; g. number of generations of the drug targeting the corresponding target.
[0022] Preferably, in step S2, total RNA is extracted from the tissue sample and processed to obtain raw expression counts. The obtained raw expression counts are then used to calculate the standardized counts per million (CPM) using the M-value trimmed mean (TMM) method in the edgeR package. The standardized count data are used for subsequent differential expression analysis.
[0023] Preferably, the original expression counts of the sample and the counts of the corresponding tumor are standardized together to obtain a standardized CPM value, thereby normalizing the data and eliminating errors in different platforms and time points.
[0024] Preferably, the formula of the target comprehensive quantification function in step S3 is:
[0025]
[0026] in
[0027]
[0028] Among them, HR is the relationship between the target and overall survival in the tissue sample, that is, the corresponding risk value, which is calculated using the survival package in the R package;
[0029] log2CPM s CPM value of tissue samples after log2 processing;
[0030] log2CPM T is the log2-processed CPM value of all tumor samples in the comprehensive tumor information database that have the same tumor as sample S;
[0031] nt is the number of tumor samples in the comprehensive tumor information database;
[0032] log2CPM Nis the CPM value of all normal tissues corresponding to tissue samples in the tumor comprehensive information database after log2 processing;
[0033] nc is the number of normal samples in the comprehensive tumor information database;
[0034] TPM S TPM is the transcript per million (TPM) value of gene expression in tissue samples;
[0035] is the tissue sample expression ranking function, which is used to calculate the expression ranking percentage in the tissue sample;
[0036] It is the target positive level function, which is used to calculate whether the target is positively expressed.
[0037] Preferably: if the Indicates that the gene is negative, represented by (-);
[0038] If the It indicates that the gene is weakly positive, represented by (+);
[0039] If the It indicates that the gene is positive, which is indicated by (++);
[0040] If the It indicates that the gene is highly positive, which is represented by (+++);
[0041] Wherein a, b, c, d, e, f, g, h, k, p1, p2, and p3 in the target comprehensive quantification function are all set parameters;
[0042] Wherein a=2.0, b=6.0, c=1.0, d=1.0, e=2.0, f=0.15, g=0.25, h=0.5, k=0.26, p1=1.0, p2=2.0, p3=4.0.
[0043] Preferably, in step S4, arriba software is used to identify gene fusion events of each target using the transcriptome of the tissue sample, and the gene fusion results are obtained;
[0044] In the S5 step, GATK2 and bcltools are used to simultaneously identify gene mutations of each target point in the high-depth transcriptome of the tissue sample and obtain gene mutation results.
[0045] Preferably, the formula of the drug comprehensive quantification function in step S6 is:
[0046]
[0047]
[0048] Where n is the number of core targets;
[0049] k is the number of non-core targets;
[0050] Indicates the comprehensive score of the core target corresponding to the drug;
[0051] Indicates the comprehensive score of the non-core targets corresponding to the drug;
[0052] Indicates the mutation status of the target;
[0053] F i core Indicates the fusion status of the target;
[0054] β is the weight parameter of the core target;
[0055] γ is the weight parameter of non-core targets;
[0056] ε is the harmonic parameter of multiple targets, -1≤ε≤0;
[0057] α is the degree of matching between the drug and the cancer indication;
[0058] θ is the drug algebra.
[0059] Preferably: Indicates the mutation status of the target. If the drug used has nothing to do with the mutation status of the target, then If the drug targets a type of mutation in the target site and the target site does not have this type of mutation, then If there is such a mutation, then
[0060] F i core Indicates the fusion status of the target. If the drug used has nothing to do with the fusion status of the target gene, then F i core =0, if the drug targets a type of mutation in the target site, and the target site does not have this type of mutation, then F i core =-20, if there is such a mutation, then F i core =80;
[0061] The weight parameter of the core target is β = 1;
[0062] The weight parameter γ for non-core targets is 0.25;
[0063] Multi-target reconciliation parameter ε = -0.4;
[0064] α is the degree of matching between the drug and the tumor indication. If the drug is a targeted drug and is just suitable for the tumor according to national or world standards, α = 7; if the drug is a targeted drug but the national or world standards are not yet suitable for the tumor, α = 2; for other non-targeted drugs, α = 1;
[0065] is the drug generation, which is determined by the drug year and the target generation, and is calculated as:
[0066] θ=(approval_year-2000) / 3+4×passage.
[0067] Preferably, the drug comprehensive quantification function calculates the pharmacodynamic value of each drug, arranges the pharmacodynamic values in descending order, and lists the combination of solutions with the maximum pharmacodynamic value.
[0068] Compared with the existing technology, the beneficial effects of the present invention are as follows: the present invention establishes and applies a tumor target and drug quantification system and method based on multi-gene expression profiles, which will provide innovative tools and methods for clinical practice. This method can be applied to the treatment of 16 types of tumors, and can provide a scientific basis for clinical prognosis assessment, treatment plans and drug selection for tumors, thereby improving the targeted treatment, reducing side effects, improving the quality of life of patients and prolonging survival. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0070] The present invention will be further described below with reference to the embodiments and accompanying drawings.
[0071] The tumor target and drug quantification method based on multi-gene expression profile, whose gene expression profile is an important data source reflecting the patient's molecular characteristics, can provide a basis for evaluating the molecular characteristics of the tumor. The present invention combines multi-dimensional gene expression data to not only reveal individual tumor characteristics, but also evaluate the potential effects of different drugs by integrating biomarkers related to drug response. This multi-gene expression profile analysis helps to identify molecular targets that are closely related to disease progression and drug response. On this basis, the present invention provides a new reference system for personalized tumor target and drug evaluation through the advancement of artificial intelligence and machine learning technology. Through large-scale data training and optimization algorithms, machine learning models can quickly analyze large amounts of gene expression profile data and drug-related data to mine potential drug response patterns.
[0072] like Figure 1As shown, a tumor target and drug quantification system based on multi-gene expression profiling includes:
[0073] A comprehensive database module is used to establish a comprehensive tumor information database, which includes tumor transcriptome data, tumor clinical information data, and anti-tumor drug and target benchmark data;
[0074] The measurement and standardization processing module measures the transcriptome of the tissue sample to obtain the original expression counts of the tissue sample, and standardizes the obtained original expression counts to obtain stable, reliable, and comparable gene expression levels;
[0075] Target quantification module quantifies and identifies the target of each gene and obtains the quantitative value of each target;
[0076] Gene fusion information identification module, which uses the transcriptome of tissue samples to identify the gene fusion information of each target and output the results;
[0077] Gene mutation information identification module, which uses the high-depth transcriptome of tissue samples to identify gene mutation information of each target and output the results;
[0078] The drug quantification module uses the results of the target quantification module, the gene fusion information identification module and the gene mutation information identification module to conduct a comprehensive quantitative analysis of the drugs, and outputs the drug quantification results in descending order.
[0079] The method for quantifying tumor targets and drugs based on multi-gene expression profiling comprises the following steps:
[0080] S1 comprehensive database module
[0081] 1.1 Acquisition of tumor transcriptome data
[0082] Through authoritative databases such as The Cancer Genome Atlas (TCGA) and the International Cancer Genome Consortium (ICGC), we obtain transcriptome data from cancer patients and their adjacent normal tissues worldwide. This data is derived from RNA sequencing (RNA-Seq) analysis and covers a large amount of gene expression information from cancer patients. The TCGA database provides a wealth of transcriptome data. To ensure data quality and consistency, we will perform quality control on the downloaded data and handle missing values to ensure reliability in subsequent analyses.
[0083] The tumor types include 15 solid tumors: lung cancer (including non-small cell lung cancer and small cell lung cancer), breast cancer, liver cancer, colorectal cancer, gastric cancer, pancreatic cancer, endometrial cancer, prostate cancer, ovarian cancer, bladder cancer, kidney cancer (including clear cell renal carcinoma and papillary renal cell carcinoma), esophageal cancer, brain tumors, head and neck squamous cell carcinoma, and malignant sarcoma; and 1 type of non-solid tumor: acute leukemia.
[0084] 1.2 Acquisition of tumor clinical information data
[0085] Integrate clinical data from public databases such as TCGA and ICGC. This data includes basic patient information (such as age, gender, and race), pathological stage, degree of differentiation, treatment options, and prognosis. This data will be used to build personalized target quantification models, provide targeted treatment plans for different patients, and help verify the clinical utility of biomarkers.
[0086] 1.3 Acquisition of Benchmark Data on Anti-tumor Drugs and Targets
[0087] To obtain target information for tumor-related anti-tumor drugs, we collected and integrated data on all approved anti-tumor drugs and their targets through DrugBank, the FDA (U.S. Food and Drug Administration) database, and the list of anti-tumor drugs approved by the China National Medical Products Administration (NMPA). Drug data includes information such as targeted drugs for tumor indications, immunotherapy drugs and their mechanisms of action, target names, gene symbols, and drug classifications. In addition, databases such as MyCancerGenome were referenced to ensure the accuracy and timeliness of the selected drug and target data. For drug targets that have not yet been identified, drug target prediction models are used to further explore potential targets to supplement the benchmark dataset.
[0088] 1.4 Data integration and standardization
[0089] After integrating the above multi-source data, data standardization and preprocessing are performed, including unifying gene symbols, removing redundant data, processing missing values, and correcting batch effects to ensure comparability and compatibility between data and establish a database.
[0090] S2: Measurement and standardization processing module
[0091] 2.1 Transcriptome analysis of sample tissues
[0092] Total RNA is extracted from fresh or well-preserved tissue samples, and its quality and integrity are rigorously tested to ensure that the samples are suitable for downstream analysis. mRNA in the transcriptome is extracted and enriched to obtain a higher signal-to-noise ratio. The enriched mRNA is reverse transcribed into cDNA, and through the library construction step, the cDNA fragments are sheared, adapters are added, and PCR amplified to generate a library suitable for reading by the Illumina sequencer. After the library construction is completed, the sample is loaded onto a sequencer such as the NovaSeq 6000, so that the overall sequencing quantity is >6G. Fastp is used to perform quality control on the original reads, and the high-quality reads are aligned to the reference genome (such as GRCH38). Transcriptome analysis tools (such as STAR) are used to align the reads to obtain the original expression counts of the patient samples.
[0093] 2.2 Obtaining normalized gene expression
[0094] To obtain stable, reliable, and comparable gene expression levels, we calculated normalized counts per million (CPM) using the trimmed mean M-value (TMM) method in the edgeR package for subsequent analysis. This method combines the raw expression counts of the sample with the corresponding tumor counts to calculate a normalized CPM value to eliminate errors across different platforms and time points.
[0095] S3: Target Quantification Module
[0096] The systematic quantification and identification of targets should be considered from multiple perspectives, including the following factors:
[0097] a. The expression difference (log2FC) of each target in the sample and the corresponding normal tissue in the tumor comprehensive information database;
[0098] b. The expression level ranking percentage of each target among the same tumor patients in the world (perc_expr%);
[0099] c, absolute values of expression levels (log2CPM and TPM);
[0100] d. Overall expression positive level (Positive);
[0101] e. The relationship between this target and overall survival (Harzard ratio, HR).
[0102] For each gene, the target must be comprehensively quantified through the target comprehensive quantification function. The formula of the target comprehensive quantification function is:
[0103]
[0104] in
[0105]
[0106] Among them, HR is the relationship between the target and overall survival in the tissue sample, that is, the corresponding risk value, which is calculated using the survival package in the R package;
[0107] log2CPM s CPM value of tissue samples after log2 processing;
[0108] log2CPM T is the log2-processed CPM value of all tumor samples with the same tumor as the tissue sample in the comprehensive tumor information database;
[0109] nt is the number of tumor samples in the construction of the comprehensive tumor information database;
[0110] log2CPM N is the log2-processed CPM value of all normal tissues corresponding to tissue samples in the tumor comprehensive information database;
[0111] nc is the number of normal samples in the comprehensive tumor information database;
[0112] TPM S TPM is the transcript per million (TPM) value of gene expression in tissue samples;
[0113] is the tissue sample expression ranking function, which is used to calculate the expression ranking percentage in the tissue sample;
[0114] is the target positive level function, which is used to calculate whether the target is positively expressed. Indicates that the gene is negative, represented by (-). Indicates that the gene is weakly positive, represented by (+). Indicates that the gene is positive, represented by (++). It indicates that the gene is highly positive, which is represented by (+++).
[0115] Among them, a, b, c, d, e, f, g, h, k, p1, p2, and p3 are parameters that need to be set, and they need to be optimized based on experience and large-scale simulation calculations. Here, through the experience and simulation calculations of biological and medical experts, the parameters are determined to be: a=2.0, b=6.0, c=1.0, d=1.0, e=2.0, f=0.15, g=0.25, h=0.5, k=0.26, p1=1.0, p2=2.0, and p3=4.0.
[0116] Based on the formula of the above target comprehensive quantification function, the comprehensive score TargetScore corresponding to each target can be obtained. Then, this comprehensive score can be used to sort all targets from large to small.
[0117] S4: Gene fusion information identification module
[0118] Since many targeted drugs target gene fusions, transcriptomes are used to identify the gene fusion information of each target. The identification method is to use arriba software to use the transcriptome of tissue samples to identify all gene fusion events and obtain the gene fusion results.
[0119] S5: Gene mutation information identification module
[0120] Similarly, because many targeted drugs target certain gene mutations, the transcriptome of tissue samples is also used to identify mutation information for each target. To ensure the accuracy of the identification results, two different methods, GATK2 and bcltools, are used to identify the target gene mutations. Those identified by both methods are considered to have the corresponding gene mutation.
[0121] S6: Drug Quantification Module
[0122] By integrating big data from medical biology and cheminformatics, and incorporating patient risk information, comprehensive evaluation information on small molecule drug targets, gene mutation events, gene fusion events, and other information into machine learning and artificial intelligence analysis models, we conduct a comprehensive efficacy evaluation and ranking of all targeted drugs (over 200) to scientifically evaluate personalized medicines and provide a scientific basis for treatment plans and drug selection. Based on the scores of all current targets, we can comprehensively quantify the potential efficacy of each drug. The comprehensive quantitative score of each drug is calculated using the drug comprehensive quantification function.
[0123] The comprehensive quantification will take into account the following information:
[0124] a. Mutation information of each target corresponding to the drug;
[0125] b. Gene fusion information of each drug target;
[0126] c. Classification information of each target (core target or non-core target);
[0127] d. Comprehensive score corresponding to the target;
[0128] e. The number of targets corresponding to the drug;
[0129] f. The degree to which the drug's indication matches the current national or world standards;
[0130] g. The number of generations of the drug targeting the corresponding target.
[0131] The formula of drug comprehensive quantification function is:
[0132]
[0133] Where n is the number of core targets;
[0134] k is the number of non-core targets;
[0135] Indicates the comprehensive score of the core target corresponding to the drug;
[0136] Indicates the comprehensive score of the non-core targets corresponding to the drug;
[0137] Indicates the mutation status of the target, Indicates the mutation status of the target. If the drug used has nothing to do with the mutation status of the target, then If the drug targets a type of mutation in the target site and the target site does not have this type of mutation, then If there is such a mutation, then
[0138] F i core Indicates the fusion status of the target. If the drug used has nothing to do with the fusion status of the target gene, then F i core =0, if the drug targets a type of mutation in the target site, and the target site does not have this type of mutation, then F i core =-20, if there is such a mutation, then F i core =80;
[0139] β is the weight parameter of the core target, β = 1;
[0140] γ is the weight parameter of non-core targets, γ = 0.25;
[0141] ε is the harmonic parameter of multiple targets, -1≤ε≤0, and this formula uses ε=-0.4;
[0142] α is the degree of matching between the drug and the tumor indication. If the drug is a targeted drug and is just suitable for the tumor according to national or world standards, α = 7; if the drug is a targeted drug but the national or world standards are not yet suitable for the tumor, α = 2; for other non-targeted drugs, α = 1;
[0143] θ is the drug generation, which is determined by the drug year and target generation, and the calculation formula is:
[0144] θ=(approval_year-2000) / 3+4×passage.
[0145] The drug comprehensive quantification function can be used to obtain the corresponding pharmacodynamic value of each drug, and its pharmacodynamic value can be arranged in descending order, so as to realize personalized drug recommendation for patients according to its pharmacodynamic value.
[0146] The tumor target and drug quantification system and method based on multi-gene expression profiling can simultaneously integrate drug informatics to comprehensively evaluate personalized drug targets for 16 types of tumors by utilizing artificial intelligence and biomedical big data. It can provide a scientific basis for clinical prognosis assessment, treatment plans and drug selection for tumors, and list the combination of solutions with the maximum pharmacodynamic value. The use of this system and method can improve the efficiency of optimized drug screening.
[0147] Finally, it should be noted that the above description is only a preferred embodiment of the present invention. Under the guidance of the present invention, ordinary technicians in this field can make various similar expressions without violating the purpose and claims of the present invention. Such changes fall within the scope of protection of the present invention.
Claims
1. A tumor target and drug quantification system based on multi-gene expression profiling, characterized by: The system comprises: A comprehensive database module is used to establish a comprehensive tumor information database, which includes tumor transcriptome data, tumor clinical information data, and anti-tumor drug and target benchmark data; The measurement and standardization processing module measures the transcriptome of the tissue sample to obtain the original expression counts of the tissue sample, and standardizes the obtained original expression counts to obtain stable, reliable, and comparable gene expression levels; Target quantification module quantifies and identifies the target of each gene and obtains the quantitative value of each target; Gene fusion information identification module, which uses the transcriptome of tissue samples to identify the gene fusion information of each target and output the results; Gene mutation information identification module, which uses the high-depth transcriptome of tissue samples to identify gene mutation information of each target and output the results; The drug quantification module uses the results of the target quantification module, the gene fusion information identification module and the gene mutation information identification module to conduct a comprehensive quantitative analysis of the drugs, and outputs the drug quantification results in descending order.
2. A method for quantifying tumor targets and drugs based on multi-gene expression profiling, characterized by: The method comprises the tumor target and drug quantification system based on multi-gene expression profile as claimed in claim 1, wherein the steps are as follows: S1: Comprehensive database module, which collects tumor benchmark data and establishes a comprehensive tumor information database, which includes tumor transcriptome data, tumor clinical information dataset, anti-tumor drug and target benchmark data; S2: Measurement and standardization processing module, which measures the transcriptome of the tissue sample to obtain the original expression counts of the sample, and standardizes the original expression counts to obtain stable, reliable, and comparable gene expression levels; S3: Target quantification module, which quantifies and identifies targets for each gene and calculates the comprehensive score corresponding to each target through the target comprehensive quantification function; Factors considered in target quantification and identification include: a) the expression difference (log2FC) of each target in the tissue sample and the corresponding normal tissue in the tumor comprehensive information database; b) the expression level ranking percentage of each target among the same tumor patients worldwide (perc_expr%); c) the absolute value of the expression level (log2CPM and TPM); d) the overall expression positive level (Positive); e) the relationship between the target and overall survival (Harzard ratio, HR); S4: Gene fusion information identification module, which uses the transcriptome of tissue samples to identify the gene fusion information of each target; S5: Gene mutation information identification module, which uses the high-depth transcriptome of tissue samples to identify gene mutation information of each target; S6: Drug quantification module, which comprehensively quantifies the potential therapeutic effect of each drug and calculates the comprehensive efficacy value of the drug through the drug comprehensive quantification function; Factors to be considered in comprehensive drug quantification include: a. mutation information of each target corresponding to the drug; b. gene fusion information of each target corresponding to the drug; c. classification information of each target, whether it is a core target; d. comprehensive scoring score corresponding to the target; e. number of targets corresponding to the drug; f. degree of indication matching of the drug with the current national or world standards; g. number of generations of the drug targeting the corresponding target.
3. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 2, characterized in that: In step S2, total RNA is extracted from the tissue sample and processed to obtain raw expression counts. The obtained raw expression counts are calculated using the M-value trimmed mean TMM method in the edgeR package to calculate the standardized counts per million (CPM). The standardized count data are used for subsequent differential expression analysis.
4. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 3, characterized in that: The original expression counts of the tissue sample and the corresponding tumor counts were standardized together to obtain the standardized CPM value to achieve data normalization.
5. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 2, characterized in that: The formula of the target comprehensive quantification function in step S3 is: in Among them, HR is the relationship between the target and overall survival in the tissue sample, that is, the corresponding risk value, which is calculated using the survival package in the R package; log2CPM s CPM value of tissue samples after log2 processing; log2CPM T is the log2-processed CPM value of all tumor samples with the same tumor as the tissue sample in the comprehensive tumor information database; nt is the number of tumor samples in the comprehensive tumor information database; log2CPM N is the log2-processed CPM value of all normal tissues corresponding to tissue samples in the tumor comprehensive information database; nc is the number of normal samples in the comprehensive tumor information database; TPM S TPM is the transcript per million (TPM) value of gene expression in tissue samples; is the tissue sample expression ranking function, which is used to calculate the expression ranking percentage in the tissue sample; It is the target positive level function, which is used to calculate whether the target is positively expressed.
6. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 5, characterized in that: If the Indicates that the gene is negative, represented by (-); If the It indicates that the gene is weakly positive, represented by (+); If the It indicates that the gene is positive, which is indicated by (++); If the It indicates that the gene is highly positive, which is represented by (+++); Wherein a, b, c, d, e, f, g, h, k, p1, p2, and p3 in the target comprehensive quantification function are all set parameters; Wherein a=2.0, b=6.0, c=1.0, d=1.0, e=2.0, f=0.15, g=0.25, h=0.5, k=0.26, p1=1.0, p2=2.0, p3=4.
0.
7. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 2, characterized in that: In the S4 step, arriba software is used to identify gene fusion events of each target using the transcriptome of the tissue sample and obtain the gene fusion results; In the S5 step, GATK2 and bcltools are used to simultaneously identify gene mutations of each target point in the high-depth transcriptome of the tissue sample and obtain gene mutation results.
8. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 7, characterized in that: The formula of the drug comprehensive quantification function in step S6 is: Where n is the number of core targets; k is the number of non-core targets; Indicates the comprehensive score of the core target corresponding to the drug; Indicates the comprehensive score of the non-core targets corresponding to the drug; Indicates the mutation status of the target; F i core Indicates the fusion status of the target; β is the weight parameter of the core target; γ is the weight parameter of non-core targets; ε is the harmonic parameter of multiple targets, -1≤ε≤0; α is the degree of matching between the drug and the tumor indication; θ is the drug algebra.
9. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 8, characterized in that: described Indicates the mutation status of the target. If the drug used has nothing to do with the mutation status of the target, then If the drug targets a type of mutation in the target site and the target site does not have this type of mutation, then If there is such a mutation, then F i core Indicates the fusion status of the target. If the drug used has nothing to do with the fusion status of the target gene, then F i core =0, if the drug targets a type of mutation in the target site, and the target site does not have this type of mutation, then F i core =-20, if there is such a mutation, then F i core =80; The weight parameter of the core target is β = 1; The weight parameter γ for non-core targets is 0.25; Multi-target reconciliation parameter ε = -0.4; α is the degree of matching between the drug and the tumor indication. If the drug is a targeted drug and is just suitable for the tumor according to national or world standards, α = 7; if the drug is a targeted drug but the national or world standards are not yet suitable for the tumor, α = 2; for other non-targeted drugs, α = 1; θ is the drug generation, which is determined by the drug year and the target generation, and is calculated as: θ=(approval_year-2000) / 3+4×passage.
10. The method for quantifying tumor targets and drugs based on multi-gene expression profiling according to claim 8, characterized in that: The drug comprehensive quantification function calculates the pharmacodynamic value of each drug, arranges the pharmacodynamic values in descending order, and lists the combination of solutions with the maximum pharmacodynamic value.