Mutated gene set for tumor molecular typing and application thereof

By providing a tumor molecular subtyping scheme that includes multiple gene sets, combined with the combined application of targeted therapy drugs, the problem of poor treatment efficacy for PTCL has been solved, achieving precision medicine and improved survival rates.

CN117253542BActive Publication Date: 2026-05-29RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2023-10-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Current technologies lack effective molecular subtyping schemes and targets, resulting in poor treatment outcomes for peripheral intranodal T-cell lymphoma (PTCL), low 5-year survival rates, and a lack of targeted treatment options.

Method used

This invention provides a set of mutated genes, including genes related to DNA methylation, histone modification, chromatin remodeling, TCR signaling pathway, PI3K-AKT, JAK-STAT, immune escape, p53 signaling pathway, tumor suppression, and NOTCH signaling pathway, for use in tumor molecular subtyping, in combination with targeted therapies such as 5-azacytidine, dasatinib, chidamide, apatinib, and PD-1 antibodies.

Benefits of technology

This approach enables precise molecular subtyping of PTCL, improving the targeting of treatment, significantly increasing patient survival rates and treatment outcomes, and reducing unnecessary overtreatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117253542B_ABST
    Figure CN117253542B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biological diagnosis, and particularly relates to a mutant gene set for tumor molecular typing and application thereof, wherein the mutant gene set comprises 82 mutant genes including DNA methylation modification related genes, histone modification related genes, chromatin remodeling related genes, TCR signal pathway related genes, PI3K-AKT signal pathway related genes, JAK-STAT signal pathway related genes, immune escape related genes, P53 signal pathway related genes, tumor suppressor genes and NOTCH signal pathway related genes. The mutant gene set can be used for molecular typing, targeted therapy and overall survival prediction of intranodal peripheral T-cell lymphoma. The mutant gene set is verified by clinical trials, and is suitable for all primary and relapsed PTCL patients, relapsed or refractory peripheral T-cell lymphoma umbrella study based on genomics typing and PTCL patients receiving different treatment schemes, and has a very wide application range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biodiagnostic technology, and in particular to a set of mutated genes for tumor molecular subtyping and its applications. Background Technology

[0002] Peripheral T-cell lymphomas (PTCLs) are a heterogeneous and aggressive group of malignant tumors, accounting for 10–15% of all non-Hodgkin's lymphoma (NHL) cases. PTCLs originate from mature T lymphocytes and mainly include anaplastic large cell lymphoma (ALCL), peripheral T-cell lymphoma of intranodal follicular helper T cell origin (TFH), and peripheral T-cell lymphoma-nonspecific type (PTCL-NOS, 30%). ALCL can be further divided into anaplastic lymphoma kinase (ALK) positive (ALK+) and ALK negative (ALK-). TFH lymphoma includes three subtypes: angioimmunoblastic, follicular, and NOS. PTCL-NOS mainly includes two subsets: PTCL-TBX21 and PTCL-GATA3, characterized by high expression of Th1 and Th2 cell differentiation regulators, respectively. Previous studies have shown that the heterogeneity of PTCL stems from complex molecular biological mechanisms. For example, in ALCL, frequent mutations in epigenetic genes (such as TET2, DNMT3A, and IDH2) and TCR signaling-related genes (such as RHOA) are common; while in PTCL-NOS, histone modification genes, mainly methylation and acetylation, are relatively common. The mutation characteristics of TFH-derived PTCL-NOS are highly similar to those of AITL, suggesting that different pathological subtypes of PTCL exhibit both genetic differences and similarities.

[0003] Currently, the treatment paradigm for PTCL is derived from therapies developed for non-Hodgkin's B-cell tumors, namely the traditional chemotherapy regimen CHOP (cyclophosphamide, doxorubicin, vincristine, and prednisone). Except for ALK+ALCL, which has a relatively good prognosis, the efficacy of other PTCL subtypes is poor, with a 5-year survival rate of only 25%–35%, and an even lower progression-free survival rate. At present, PTCL genomics research both domestically and internationally is limited, and there is a lack of molecular subtyping schemes and effective molecular therapeutic targets for PTCL. Therefore, molecular subtyping of PTCL based on genomic data is particularly important for promoting targeted therapy and improving patient survival rates.

[0004] Chinese patent CN112430658A discloses a detection kit and library construction method for genes related to intranodal and peripheral T-cell lymphoma. The detection kit for genes related to intranodal and peripheral T-cell lymphoma includes capture probes for 76 target genes. The library construction has high coverage, high alignment rate, good uniformity and simple operation, but it lacks clinical support for prognosis and its application is limited. Summary of the Invention

[0005] A first aspect of the present invention provides a set of mutated genes for tumor molecular subtyping, the set of mutated genes comprising the following genes:

[0006] DNA methylation-related genes: TET2, DNMT3A, IDH2, SALL3, TET1, TET3;

[0007] Histone modification-related genes: KMT2C, KMT2D, KDM6B, SETD2, KMT2A, SETBP1, SETD1B, CREBBP, EP300, NCOR2, MEF2A, TRRAP, SPEN, BCOR, YEATS2;

[0008] Chromatin remodeling-related genes: ARID1A, ARID1B, ARID2, ASXL3, SMARCA2, SMARCA4, CHD8, ZEB1;

[0009] TCR signaling pathway related genes: RHOA, CARD11, FYN, YTHDF2, BIRC3, BIRC6, PLCG1, PLCG2, TAL1;

[0010] Genes related to the PI3K-AKT signaling pathway: TSC2, VAV1, PIK3R1, ITPR3, IKBKB, ITPKB, MTOR;

[0011] JAK-STAT signaling pathway related genes: PTPN13, JAK1, JAK2, JAK3, STAT3, STAT5B, PTPRC, PTPRD, PTPRS, SOCS1;

[0012] Immune escape-related genes: HLA-A, HLA-B, CD58, CIITA, PDCD1;

[0013] P53 signaling pathway related genes: TP53, ATM, JMY, CDKN2A, CHEK2, MSH2, MSH3, MSH6, PMS1, REV3L, CCND3;

[0014] Tumor suppressor genes: MGA, CIC, APC, APC2, LRP1B, NF1, PHLPP1, PTEN;

[0015] Genes related to the NOTCH signaling pathway: NOTCH1, NOTCH2, NOTCH3.

[0016] A second aspect of the present invention provides a method for screening a set of mutated genes for tumor molecular subtyping, the screening method comprising the following steps:

[0017] S1. Perform whole-exome sequencing on the sample to filter out non-tumor somatic mutations;

[0018] S2. Screening for reproducible mutations;

[0019] S3. Screen for lymphoma-related pathogenic and drug resistance genes.

[0020] A third aspect of the invention provides the application of a set of mutated genes for tumor molecular typing in the preparation of a detection product for tumor molecular typing.

[0021] Furthermore, the detection product may include at least one of a reagent kit, a gene chip, or a reagent combination.

[0022] A fourth aspect of the present invention provides the application of a set of mutant genes for tumor molecular typing in the preparation of a gene chip for tumor molecular typing, said gene chip comprising a solid support and probes.

[0023] In some embodiments, the tumor includes intranodal and extranodal T-cell lymphomas.

[0024] Furthermore, the intranodal and peripheral T-cell lymphomas include T1 type intranodal and peripheral T-cell lymphoma, T2 type intranodal and peripheral T-cell lymphoma, T3.1 type intranodal and peripheral T-cell lymphoma, and T3.2 type intranodal and peripheral T-cell lymphoma.

[0025] Furthermore, compared to other types of patients, those with T1 intranodal and peripheral T-cell lymphoma exhibit mutations in DNA methylation-related genes and TCR signaling pathway-related genes, characterized by RHOA and TET2 mutations. TET2 is a dioxygenase that plays a role in DNA demethylation; RHOA is involved in regulating cell proliferation and TCR signal transduction. Based on the characteristics of the coexisting mutated genes in patients, this type of patient may be considered for targeted therapy using the demethylating agent 5-azacytidine and the multi-kinase inhibitor dasatinib.

[0026] Furthermore, compared to other types of patients, those with T2 intranodal and peripheral T-cell lymphoma are characterized by mutations in DNA methylation-related genes and / or PI3K signaling pathway-related genes. Based on the characteristics of the coexisting mutated genes, this type of patient can be treated with a combination of DNA methylation-targeting 5-azacytidine and the PI3K inhibitor duvelisib.

[0027] Furthermore, compared to other types of patients, those with T3.1 intranodal and peripheral T-cell lymphoma have histone modification-related gene mutations, characterized by KMT2C / KMT2D mutations. These patients can be treated with a combination of epigenetic drugs, including chidamide and decitabine.

[0028] Furthermore, compared to other types of patients, those with T3.2 intranodal and peripheral T-cell lymphoma have immune escape-related gene mutations, characterized by mutations such as HLA-A / HLA-B / PTPN13. These patients can be treated with a combination of epigenetic drugs, apatinib, and anti-PD-1 antibodies.

[0029] A fifth aspect of the present invention provides a detection product for tumor molecular subtyping, the detection product comprising a detection reagent for detecting a set of mutated genes.

[0030] In some implementations, the detection product is an in vitro detection product.

[0031] A sixth aspect of the present invention provides a kit for tumor molecular typing, the kit comprising nucleic acid, oligonucleotide strands, and PCR primer set for detecting a set of mutated genes.

[0032] A seventh aspect of the present invention provides a method for detecting tumor molecular subtyping, the method comprising the following steps:

[0033] S1. Contact the biological sample with the set of mutated genes;

[0034] S2. Detect the expression pattern and level of the mutated gene set in the biological sample after contact, and determine the tumor type of the biological sample according to the PTCL algorithm.

[0035] In some embodiments, the tumor types include T1 type intranodal and peripheral T-cell lymphoma, T2 type intranodal and peripheral T-cell lymphoma, T3.1 type intranodal and peripheral T-cell lymphoma, and T3.2 type intranodal and peripheral T-cell lymphoma; the targeted therapy for T1 type intranodal and peripheral T-cell lymphoma is a combination of azacitidine and dasatinib; the targeted therapy for T2 type intranodal and peripheral T-cell lymphoma is a combination of azacitidine and duvelixe; the targeted therapy for T3.1 type intranodal and peripheral T-cell lymphoma is a combination of chidamide and decitabine; and the targeted therapy for T3.2 type intranodal and peripheral T-cell lymphoma is a combination of apatinib and a PD-1 inhibitor.

[0036] Based on the patient mutation gene set data obtained from the above steps, PCA analysis using the "prcomp" R software package and unsupervised hierarchical clustering analysis using the "pheatmap" R software package were performed. Patients were divided into four types ( Figure 1 ):

[0037] Patients with type T1 carry both RHOA and TET2 gene mutations;

[0038] T2 patients carry the TET2 mutation but do not have the RHOA gene mutation;

[0039] Patients with type T3.1 carry mutations in histone modification genes such as KMT2C and KMT2D;

[0040] Patients with type T3.2 carry mutations in immune escape-related genes such as HLA-A / HLA-B and / or PTPN13.

[0041] In some implementations, the PTCL algorithm specifically includes the following steps:

[0042] 1: First, use the library() function to call the R packages shiny, shinythemes, readxl, and dplyr.

[0043] 2: Use the read_excel function in the readxl package to read data from the file "PTCL.xlsx". The sheet specifies that the data should be read from the worksheet named "Gene" in the Excel file and assigned to a variable named PTCL.

[0044] 3: Extract the first column "Gene" of the "PTCL.xlsx" file using PTCL[,1], then transpose it using t(PTCL[,1]) and assign it to G.PTCL.

[0045] 4: By setting the value of the variable mut.PTCL to NULL by setting mut.PTCL to NULL, mut.PTCL is initialized to an empty value so that it can be used in subsequent code.

[0046] 5. Define the function MyPTCL to perform PTCL classification based on given mutation information. This function includes a series of operations, such as gene screening, data frame concatenation, and maximum value calculation. The specific operations are as follows:

[0047] A custom function `MyPTCL` accepts three parameters: `mut`, `unknown`, and `gene`. The default value of parameter `mut` is set to NULL using `mut = NULL`, and the default value of parameter `unknown` is set to NULL using `unknown = NULL`. Parameter `gene` has no default value. A conditional statement checks if the input mutation data `mut` is NULL (i.e., whether a checkbox on the UI page is selected). If no checkbox is selected, `mut` is set to NULL, and the string "which" is assigned to the variable `out`.

[0048] If the mutated data mut is not empty, the function will perform the following operations:

[0049] (1) First, filter the gene information data frame gene based on the unknown mutation information unknown. Keep the rows in the gene data frame whose first column value is not in unknown. The filtered results will be returned as a new data frame gene.

[0050] The specific steps are as follows: Use gene[,1] to extract the first column of the gene data frame, use the unlist() function to convert the first column into a one-dimensional vector, use the %in% operator to determine whether the value in the first column is in the unknown vector and return a logical vector (0 means not in unknown, 1 means in unknown), and reassign the filtered result to the gene variable. The resulting gene data frame is the result of removing the rows in the unknown vector from the first column.

[0051] (2) Convert the `mut` data to a data frame format using `mut = data.frame(mut)`, and set the column name of the `mut` data frame to "Gene" using `colnames(mut) = "Gene"`. Perform a left join operation using the `left_join()` function to match and merge the `mut` and `gene` data frames based on the values ​​in the "Gene" column. The left join operation retains all rows from the `mut` data frame and merges the rows from the matching `gene` data frames. The join operation is based on the values ​​in the "Gene" column; if the "Gene" values ​​on the left and right sides match, they are merged into one row. If there is no matching row on the right side, a missing value will be displayed at the corresponding position. By performing a left join operation, all rows in the `mut` data frame are retained, and the columns from the matching `gene` data frames are added to the result. In this way, each gene in the mutation data will be merged along with its corresponding grouping information.

[0052] (3) For the merged `mut` dataframe, extract the second column and calculate its minimum value, storing it in the variable `key`. Based on the value of `key`, filter the rows in the second column of the `mut` dataframe that are equal to `key`, and group them according to the "Group" column. Then, use the `summarise()` function to calculate the number of rows in each group and create a column named "Counts" to store the number of rows in each group. Finally, assign the result to the variable `jud`. This results in a dataframe `jud` grouped according to the "Group" column, where each group has a corresponding row count. Filter the maximum value in the "Counts" column of the `jud` dataframe, keeping only the rows where the "Counts" column equals the maximum value, and then extract the first column value of these rows, i.e., the value of the "Group" column, to obtain a result vector `res` storing the "Group" column values ​​that meet the conditions.

[0053] (4) Based on the code `if((unlist(res)%>%length()!=1)&(unique(mut[,2])%>%length()!=1)==1)`, a condition is determined. If the number of values ​​in `res` is not equal to 1 and the number of unique values ​​in the second column of the `mut` data frame is not 1, then the condition is met. Increment the value of the variable `key` by 1 and assign the result to the new variable `key.sub`. Based on `mut[,2]==key.sub`, filter out the rows in the second column of the `mut` data frame that are equal to `key.sub`. Use `group_by()` to group the data according to the `Group` column, use `summarise()` to perform statistics on each group, and add a new column `Counts` to count the number of rows in each group. The resulting `jud.sub` data frame is the result after filtering and statistics.

[0054] (5) Use `jud$Group` to get the value of the Group column in the `jud` data frame, and `jud.sub$Group` to get the value of the Group column in the `jud.sub` data frame. Use the `intersect()` function to calculate the intersection of the Group columns in `jud` and `jud.sub`, and assign the result to the variable `group.sub`. The resulting `group.sub` is a vector containing the Group values ​​that appear in both the `jud` and `jud.sub` data frames.

[0055] (6) group.sub stores the intersection of the Group column in two data frames jud and jud.sub. group.sub%>%length() passes group.sub to the length() function through the %>% operator to calculate the length of group.sub. If the length of group.sub is not zero, that is, the intersection result exists, then the rows in the res variable whose Group column is equal to group.sub are filtered according to the code res[res$Group==group.sub,], and the filtered result is updated to the new res.

[0056] (7) The code `out = paste0(unlist(res), collapse = "", recycle0 = T, sep = "~")` concatenates the values ​​in the variable `res`, specifying the delimiter and other parameters. The `unlist()` function converts the elements in the variable `res` into a single vector. The `paste0()` function concatenates multiple vectors or strings. `collapse = ""` specifies that the delimiter between the concatenated elements is an empty string. The `recycle0 = T` parameter repeats the concatenation of characters that are not long enough so that the final concatenation result has the same length as the original `res`. `sep = "~` specifies that the delimiter between the concatenated elements is a tilde (~). Finally, the concatenation result is stored in the variable `out`.

[0057] 6: Define the UI of the Shiny application, including a title panel, a checkbox group, and a main panel for displaying results. The specific steps are as follows:

[0058] The `fluidPage` function creates a page for the Shiny application, defining the application's appearance and layout in the UI section. `shinytheme("cerulean")` sets the application's theme to `cerulean`. `titlePanel("Classification for PTCL")` creates a title panel displaying the application's title as "Classification for PTCL". `checkboxGroupInput()` creates a checkbox component; `mut.PTCL` specifies the input for this checkbox component, `label="Mutations in"` specifies the label text, `choices=G.PTCL` specifies that the values ​​in the variable `G.PTCL` will be used as the checkbox options, and `inline=T` specifies that the checkboxes will be displayed inline. The `mainPanel` function creates a main panel, and the `uiOutput` function dynamically renders a UI output named "GroupN". `uiOutput("GroupN")` specifies a UI output object named "GroupN", which will be dynamically generated and rendered in subsequent server-side code.

[0059] 7. Define the server logic for the Shiny application, which includes an output function for rendering results. In this function, select mutation information is retrieved from user input, the MyPTCL function is called for classification, and the results are displayed as text with the title "Group N". The specific steps are as follows:

[0060] The Server section of the Shiny application defines the application's backend logic, used to process user input and generate corresponding output. `server<-function(input,output){...}` defines a function `server` that accepts two parameters, `input` and `output`. This is the server-side function of the Shiny application, used to process user input and generate output. `output$GroupN=renderUI({})` creates an output object named "GroupN" and dynamically generates and renders the UI output using the `renderUI` function. `mut.PTCL=input$mut.PTCL` assigns the user-selected `mut.PTCL` input value to the variable `mut.PTCL`. `Group.PTCL=MyPTCL(mut.PTCL,unknown=NULL,PTCL)` calls the custom function `MyPTCL`, passing the value of `mut.PTCL` and other parameters to generate the value of `Group.PTCL`. `h2(sprintf("Your patient is in%s\n group",Group.PTCL))` uses the `sprintf` function to create a header 2 element displaying information about the patient's `Group.PTCL` group.

[0061] 8: Finally, combine the defined UI and server-side functions using shinyApp(ui=ui, server=server) to create a complete Shiny application.

[0062] When the algorithm calculates that the gene weight is highest in group T1, it is identified as type T1; the gene weight is highest in group T2, it is identified as type T2; the gene weight is highest in group T3.1, it is identified as type T3.1; and the gene weight is highest in group T3.2, it is identified as type T3.2.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] 1. The mutant gene set of the present invention can be used for molecular subtyping, targeted therapy and prediction of overall survival of patients with intranodal and extranodal peripheral T-cell lymphoma. It has been verified by clinical trials that it is applicable to all newly diagnosed and relapsed PTCL patients, umbrella studies of relapsed or refractory peripheral T-cell lymphoma based on genomic subtyping, and PTCL patients receiving different treatment regimens, and has a very wide range of applications.

[0065] 2. This invention obtains raw data based on DNA sequencing of tumor tissue samples. It is applicable to various specimen formats such as paraffin and fresh samples, has low requirements for sample quality, and has the potential to be promoted to various hospitals. This will help doctors provide medication guidance, achieve precision medicine, improve the survival rate of patients with metastatic cancer of unknown primary lesion, and improve the patients' living conditions. It has extremely high clinical application prospects.

[0066] 3. The mutant gene set of this invention can directly reflect the molecular subtype genomic characteristics of PTCL patients: T1 patients carry RHOA and TET2 gene mutations, T2 patients carry only TET2 mutations, T3.1 patients carry histone modification gene mutations such as KMT2C, KMT2D and / or CREBBP, and T3.2 patients carry immune escape-related gene mutations such as HLA-A / HLA-B and / or PTPN13. This invention helps PTCL patients to accurately select targeted therapy drugs, which is of great significance for improving patient prognosis.

[0067] 4. The present invention provides clear criteria for molecular subtyping of patients with initial and recurrent PTCL, and the results are accurate, which can provide doctors with more precise and reliable treatment guidance, thereby reducing unnecessary overtreatment.

[0068] 5. This invention has excellent prospects for promotion from the perspectives of both effectiveness and economic benefits. Attached Figure Description

[0069] Figure 1 This represents the genetic characteristics of 221 PTCL patients in the experimental case, categorized into four types. Specifically, it shows the mutation rate of the characteristic gene in each of the four groups.

[0070] Figure 2 The genetic characteristics of 221 PTCL patients in the experimental case were divided into four types. The specific characteristics of each type are described as follows: Type T1 patients all carry RHOA and TET2 gene mutations, some with DNMT3A and IDH2 mutations; Type T2 patients carry TET2 mutations but no RHOA gene mutations; Type T3.1 patients carry histone modification gene mutations such as KMT2C, KMT2D, and CREBBP; Type T3.2 patients carry immune escape-related gene mutations such as HLA-A / HLA-B and / or PTPN13.

[0071] Figure 3 To compare the characteristics of different subtypes of PTCL at the transcriptome level. Compared with T2 patients, T1 and T2 patients showed activation of the TCR signaling pathway, while the PI3K signaling pathway was relatively suppressed; compared with T3.1 patients, T3.2 patients showed significant activation of the DNA replication pathway, while T3.2 patients showed activation of the VEGF signaling pathway.

[0072] Figure 4To illustrate the efficacy of combination therapy with different targeted drugs against different subtypes in in vitro and in vivo models. Figure A: A significant synergistic effect was observed in the combination of azacitidine and dasatinib in cell lines characteristic of T1 PTCL patients. In zebrafish characteristic of T1 PTCL patients, the survival time of zebrafish treated with the combination of azacitidine and dasatinib was significantly longer than that of the single-drug group and the untreated group. Figure B: A significant synergistic effect was observed in the combination of azacitidine and duveloxetine in cell lines characteristic of T2 PTCL patients. In zebrafish characteristic of T2 PTCL patients, the survival time of zebrafish treated with the combination of azacitidine and duveloxetine was significantly longer than that of the single-drug group and the untreated group. Figure C: A significant synergistic effect was observed in the combination of chidamide and decitabine in cell lines characteristic of T3.1 PTCL patients. In zebrafish exhibiting characteristics of T3.1 type PTCL patients, those treated with a combination of chidamide and decitabine showed significantly longer survival times than the single-drug and untreated groups. Figure D: In cell lines exhibiting characteristics of T3.2 type PTCL patients, the proliferation index was significantly reduced and Th2 cells were significantly decreased in the tumor microenvironment after the combined application of PD-1 inhibitors and apatinib.

[0073] Figure 5-7 This is a flowchart of the PTCL patient typing gene screening in Embodiment 2 of the present invention. Detailed Implementation

[0074] Example 1

[0075] This embodiment provides a set of mutated genes for tumor molecular subtyping and a screening method thereof.

[0076] The set of mutated genes includes the following genes:

[0077] DNA methylation-related genes: TET2, DNMT3A, IDH2, SALL3, TET1, TET3;

[0078] Histone modification-related genes: KMT2C, KMT2D, KDM6B, SETD2, KMT2A, SETBP1, SETD1B, CREBBP, EP300, NCOR2, MEF2A, TRRAP, SPEN, BCOR, YEATS2;

[0079] Chromatin remodeling-related genes: ARID1A, ARID1B, ARID2, ASXL3, SMARCA2, SMARCA4, CHD8, ZEB1;

[0080] TCR signaling pathway related genes: RHOA, CARD11, FYN, YTHDF2, BIRC3, BIRC6, PLCG1, PLCG2, TAL1;

[0081] Genes related to the PI3K-AKT signaling pathway: TSC2, VAV1, PIK3R1, ITPR3, IKBKB, ITPKB, MTOR;

[0082] JAK-STAT signaling pathway related genes: PTPN13, JAK1, JAK2, JAK3, STAT3, STAT5B, PTPRC, PTPRD, PTPRS, SOCS1;

[0083] Immune escape-related genes: HLA-A, HLA-B, CD58, CIITA, PDCD1;

[0084] P53 signaling pathway related genes: TP53, ATM, JMY, CDKN2A, CHEK2, MSH2, MSH3, MSH6, PMS1, REV3L, CCND3;

[0085] Tumor suppressor genes: MGA, CIC, APC, APC2, LRP1B, NF1, PHLPP1, PTEN;

[0086] Genes related to the NOTCH signaling pathway: NOTCH1, NOTCH2, NOTCH3.

[0087] The screening method includes the following steps:

[0088] S1. Perform whole-exome sequencing on the sample to filter out non-tumor somatic mutations;

[0089] S2. Screening for reproducible mutations;

[0090] S3. Screen for lymphoma-related pathogenic and drug resistance genes.

[0091] Example 2

[0092] This embodiment provides a detection method for tumor molecular subtyping.

[0093] The detection method includes the following steps:

[0094] S1. Contact the biological sample with the set of mutated genes;

[0095] S2. Detect the expression pattern and level of the mutated gene set in the biological sample after contact, and determine the tumor type of the biological sample according to the PTCL algorithm.

[0096] The PTCL algorithm specifically includes the following steps:

[0097] 1: First, use the library() function to call the R packages shiny, shinythemes, readxl, and dplyr.

[0098] 2: Use the read_excel function in the readxl package to read data from the file "PTCL.xlsx". The sheet specifies that the data should be read from the worksheet named "Gene" in the Excel file and assigned to a variable named PTCL.

[0099] 3: Extract the first column "Gene" of the "PTCL.xlsx" file using PTCL[,1], then transpose it using t(PTCL[,1]) and assign it to G.PTCL.

[0100] 4: By setting the value of the variable mut.PTCL to NULL by setting mut.PTCL to NULL, mut.PTCL is initialized to an empty value so that it can be used in subsequent code.

[0101] 5. Define the function MyPTCL to perform PTCL classification based on given mutation information. This function includes a series of operations, such as gene screening, data frame concatenation, and maximum value calculation. The specific operations are as follows:

[0102] A custom function `MyPTCL` accepts three parameters: `mut`, `unknown`, and `gene`. The default value of parameter `mut` is set to NULL using `mut = NULL`, and the default value of parameter `unknown` is set to NULL using `unknown = NULL`. Parameter `gene` has no default value. A conditional statement checks if the input mutation data `mut` is NULL (i.e., whether a checkbox on the UI page is selected). If no checkbox is selected, `mut` is set to NULL, and the string "which" is assigned to the variable `out`.

[0103] If the mutated data mut is not empty, the function will perform the following operations:

[0104] (1) First, filter the gene information data frame gene based on the unknown mutation information unknown. Keep the rows in the gene data frame whose first column value is not in unknown. The filtered results will be returned as a new data frame gene.

[0105] The specific steps are as follows: Use gene[,1] to extract the first column of the gene data frame, use the unlist() function to convert the first column into a one-dimensional vector, use the %in% operator to determine whether the value in the first column is in the unknown vector and return a logical vector (0 means not in unknown, 1 means in unknown), and reassign the filtered result to the gene variable. The resulting gene data frame is the result of removing the rows in the unknown vector from the first column.

[0106] (2) Convert the `mut` data to a data frame format using `mut = data.frame(mut)`, and set the column name of the `mut` data frame to "Gene" using `colnames(mut) = "Gene"`. Perform a left join operation using the `left_join()` function to match and merge the `mut` and `gene` data frames based on the values ​​in the "Gene" column. The left join operation retains all rows from the `mut` data frame and merges the rows from the matching `gene` data frames. The join operation is based on the values ​​in the "Gene" column; if the "Gene" values ​​on the left and right sides match, they are merged into one row. If there is no matching row on the right side, a missing value will be displayed at the corresponding position. By performing a left join operation, all rows in the `mut` data frame are retained, and the columns from the matching `gene` data frames are added to the result. In this way, each gene in the mutation data will be merged along with its corresponding grouping information.

[0107] (3) For the merged `mut` dataframe, extract the second column and calculate its minimum value, storing it in the variable `key`. Based on the value of `key`, filter the rows in the second column of the `mut` dataframe that are equal to `key`, and group them according to the "Group" column. Then, use the `summarise()` function to calculate the number of rows in each group and create a column named "Counts" to store the number of rows in each group. Finally, assign the result to the variable `jud`. This results in a dataframe `jud` grouped according to the "Group" column, where each group has a corresponding row count. Filter the maximum value in the "Counts" column of the `jud` dataframe, keeping only the rows where the "Counts" column equals the maximum value, and then extract the first column value of these rows, i.e., the value of the "Group" column, to obtain a result vector `res` storing the "Group" column values ​​that meet the conditions.

[0108] (4) Based on the code `if((unlist(res)%>%length()!=1)&(unique(mut[,2])%>%length()!=1)==1)`, a condition is determined. If the number of values ​​in `res` is not equal to 1 and the number of unique values ​​in the second column of the `mut` data frame is not 1, then the condition is met. Increment the value of the variable `key` by 1 and assign the result to the new variable `key.sub`. Based on `mut[,2]==key.sub`, filter out the rows in the second column of the `mut` data frame that are equal to `key.sub`. Use `group_by()` to group the data according to the `Group` column, use `summarise()` to perform statistics on each group, and add a new column `Counts` to count the number of rows in each group. The resulting `jud.sub` data frame is the result after filtering and statistics.

[0109] (5) Use `jud$Group` to get the value of the Group column in the `jud` data frame, and `jud.sub$Group` to get the value of the Group column in the `jud.sub` data frame. Use the `intersect()` function to calculate the intersection of the Group columns in `jud` and `jud.sub`, and assign the result to the variable `group.sub`. The resulting `group.sub` is a vector containing the Group values ​​that appear in both the `jud` and `jud.sub` data frames.

[0110] (6) group.sub stores the intersection of the Group column in two data frames jud and jud.sub. group.sub%>%length() passes group.sub to the length() function through the %>% operator to calculate the length of group.sub. If the length of group.sub is not zero, that is, the intersection result exists, then the rows in the res variable whose Group column is equal to group.sub are filtered according to the code res[res$Group==group.sub,], and the filtered result is updated to the new res.

[0111] (7) The code `out = paste0(unlist(res), collapse = "", recycle0 = T, sep = "~")` concatenates the values ​​in the variable `res`, specifying the delimiter and other parameters. The `unlist()` function converts the elements in the variable `res` into a single vector. The `paste0()` function concatenates multiple vectors or strings. `collapse = ""` specifies that the delimiter between the concatenated elements is an empty string. The `recycle0 = T` parameter repeats the concatenation of characters that are not long enough so that the final concatenation result has the same length as the original `res`. `sep = "~` specifies that the delimiter between the concatenated elements is a tilde (~). Finally, the concatenation result is stored in the variable `out`.

[0112] 6. Define the UI of the Shiny application, including a title panel, a checkbox group, and a main panel for displaying results. The specific steps are as follows:

[0113] The `fluidPage` function creates a page for the Shiny application, defining the application's appearance and layout in the UI section. `shinytheme("cerulean")` sets the application's theme to `cerulean`. `titlePanel("Classification for PTCL")` creates a title panel displaying the application's title as "Classification for PTCL". `checkboxGroupInput()` creates a checkbox component; `mut.PTCL` specifies the input for this checkbox component, `label="Mutations in"` specifies the label text, `choices=G.PTCL` specifies that the values ​​in the variable `G.PTCL` will be used as the checkbox options, and `inline=T` specifies that the checkboxes will be displayed inline. The `mainPanel` function creates a main panel, and the `uiOutput` function dynamically renders a UI output named "GroupN". `uiOutput("GroupN")` specifies a UI output object named "GroupN", which will be dynamically generated and rendered in subsequent server-side code.

[0114] 7. Define the server logic for the Shiny application, which includes an output function for rendering results. In this function, select mutation information is retrieved from user input, the MyPTCL function is called for classification, and the results are displayed as text with the title "Group N". The specific steps are as follows:

[0115] The Server section of the Shiny application defines the application's backend logic, used to process user input and generate corresponding output. `server<-function(input,output){...}` defines a function `server` that accepts two parameters, `input` and `output`. This is the server-side function of the Shiny application, used to process user input and generate output. `output$GroupN=renderUI({})` creates an output object named "GroupN" and dynamically generates and renders the UI output using the `renderUI` function. `mut.PTCL=input$mut.PTCL` assigns the user-selected `mut.PTCL` input value to the variable `mut.PTCL`. `Group.PTCL=MyPTCL(mut.PTCL,unknown=NULL,PTCL)` calls the custom function `MyPTCL`, passing the value of `mut.PTCL` and other parameters to generate the value of `Group.PTCL`. `h2(sprintf("Your patient is in%s\n group",Group.PTCL))` uses the `sprintf` function to create a header 2 element displaying information about the patient's `Group.PTCL` group.

[0116] 8: Finally, combine the defined UI and server-side functions using shinyApp(ui=ui, server=server) to create a complete Shiny application.

[0117] When the algorithm calculates that the gene weight is the highest in group T1, it is identified as type T1; the gene weight is the highest in group T2, it is identified as type T2; the gene weight is the highest in group T3.1, it is identified as type T3.1; and the gene weight is the highest in group T3.2, it is identified as type T3.2.

[0118] Experimental Example

[0119] The mutant gene set obtained in Example 1 and the detection method in Example 2 were used in a clinical application trial.

[0120] 1. Tumor tissues were obtained from 221 PTCL patients, and the mutated gene set obtained in Example 1 was detected to determine the mutation characteristics of the patients, as detailed below:

[0121] 1.1 Data Preparation

[0122] The “PTCL.xlsx” file contains three columns: “Gene,” “Key,” and “Group.” The “Gene” column represents the name of the gene, with each row corresponding to a specific gene. The “Key” column uniquely identifies each gene; each gene is assigned a unique “Key” for quick referencing and indexing during data processing and computation, ensuring accurate association between gene mutation and classification information. The “Group” column represents the group or category to which the gene belongs.

[0123] 1.2 Algorithm Principle

[0124] A custom function, MyPTCL, is used to classify patients by counting the frequency of gene types. The most frequent group is used as the initial classification criterion, and multiple most frequent groups are resolved according to certain rules. This function uses a gene's "Key" value to associate gene mutation information with gene classification information, ensuring accurate matching between gene classification and mutation information. The classification decision is made based on the frequency of gene types, and mainly consists of the following steps:

[0125] (1) Input data: Receive a list of gene mutation information as input. Each gene mutation information includes the gene name and its associated classification information.

[0126] (2) Initial classification: In the gene mutation information list, find the common classification standard for all genes, which is the smallest classification value among all genes. Perform initial classification of all genes according to this common classification standard.

[0127] (3) Count the number of times each group appears: Count the number of times each group appears in the gene data. This will give you a list of groups and their corresponding counts.

[0128] (4) Select the final classification result: Find the group that appears most frequently from the list of occurrences. If multiple groups have the same number of occurrences and are all the most frequent, select one of them as the final classification result.

[0129] (5) Handling multiple most frequent groups: If there are multiple most frequent groups, and the number of groups in the gene data is greater than 1, further processing is required. Find the next classification standard in the gene data and classify the genes again according to this new standard.

[0130] (6) Output: The final classification result will be output. If there are multiple groups that appear most frequently, output a list containing these groups. Otherwise, output a single final classification result.

[0131] 1.3 Creating the Program

[0132] This application uses the Shiny library to create an interactive web interface.

[0133] The user interface (UI) is defined: The UI uses the fluidPage function to define the application's user interface, which includes: the application's theme is set to "cerulean", the title panel is displayed as "Classification for PTCL", and checkboxes are used by the user to select gene mutation information.

[0134] The server logic is defined as follows: The server logic contains the renderUI() function, which classifies patients based on the gene mutation information selected by the user by calling the MyPTCL function and displays the results in the application's UI.

[0135] Run the application: Combine the defined UI and server-side functions to create a complete Shiny application.

[0136] 2. The patient's subtype and characteristics are determined using the PTCL molecular subtyping algorithm.

[0137] This includes T1 patients who carry both RHOA and TET2 gene mutations, and also have DNMT3A, IDH2, or CRAD11 mutation features; T2 patients who carry TET2 mutations do not have RHOA gene mutations; T3.1 patients may carry histone modification gene mutations, including KMT2C and / or KMT2D; T3.2 patients commonly have immune escape-related gene mutations, including HLA-A, HLA-B, and / or PTPN13.

[0138] 3. Different targeted therapies are applied according to different subtypes.

[0139] The application methods are as follows: For type T1, azacitidine and dasatinib are used in combination (azacitidine 100mg, once daily for days 1-7, intravenous drip; dasatinib 80mg, once daily, orally); for type T2, azacitidine and the PI3K inhibitor limprolix are used in combination (azacitidine 100mg, once daily for days 1-7, intravenous drip; limprolix 80mg, once daily, orally); for type T3.1, chidamide and SHR2554 are used in combination (chidamide 30mg, twice weekly, with an interval of at least 3 days between doses, orally; SHR2554 300mg, twice daily, orally); for type T3.2, apatinib mesylate and the PD-1 inhibitor camrelizumab are used in combination (apatinib mesylate 250mg, once daily, orally after meals; camrelizumab 200mg / dose, once every two weeks, intravenous drip).

[0140] Transcriptomic analysis revealed that the TCR signaling pathway was activated in group T1, and the PI3K signaling pathway was activated in group T2; DNA replication and transcription-related pathways were activated in group T3.1, and the VEGF signaling pathway was activated in group T3.2. Figure 3 Based on genomic and transcriptomic abnormalities, corresponding targeted drugs are selected for treatment. For T1 type, azacitidine, a hypomethylating drug targeting TET2 mutations, and dasatinib, a multi-kinase inhibitor that inhibits TCR signaling, are used in combination; for T2 type, azacitidine and PI3K signaling pathway inhibitor duvelixe are used in combination; for T3.1 type, the epigenetic modification drugs chidamide and decitabine are used in combination; and for T3.2 type, the VEGF inhibitor apatinib and a PD-1 inhibitor are used in combination. In vitro cell line models and in vivo zebrafish models have confirmed the synergistic anti-tumor effect of the two-drug combination in each group, improving the prognosis of patients in each subtype. Figure 4 ).

Claims

1. A set of mutated genes for tumor molecular subtyping, characterized in that, The set of mutated genes includes the following genes: DNA methylation-related genes: TET2, DNMT3A, IDH2, SALL3, TET1, TET3; Histone modification-related genes: KMT2C, KMT2D, KDM6B, SETD2, KMT2A, SETBP1, SETD1B, CREBBP, EP300, NCOR2, MEF2A, TRRAP, SPEN, BCOR, YEATS2; Chromatin remodeling-related genes: ARID1A, ARID1B, ARID2, ASXL3, SMARCA2, SMARCA4, CHD8, ZEB1; TCR signaling pathway related genes: RHOA, CARD11, FYN, YTHDF2, BIRC3, BIRC6, PLCG1, PLCG2, TAL1; Genes related to the PI3K-AKT signaling pathway: TSC2, VAV1, PIK3R1, ITPR3, IKBKB, ITPKB, mTOR; JAK-STAT signaling pathway related genes: PTPN13, JAK1, JAK2, JAK3, STAT3, STAT5B, PTPRC, PTPRD, PTPRS, SOCS1; Immune escape-related genes: HLA-A, HLA-B, CD58, CIITA, PDCD1; P53 signaling pathway related genes: TP53, ATM, JMY, CDKN2A, CHEK2, MSH2, MSH3, MSH6, PMS1, REV3L, CCND3; Tumor suppressor genes: MGA, CIC, APC, APC2, LRP1B, NF1, PHLPP1, PTEN; Genes related to the NOTCH signaling pathway: NOTCH1, NOTCH2, NOTCH3; The method for screening a set of mutated genes for tumor molecular subtyping includes the following steps: S1. Perform whole-exome sequencing on the sample to filter out non-tumor somatic mutations; S2. Screening for reproducible mutations; S3. Screening for lymphoma pathogenic and drug resistance-related genes; The tumor is a nodal and peripheral T-cell lymphoma.

2. A method for screening a set of mutated genes for tumor molecular subtyping as described in claim 1, characterized in that, The screening method includes the following steps: S1. Perform whole-exome sequencing on the sample to filter out non-tumor somatic mutations; S2. Screening for reproducible mutations; S3. Screening for lymphoma pathogenic and drug resistance-related genes; 3. The application of the mutant gene set for tumor molecular typing as described in claim 1 in the preparation of a detection product for tumor molecular typing.

4. The application of the mutant gene set for tumor molecular subtyping as described in claim 1 in the preparation of a gene chip for tumor molecular subtyping, characterized in that, The gene chip includes a solid support and probes.

5. A detection product for tumor molecular subtyping, characterized in that, The detection product includes a detection reagent for detecting the set of mutated genes as described in claim 1.

6. The testing product according to claim 5, characterized in that, The testing product is an in vitro testing product.

7. A reagent kit for tumor molecular subtyping, characterized in that, The kit includes nucleic acids, oligonucleotide chains, and PCR primers for detecting the mutant gene set as described in claim 1.