Methylation marker site and brainstem glioma typing and prognosis model
By screening out specific methylation marker sites and binding gene mutation information, a methylation model is constructed to perform typing and prognosis prediction of brainstem glioma, solving the inaccuracy and risk of typing and prognosis in the prior art, and achieving high accuracy and safety diagnosis and monitoring.
Patent Information
- Application Number
- CN202510551085.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art is difficult to effectively type and prognose brain stem gliomas, especially in the process of obtaining tumor tissues, which have high risks, inaccurateness and inability to monitor dynamically.
By screening out specific methylation marker sites and combining gene mutation information, a methylation model is constructed to perform typing and prognosis prediction of brainstem glioma. The model uses cerebrospinal fluid samples for detection, avoiding invasive operations and improving diagnostic accuracy and safety.
Highly accurate typing and prognosis prediction of brainstem gliomas is achieved, which reduces the risk of surgery, provides the possibility of dynamic monitoring, and improves the guiding significance of the treatment plan.
Smart Images

Figure CN120060477A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene detection, and particularly relates to a methylation biomarker locus, a brainstem glioma classification, and a prognosis model. Background Art
[0002] Brainstem gliomas (BSG), as one of the malignant brain tumors, have a median survival of only 10 months. As tumors with high heterogeneity, molecular pathological classification has gradually become an essential basis for guiding the precise diagnosis and treatment of BSG. According to the main molecular markers, BSG can be divided into three major subtypes (H3K27M, IDH, and Double-negative). In the H3K27M subtype tumors, genes encoding histones (HIST1H3B / C, HIST2H3C, H3F3A) mutate, resulting in the inability of the PRC2 protein to methylate histone H3K27, thereby leading to a decrease in the overall methylation level of cells; in IDH subtype tumors, the mutation of the IDH1 / 2 gene leads to the accumulation of 2-hydroxyglutaric acid (2-HG), which in turn affects the activity of the DNA demethylase TET protein, resulting in an increase in the overall methylation level of cells; the Double-negative subtype lacks obvious biomarkers, and the prognosis of this subtype is significantly better than that of other subtypes.
[0003] Currently, the main methods for obtaining tumor tissues for molecular pathological diagnosis are surgery or stereotactic puncture biopsy, but both of these methods have many disadvantages. First, they are both invasive operations with high risks. Second, patients can only obtain the molecular pathological diagnosis results after undergoing invasive examinations, which reduces the guiding significance of molecular pathological diagnosis in guiding patients to choose the correct treatment plan, including the guiding significance of avoiding unnecessary invasive operations. Third, there is sampling bias, making it difficult to reflect the spatial heterogeneity within the tumor. Fourth, it is difficult to perform repeatedly and impossible to dynamically monitor the tumor. The brainstem, as the core region of the central nervous system, encompasses key vital centers such as respiration and heartbeat, and the surgical risk is extremely high. A slight mistake during the operation may lead to serious complications or even death. In addition, brainstem gliomas usually grow infiltratively with blurred boundaries, further increasing the difficulty of surgical resection. Therefore, the tissue accessibility of brainstem glioma surgery is poor, while the accessibility of cerebrospinal fluid is relatively high. In contrast, cerebrospinal fluid detection shows significant advantages. Liquid biopsy technology does not require invasive operations and provides a reliable basis for precise diagnosis and treatment. For patients with brainstem gliomas for whom it is difficult to obtain tissues through surgery, cerebrospinal fluid detection is undoubtedly a better choice. Secondly, there are also problems inconsistent with clinical practice in using molecular pathological diagnosis. For example, patients with the Double negative subtype are usually considered to have a significantly better prognosis than other subtypes, but other clinical information indicates that some patients with this subtype have a poor prognosis, thus reducing the practical guiding significance of the classification. Therefore, there is an urgent need to develop new diagnostic methods with high accuracy, non-invasive or minimally invasive, and simple in clinical practice. Summary of the Invention
[0004] To solve the above problems, the present invention provides a methylation marker site, a classification of brainstem gliomas, and a prognostic model.
[0005] To achieve the above object, the technical solutions adopted by the present invention are as follows: The present invention provides methylation marker sites for the classification and prognostic stratification of brainstem gliomas, including: with hg19 as the reference genome, the methylation marker sites include: chr6: 36355479-36355532, chr17: 74497236-74497334, chr17: 79480435-79480471, chr1: 13910793-13910796, chr5: 132082727-132082730, chr8: 82193629-82193706, chr10: 72200923-72200938, chr10: 130339528-130339595, chr12: 123380366-123380410, chr2: 171568407-171568410, chr3: 71631220-71631223, chr3: 101568920-101568923, chr3: 139258596-139258642, chr5: 176170373-176170376, chr9: 82188499-82188502, chr12: 53614079-53614082, chr17: 38334060-38334063, chr19: 11450022-11450036, chr19: 52207581-52207591, chr1: 13910566-13910609, chr1: 13910697-13910737, chr3: 134032302-134032392, chr5: 132082823-132082826, chr1: 11752202-11752209, chr1: 155043729-155043732, chr4: 54965828-54965831, chr11: 72463414-72463417, chr12: 52445120-52445197, chr19: 52207340-52207343, chr9: 109623014-109623066, chr19: 19281254-19281272, chr1: 210466206-210466209, chr1: 228783348-228783351, chr10: 130339689-130339692, chr10: 130339691-130339694, chr12: 49740749-49740752, chr17: 74497720-74497723, chr2: 201983198-201983201,chr2: 201983331 - 201983334, chr3: 197183565 - 197183568, chr11: 19735659 - 19735662, chr11: 72387963 - 72387966, chr11: 463092 - 72463095, chr19: 19281042 - 19281045, chr1: 220921501 - 220921504, chr4: 108745592 - 108745595, chr20: 10652811 - 10652814.,
[0006] An embodiment of the present invention also provides a method for screening the above methylation marker sites, including the following steps: collecting methylation chip data related to brainstem glioma through a public database; performing normalization processing on the methylation chip data to obtain a β matrix; respectively performing pairwise comparisons for H3K27M, IDH, and Double - negative of BSG to obtain differentially methylated sites; screening out methylated sites with relatively large differences by setting conditions of Δβ≥0.3 and SD≤0.1; using a variety of feature selection algorithms to screen important CpGs, and finally obtaining the methylation marker sites.
[0007] Further, the variety of feature selection algorithms include: Random Forest, Support Vector Machine, and Lasso. Screen the CpGs jointly retained by at least two algorithms to obtain the important CpGs.
[0008] An embodiment of the present invention also provides a method for constructing a methylation model for non - diagnostic and therapeutic purposes, including the following steps: for known samples, sequencing the above methylation marker sites to obtain sequencing data; performing outlier processing on the sequencing data; using the processed sequencing data for model training, the model being a random forest model, and the random forest model outputting the prediction probabilities of three subtypes of BSG; using test samples to verify the trained random forest model.
[0009] An embodiment of the present invention also provides a methylation model, and the methylation model is constructed by the above construction method.
[0010] An embodiment of the present invention also provides a method for classifying brainstem glioma according to the above methylation model and gene mutations, and the method is for non - diagnostic and therapeutic purposes, including: collecting cerebrospinal fluid of a patient; performing gene mutation detection and methylation level detection on the cerebrospinal fluid; based on the methylation detection data, calculating the prediction probabilities of three subtypes of BSG through the methylation model; determining the classification of the patient based on the gene mutation information and the prediction probabilities.
[0011] Furthermore, the genes include IDH1 / 2, HIST1H3B / C, HIST2H3C, and H3F3A genes. Furthermore, the methylation level detection region includes methylation marker sites.
[0012] Furthermore, the typing of the patient is determined based on gene mutation information and methylation information, including: The IDH1 / 2, HIST1H3B / C, HIST2H3C, and H3F3A genes are used for typing the sample, marking the mutant type as 1 and the wild type as 0; the predicted probabilities of the three subtypes are obtained based on the methylation model respectively, and the value range is between 0 and 1; the mutation result is added to the predicted probability to obtain the scores of the three subtypes, and the subtype with the highest score is the final subtype of the sample.
[0013] An embodiment of the present invention also provides a computer storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0014] An embodiment of the present invention also provides a prognostic stratification model for BSG patients. Based on cerebrospinal fluid, the predicted probability of the H3K27M subtype is calculated through the above methylation model; a survival curve is drawn based on the survival package, and the predicted probability value when the Youden index is the largest is selected as the threshold for stratification; when the predicted probability is greater than the threshold, the sample is at high risk; when the predicted probability is not greater than the threshold, the sample is at low risk.
[0015] The beneficial effects brought by the technical solution provided by the embodiment of the present invention include: The methylation marker sites proposed in the embodiment of the present invention are closely related to the typing and prognosis of brainstem glioma. First, abnormal DNA methylation is highly correlated with the occurrence and development of tumors. Central nervous system tumors have obvious differences in DNA methylation characteristics, and typing based on DNA methylation characteristics can better reflect the actual clinical situation. Research has found that DNA methylation related to BSG shows obvious overall methylation characteristics of different subtypes and has good discrimination. However, there are also corresponding problems: how to select the corresponding methylation marker sites is a technical problem to be solved. The applicant uses a specific selection method to select the corresponding methylation marker sites, which can well predict the typing and prognosis of glioma and is more in line with the actual clinical situation; finally, the methylation markers screened by the present invention have high specificity in different subtypes of BSG. The machine learning model trained by the methylation markers combined with mutation information has very high accuracy and specificity, and the model constructed based on the above methylation markers also has high accuracy and specificity for cerebrospinal fluid samples, providing the possibility for subsequent tumor dynamic monitoring. The methylation data prediction results have guiding significance for patient prognosis stratification. Brief Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It shows the methylation status of the methylation marker in the BSG sample provided by Embodiment 1 of the present invention; Figure 2 It is a schematic diagram for jointly determining the sample typing by the gene mutation and methylation model provided by Embodiment 2 of the present invention; Figure 3 It is the stratified prognostic survival curve of the methylation model for the H3K27M subtype provided by Embodiment 3 of the present invention. Detailed Description of the Embodiments
[0018] The present invention will be further described in detail below through specific embodiments. However, those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be construed as limiting the scope of the present invention. For those not specified in the embodiments regarding specific techniques or conditions, they shall be carried out according to the techniques or conditions described in the literature in this field or according to the product specifications. For reagents or instruments not specified as to the manufacturer, they are all conventional products that can be obtained commercially.
[0019] As used herein, the words "comprising", "including", "having" or any other variant thereof are intended to cover non-exclusive inclusion. For example, a process, method, article or apparatus that includes the recited elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article or apparatus. Unless the context clearly dictates otherwise, the singular forms "a / an" and "the" include plural referents.
[0020] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art. In addition to the specific methods, equipment, and materials used in the embodiments, according to the knowledge of those skilled in the art in the prior art and the description of the present invention, any methods, equipment, and materials similar or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0021] Unless otherwise stated, the experimental methods, detection methods, and preparation methods disclosed in the present invention all adopt the conventional techniques in the fields of molecular biology, immunology, laboratory medicine, gene sequencing technology, bioinformatics technology, and related fields in this technical field.
[0022] An embodiment of the present invention provides a methylation marker locus for the classification and prognostic stratification of brainstem gliomas, which is closely related to the classification and prognosis of brainstem gliomas. First, abnormal DNA methylation is highly correlated with the occurrence and development of tumors. Central nervous system tumors have obvious differences in DNA methylation characteristics, and classification based on DNA methylation characteristics can better reflect the actual clinical situation. Research has found that DNA methylation studies related to BSG show that the overall methylation characteristics of different subtypes are obvious and have good discrimination. However, there are also corresponding problems: how to select the corresponding methylation marker locus is a technical problem that needs to be solved. The applicant uses a specific selection method to select the corresponding methylation marker locus, which can well predict the classification and prognosis of gliomas and is more in line with the actual clinical situation. Finally, the methylation markers screened by the present invention have high specificity in different subtypes of BSG. The machine learning model trained using methylation markers combined with mutation information has very high accuracy and specificity, and the model constructed based on the above methylation markers also has high accuracy and specificity for cerebrospinal fluid samples, providing the possibility for subsequent tumor dynamic monitoring. The prediction results of methylation data have guiding significance for the prognostic stratification of patients.
[0023] The screening method of the methylation marker locus includes the following steps: S1. Collect methylation chip data related to brainstem gliomas through public databases.
[0024] Methylation chip data related to brainstem glioma tissue samples, including GSE90496, GSE109379, GSE161944, GSE50022, GSE64509, and EGAS00001004341, were collected through public databases.
[0025] S2. Perform standardization processing on the methylation chip data to obtain a β matrix.
[0026] The original methylation chip data was processed using the minfi (v1.50.0) package in R language (v4.3.2) to obtain a standardized β matrix.
[0027] S3. Perform pairwise comparisons for H3K27M, IDH, and Double-negative of BSG respectively to obtain differentially methylated sites.
[0028] S4. Screen out methylation sites with larger differences by setting conditions of Δβ≥0.3 and SD≤0.1.
[0029] S5. Use various feature selection algorithms to screen important CpGs, and finally obtain the methylation marker locus.
[0030] The multiple feature selection algorithms include: Random Forest (RF), Support Vector Machine (SVM), and Lasso. At least two algorithms are screened to jointly retain CpGs, and the important CpGs are obtained. In the above algorithms, SVM and Lasso will fit the CpGs and retain the CpGs with non-zero coefficients, while RF retains them according to the importance of CpGs.
[0031] The methylation marker site information for brainstem glioma typing and prognosis stratification is specifically shown in Table 1 (using hg19 as the reference genome). Table 1 Methylation site information
[0032] The embodiment of the present invention also discloses a method for constructing a methylation model for non-diagnostic and therapeutic purposes, including the following steps: S10. For known samples, sequence the above methylation marker sites to obtain sequencing data.
[0033] The methylation detection and data analysis process is as follows: (1) Sample DNA extraction; For tumor tissue samples, use a nucleic acid extraction and purification kit (QIAamp DNA Mini Kit 250, QIAGEN) to extract and purify DNA. Use a DNA nucleic acid quantification kit (Qubit dsDNA HS Assay Kits, ThermoFisher Scientific) to detect the concentration of the extracted DNA. Samples with a total DNA amount greater than 10 ng go through the subsequent library construction process, and samples with a total amount less than 10 ng are re-extracted for sample DNA.
[0034] (2) Library construction and sequencing; Adopt the method of patent ZL201910983038.8 for library construction, and its process is as follows: S10. First, use methylation-sensitive restriction endonucleases (HhaI and HinpI) to digest the DNA, then fill in and add A and adapters to form a pre-library, then amplify and purify the pre-library, take 400 ng of the pre-library and use the designed primers to amplify the target region, and then form a methylation detection library according to the method of patent ZL201910983038.8. Finally, according to the instructions for the sequencing instrument, use an Illumina or BGI sequencer for on-machine sequencing.
[0035] Table 2 Methylation site primer information
[0036] Among them, A represents the primer of the positive strand, B represents the primer of the negative strand, and 1A and 2A, 1B and 2B are the primers of two different regions at the corresponding sites respectively.
[0037] S20. Perform outlier processing on the sequencing data.
[0038] Outlier processing: When there is no reads coverage in the target region, it is defined as a missing value, and samples with a missing value ratio higher than 10% are removed; for partial missing values, the average reads value at this site in the sample is used for filling.
[0039] S30. Use the processed sequencing data to train the model. The model is a random forest model, and the random forest model outputs the prediction probabilities of three subtypes of BSG.
[0040] Use the random forest model in caret (v4.3.3) in R language (v4.3.2): model = train(data = data, method = 'rf', trControl, metric = 'kappa'), where data is the training data, method defines the method adopted by the model, trControl defines the way of cross-validation, 10-fold cross-validation is adopted, metric defines the parameter for selecting the optimal model, and kappa is used as the parameter for optimizing the classification model.
[0041] S40. Use the test samples to verify the trained random forest model.
[0042] Use the samples in the test set for model testing: predict(object, testdata, type = 'prob'), where object is the trained model, testdata is the test set data, type is the format of the output prediction result, and finally the output result is the prediction probabilities of three subtypes of BSG. The subtype with the largest of the three prediction probabilities is used as the final sample subtype prediction.
[0043] The embodiment of the present invention also provides a methylation model, and the methylation model is constructed by the above construction method.
[0044] The embodiment of the present invention also provides a method for classifying brainstem gliomas according to the above methylation model and gene mutations. The method is for non-diagnostic and non-therapeutic purposes and includes: S100. Collect the cerebrospinal fluid of the patient. The cfDNA in the cerebrospinal fluid sample is extracted using the Apostle MiniMax high efficiency cfDNA isolation kit (Apostle; Pleasanton, CA, USA).
[0045] S200. Perform gene mutation detection and methylation level detection on the cerebrospinal fluid.
[0046] The genes include IDH1 / 2, HIST1H3B / C, HIST2H3C, and H3F3A genes.
[0047] The methylation level detection region includes methylation marker sites.
[0048] The library construction is carried out using the method of patent ZL201910983038.8, and its process is as follows: First, use methylation-sensitive restriction endonucleases (HhaI and HinpI) to digest the DNA, then fill in and add A and adapters to form a pre-library, then amplify and purify the pre-library. Take 1 μg of the pre-library and use the 68-gene detection product for primary brain tumors (Beijing Genepioneer Medical Laboratory Co., Ltd.) to capture the target region to form a library for mutation detection. Take 400 ng of the pre-library and use the designed primers to amplify the target region, and then form a methylation detection library according to the method of patent ZL201910983038.8. Finally, according to the instructions for loading the sequencer, use an Illumina or BGI sequencer for loading and sequencing.
[0049] S300. Based on the gene sequencing data, obtain the gene mutation detection result, mark the mutant type as 1, and the wild type as 0.
[0050] S400. Based on the methylation sequencing data, calculate the predicted probabilities of the three subtypes of BSG through the methylation model, and the value range is between 0 and 1.
[0051] S500. Determine the classification of the patient based on the gene mutation information and the predicted probability of the methylation model. Add the gene mutation detection result and the predicted probability of the methylation model to obtain the scores of the three subtypes, and take the subtype with the highest score as the final subtype of the sample.
[0052] It should be noted that although molecular pathology is the gold standard for diagnostic classification, it has certain limitations. As described in the background art, molecular pathology requires invasive sampling for analysis. However, for brainstem gliomas, invasive sampling causes great harm to the body. Therefore, ctDNA in cerebrospinal fluid is considered. This is because due to the blood-brain barrier, the content of ctDNA (circulating tumor DNA) in the blood is limited for brain tumors. As a special body fluid in the brain, cerebrospinal fluid fills the ventricles, subarachnoid space, and central canal of the spinal cord, and has a relatively high content of ctDNA, which can better represent the true information of the tumor. Therefore, liquid biopsy of cerebrospinal fluid will help overcome the clinical diagnostic dilemma of BSG and achieve dynamic monitoring. In addition, single-gene variant detection is limited by single-point and qualitative biomarker detection. In cases where the tumor content is low or the tumor release is limited, there is a high probability of missed detection; while the methylation multi-site and quantitative biomarker characteristics predict subtypes, and through the combined detection of multiple sites, the limitations of single-point gene variant detection can be compensated for probabilistically. Therefore, the combination of methylation multi-sites and gene mutations is more practically instructive.
[0053] The present invention also provides a computer storage medium storing a computer program, which when executed by a processor implements the method as described above.
[0054] The embodiment of the present invention also provides a prognostic model for BSG patients. Based on ctDNA in cerebrospinal fluid, the prediction probability of the H3K27M subtype is calculated through the above methylation model; a survival curve is drawn based on the survival package, and the prediction probability value when the Youden index is the largest is selected as the threshold for stratification; when the prediction probability is greater than the threshold, the sample is at high risk; when the prediction probability is not greater than the threshold, the sample is at low risk.
[0055] Example 1
[0056] A machine learning model of methylation markers.
[0057] The inventor collected tissue samples from 67 BSG patients. All enrolled patients signed informed consent forms. These patients were divided into a training set and a test set according to a certain ratio. The training set was used for the construction of the following machine learning model, and the test set was used for the performance test of the model. The sample sizes of each subtype in each dataset are shown in the following table.
[0058] Table 3 Sample Allocation
[0059] The methylation detection and data analysis process is as follows: (1) Sample DNA extraction; Tumor tissue samples were used for DNA extraction and purification using a nucleic acid extraction and purification kit (QIAamp DNA Mini Kit 250, QIAGEN). The concentration of the extracted DNA was detected using a DNA nucleic acid quantification kit (Qubit dsDNA HS Assay Kits, ThermoFisher Scientific). Samples with a total DNA amount greater than 10 ng were subjected to the subsequent library construction process, and samples with a total amount less than 10 ng were re-extracted for sample DNA.
[0060] (2) Library construction and sequencing; The library was constructed using the method of patent ZL201910983038.8. The general process is as follows: First, the DNA was digested with methylation-sensitive restriction endonucleases (HhaI and HinpI), then filled in and A-added and ligated with adapters to form a pre-library. Then, the pre-library was amplified and purified. 400 ng of the pre-library was used to amplify the target region using the designed primers, and then a methylation detection library was formed according to the method of patent ZL201910983038.8. Finally, according to the instructions for loading the sequencer, an Illumina or BGI sequencer was used for loading and sequencing.
[0061] (3) Analysis of the data after sequencing; The data after sequencing was quality-controlled and then aligned with the human reference genome hg19. Data analysis was performed according to the method of patent ZL201910983038.8 to obtain the methylation information of the target region. The distribution of methylation at the target sites in different samples is shown in Figure 1 .
[0062] (4) Construction and validation of the methylation model; The methylation model was constructed based on the methylation information of the target region. The specific steps are as follows: 1. Outlier handling: Samples with no reads coverage in the target region were defined as missing values. Samples with a missing value ratio higher than 10% were removed; for partial missing values, the average reads value at this site in the sample was used for filling; 2. Methylation model training: The random forest model in caret (v4.3.3) in R language (v4.3.2) was used: model = train(data = data, method = 'rf', trControl, metric = 'kappa'), where data is the training data, method defines the method used for the model, trControl defines the way of cross-validation, 10-fold cross-validation was adopted, and metric defines the parameter for selecting the optimal model. Kappa was used as the parameter for optimizing the classification model.
[0063] 3. Methylation model testing: Use the samples in the test set to test the model with the predict function: predict(object, testdata, type = 'prob'), where object is the trained model, testdata is the test set data, and type is the output prediction result format. The final output result is the predicted probabilities of the three subtypes of BSG. Use the subtype with the largest of the three predicted probabilities as the final sample subtype prediction.
[0064] 4. Statistics of methylation model performance indicators: Statistically analyze the accuracy, specificity, and AUC of the model on the training set and test set respectively. The sensitivity, specificity, and AUC of the methylation model on the training set are shown in the following table. The sensitivity and specificity of the model for the three subtypes on the training set both reach 1, and the AUC is 1.
[0065] Table 4 Performance of the methylation model on the training set
[0066] The accuracy, specificity, and AUC of the methylation model on the test set are shown in the following table. The sensitivity of the IDH subtype of the model on the test set is 0.5, and the other two subtypes are 1. Combining the mutation detection results of the IDH1 gene (VAF is 0.047) indicates that the tumor cell content in this sample is less, and usually the minimum detection limit of the methylation level is 5%, which ultimately leads to misclassification of the sample.
[0067] Table 5 Performance of the methylation model on the test set
[0068] Example 2 Detection of BSG cerebrospinal fluid samples.
[0069] Collect tissue samples and cerebrospinal fluid samples from 29 BSG patients respectively. The cerebrospinal fluid samples are used to extract cfDNA with the ApostleMiniMax high efficiency cfDNA isolation kit (Apostle; Pleasanton, CA, USA). The tissue samples are molecularly typed using the 68-gene detection product for primary brain tumors from PanCancer as the reference standard.
[0070] (1) Use the methylation marker data of cerebrospinal fluid for sample typing.
[0071] To further verify the detection performance of the above markers and methylation model for cerebrospinal fluid samples, the methylation detection and data processing of 29 cerebrospinal fluid samples are also carried out according to the method in Example 1. The specific sample information is shown in the following table.
[0072] Table 6 Number of cerebrospinal fluid samples
[0073] Sample typing prediction was performed using the methylation marker data alone. The trained methylation typing model in Example 1 above was used to perform typing prediction on the above cerebrospinal fluid samples, and the prediction results are shown in the following table. Among them, the sensitivity and specificity of the H3K27M subtype were 0.783 and 0.833 respectively, the sensitivity and specificity of the IDH subtype were 0.000 and 1.000 respectively, and the sensitivity and specificity of the Double negative subtype were 0.750 and 0.720 respectively. Given the significant differences between cerebrospinal fluid and tissue samples, firstly, the content of ctDNA in cerebrospinal fluid is low, and secondly, not all methylation markers enter cerebrospinal fluid, resulting in misclassification of individual samples.
[0074] Table 7 Detection performance of methylation model for cerebrospinal fluid samples
[0075] (2)Sample typing was performed using the gene mutation information of cerebrospinal fluid.
[0076] The same samples as in (1) were used for the detection and data processing of gene mutations.
[0077] The library construction was carried out using the method of Patent ZL201910983038.8, and the process was as follows: First, DNA was digested with methylation-sensitive restriction enzymes (HhaI and HinpI), then filled in and A and adapters were added to form a pre-library, then the pre-library was amplified and purified, 1 μg of the pre-library was taken and the target region was captured using the 68-gene detection product for primary brain tumors (Beijing Genepioneer Medical Laboratory Co., Ltd.) to form a library for mutation detection, and finally, according to the instructions for the sequencer, Illumina or BGI sequencer was used for on-machine sequencing.
[0078] The samples were typed based on the mutation results of IDH1 / 2, HIST1H3B / C, HIST2H3C, and H3F3A genes. Among them, samples with IDH1 / 2 gene mutations were of the IDH subtype, samples with any gene mutations in HIST1H3B / C, HIST2H3C, and H3F3A were of the H3K27M subtype, and samples without the above gene mutations were of the Double negative subtype. The specific typing results are shown in the following table, where the sensitivity and specificity of the H3K27M subtype were 0.913 and 1.000 respectively, the sensitivity and specificity of the IDH subtype were 1.000 and 1.000 respectively, and the sensitivity and specificity of the Double negative subtype were 1.000 and 0.920 respectively.
[0079] Table 8 Mutation detection performance of cerebrospinal fluid samples
[0080] (3)Perform sample typing using gene mutation and methylation marker data.
[0081] First, use the IDH1 / 2, HIST1H3B / C, HIST2H3C, and H3F3A genes for sample typing, mark the mutant type as 1, and the wild type as 0; second, use the trained methylation model in Example 2 to predict the methylation data to obtain the predicted probabilities of the three subtypes, with the value range between 0 and 1; finally, add the mutation results and the methylation model prediction results to obtain the scores of the three subtypes, and take the subtype with the highest score as the final subtype of the sample. The specific steps are as Figure 2 shown, and the specific typing results are shown in the following table. The sensitivity and specificity of the H3K27M subtype are 0.956 and 0.833 respectively, the sensitivity and specificity of the IDH subtype are both 1, and the sensitivity and specificity of the Double negative subtype are 0.750 and 0.960 respectively. The combination of mutation and methylation markers significantly improves the accuracy of BSG cerebrospinal fluid typing. Although 1 patient with the Double negative subtype was misclassified, other clinical information indicates that the patient is of the GBM (glioblastoma) subtype with a poor prognosis and a survival time of only 250 days, which is closer to the clinical manifestation of the H3K27M subtype.
[0082] Table 9 Methylation model + mutation detection performance of cerebrospinal fluid samples
[0083] Example 3 Prognostic stratification of cerebrospinal fluid in BSG patients.
[0084] The prognostic differences among different subtypes of BSG patients are relatively large, and the prognosis of the H3K27M subtype is the worst. To better achieve the prognostic stratification of H3K27M patients, use the cerebrospinal fluid methylation marker data of H3K27M subtype patients before surgery or during surgery in Example 2 combined with the patient survival information for prognostic analysis. After the sample methylation data is basically processed, use the trained methylation model in Example 1 to predict the samples to obtain the predicted probabilities of the H3K27M subtype of all samples, that is, the risk score of H3K27M. Use the survival package to draw the survival curve, and select the risk score when the youden index is the largest as the threshold to stratify the samples. The results are shown in Figure 3When the risk score > 0.44, the sample is of high risk, and its survival time is 328 days (95% CI, 166 - 527 days); otherwise, it is of low risk, and its survival time is 769 days (95% CI, 281 - not reached), with a P value of 0.023. The results indicate that the prediction results of the methylation classification model can be used for prognostic stratification of patients with the H3K27M subtype.
[0085] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A methylation marker site for brainstem glioma classification and prognostic stratification, characterized in that: include: Taking hg19 as the reference genome, the methylation marker sites include: chr6:36355479-36355532,chr17:74497236-74497334,chr17:79480435-79480471,chr1:13910793-13910796,chr5:132082727-132082730,chr8:82193629-82193706,chr10:72200923-72200938,chr10:130339528-130339595,chr12:123380366-123380410,chr2:171568407-171568410,chr3:71631220-71631223,chr3:101568920-101568923,chr3:139258596-139258642,chr5:176170373-176170376,chr9: 82188499-82188502,chr12:53614079-53614082,chr17:38334060-38334063,chr19:11450022-11450036,chr19:52207581-52207591,chr1:13910566-13910609,chr1:13910697-13910737,chr3:134032302-134032392,chr5:132082823-132082826,chr1:11752202-11752209,chr1:155043729-155043732,chr4:54965828-54965831,chr11:72463414-72463417,chr12:52445120-52445197,chr19:52207340-52207343,chr9:109623014-109623066,chr19:19281254-19281272,chr1:210466206-210466209,chr1:228783348-228783351,chr10:130339689-130339692,chr10:130339691-130339694,chr12:49740749-49740752,chr17:74497720-74497723,chr2:201983198-201983201,chr2:201983331-201983334,chr3:197183565-197183568,chr11:19735659-19735662,chr11:72387963-72387966,chr11:463092-72463095,chr19:19281042-19281045,chr1:220921501-220921504,chr4:108745592-108745595,chr20:10652811-10652814。, 2. The method for screening methylation marker sites according to claim 1, characterized in that: The steps include: Collect brainstem glioma-related methylation chip data through public databases; Standardizing the methylation chip data to obtain a β matrix; The H3K27M, IDH and Double-negative of BSG were compared pairwise to obtain differentially methylated sites; By setting the conditions of Δβ≥0.3 and SD≤0.1, we screened out methylation sites with large differences; A variety of feature selection algorithms are used to screen important CpGs, and finally the methylation marker sites are obtained.
3. The screening method according to claim 2, characterized in that The multiple feature selection algorithms include: Random Forest, Support Vector Machine and Lasso, and CpGs retained by at least two algorithms are screened to obtain the important CpGs.
4. A method for constructing a methylation model for non-diagnostic and non-therapeutic purposes, characterized in that: The steps include: For a known sample, sequencing the methylation marker site described in claim 1 or the methylation marker site obtained by the screening method described in any one of claims 2 to 3 to obtain sequencing data; Performing outlier processing on the sequencing data; The processed sequencing data is used to train a model, wherein the model is a random forest model, and the random forest model outputs the predicted probabilities of the three subtypes of BSG; The test samples are used to validate the trained random forest model.
5. A methylation model, wherein the methylation model is constructed using the construction method of claim 4.
6. The method for classifying brainstem gliomas using a methylation model and gene mutation according to claim 5, wherein the method is used for non-diagnostic and therapeutic purposes, and is characterized in that: include: Collect cerebrospinal fluid from patients; Performing gene mutation detection and methylation level detection on the cerebrospinal fluid; Based on the methylation detection data, the prediction probability of the three BSG subtypes was calculated by the methylation model; The patient's typing is determined based on the gene mutation information and the predicted probability.
7. The method according to claim 6, characterized in that The genes include IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes.
8. The method according to claim 6, characterized in that The patient's typing is determined based on gene mutation information and methylation information, including: The samples were typed using IDH1 / 2, HIST1H3B / C, HIST2H3C, and H3F3A genes, with the mutant type marked as 1 and the wild type as 0; Based on the methylation model, the predicted probabilities of the three subtypes were obtained, with the value range being between 0 and 1; The mutation results are added to the predicted probability to obtain the scores of the three subtypes, and the subtype with the largest score is the final subtype of the sample.
9. A computer storage medium, wherein the storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 6 to 8.
10. A prognostic model for BSG patients, characterized in that Based on the cerebrospinal fluid, the predicted probability of the H3K27M subtype is calculated by the methylation model as described in claim 5; The survival curve is drawn based on the survival package, and the predicted probability value when the Youden index is the largest is selected as the threshold for stratification; When the predicted probability is greater than the threshold, the sample is high risk; When the predicted probability is not greater than the threshold, the sample is low risk.
Citation Information
Patent Citations
High efficiency targeted in situ genome-wide profiling
CN111727248A
Method for detecting variation and methylation of tumor specific genes in ctDNA
CN112176419A
Molecular markers for glioma prognosis typing and typing method and application thereof
CN114381525A
Biomarker, kit and system for evaluating prognosis of stage II colorectal cancer
CN115820849A
Methylation marker for detecting benign and malignant pulmonary nodules, evaluation model and application
CN118166108A
Cited By
Gene methylation marker screening method, kit, storage medium and system
CN121662164A
DNA methylation marker combination, kit, device and storage medium
CN122128431A