Methylation marker site, brain stem glioma typing and prognosis model
By screening specific methylation marker sites and combining them with machine learning models, this method uses cerebrospinal fluid samples to classify and predict the prognosis of brainstem gliomas, solving the problem of non-invasive molecular pathological diagnosis. It achieves highly accurate and specific classification and prognostic judgment, dynamically monitors tumor changes, and guides personalized treatment.
Patent Information
- Application Number
- CN202510551085.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Current technologies make it difficult to perform non-invasive or minimally invasive molecular pathological diagnosis of brainstem gliomas. Surgical sampling is difficult, and cerebrospinal fluid testing has problems with inaccurate subtyping, which affects the choice of treatment plan and prognosis.
By screening specific methylation marker sites and combining them with machine learning models, brainstem gliomas are classified and their prognosis is predicted using cerebrospinal fluid samples. A random forest model is constructed to classify and predict gliomas by combining methylation marker sites with gene mutation information.
It achieves highly accurate and specific brainstem glioma classification and prognosis assessment, enables dynamic monitoring of tumor changes, guides personalized treatment, and reduces surgical risks and classification errors.
Smart Images

Figure CN120060477B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of gene detection, and particularly relates to a methylation marker site, brainstem glioma typing and prognosis model. BACKGROUND
[0002] Brainstem gliomas (BSG) is one of the brain malignancies, with a median survival of only 10 months. As a tumor with high heterogeneity, molecular pathological typing has gradually become an essential basis for guiding the precise diagnosis and treatment of BSG. According to the main molecular markers, BSG can be divided into three major subtypes (H3K27M, IDH and Double-negative). In the H3K27M subtype tumor, the genes encoding histone (HIST1H3B / C, HIST2H3C, H3F3A) are mutated, which leads to the inability of PRC2 protein to methylate histone H3K27, and further leads to the decrease of the overall methylation level of the cell; due to the mutation of IDH1 / 2 gene, the accumulation of 2-hydroxyglutarate (2-HG) in the IDH subtype tumor, which further affects the activity of DNA demethylase TET protein, leading to the increase of the overall methylation level of the cell; the Double-negative subtype lacks obvious biomarkers, and the prognosis of this subtype is significantly better than that of other subtypes.
[0003] Currently, the main way to obtain tumor tissue for molecular pathological diagnosis is surgery or stereotactic biopsy, but these two methods have many shortcomings. First, both are invasive procedures with high risk; second, patients must undergo invasive procedures to obtain molecular pathological diagnosis results, which reduces the guiding significance of molecular pathological diagnosis in guiding patients to choose the correct treatment plan, including avoiding unnecessary invasive procedures; third, there is sampling bias, which makes it difficult to reflect the spatial heterogeneity within the tumor; fourth, it is difficult to repeat multiple times and cannot dynamically monitor the tumor. The brainstem is the core area of the central nervous system, covering key life centers such as breathing and heartbeat, and the risk of surgery is extremely high. A slight mistake during surgery can lead to serious complications or even death. In addition, brainstem glioma usually grows in an infiltrative manner, with a blurred boundary, further increasing the difficulty of surgical resection. Therefore, the brainstem glioma surgery has poor tissue accessibility, while the accessibility of cerebrospinal fluid is relatively high. In contrast, cerebrospinal fluid testing has shown significant advantages. Liquid biopsy technology does not require invasive procedures and provides a reliable basis for precision diagnosis and treatment. For brainstem glioma patients who cannot obtain tissue through surgery, cerebrospinal fluid testing is undoubtedly a better choice. Second, molecular pathological diagnosis also has the problem of inconsistency with clinical practice. For example, Double negative subtype patients are usually considered to have significantly better prognosis than other subtypes, but other clinical information indicates that some patients in this subtype have poor prognosis, thus reducing the actual guiding significance of the subtype. Therefore, there is an urgent need in the clinic to develop new diagnostic methods with high accuracy, non-invasive or minimally invasive, and simple. SUMMARY
[0004] To solve the above problems, the present application provides a methylation marker site, brainstem glioma typing and prognosis model.
[0005] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:
[0006] The application provides a methylation marker site for brain stem glioma typing and prognosis stratification, including: taking hg19 as a reference genome, the methylation marker site includes: chr6:36355479-36355532, chr17:74497236-74497334, chr17:79480435-79480471, chr1:13910793-13910796, chr5:132082727-132082730, chr8:82193629-82193706, chr10:72200923-72200938, chr10:130339528-130339595, chr12:123380366-123380410, chr2:171568407-171568410, chr3:71631220-71631223, chr3:101568920-101568923, chr3:139258596-139258642, chr5:176170373-176170376, chr9: 82188499-82188502, chr12:53614079-53614082, chr17:38334060-38334063, chr19:11450022-11450036, chr19:52207581-52207591, chr1:13910566-13910609, chr1:13910697-13910737, chr3:134032302-134032392, chr5:132082823-132082826, chr1:11752202-11752209, chr1:155043729-155043732, chr4:54965828-54965831, chr11:72463414-72463417, chr12:52445120-52445197, chr19:52207340-52207343, chr9:109623014-109623066, chr19:19281254-19281272, chr1:210466206-210466209, chr1:228783348-228783351, chr10:130339689-130339692, chr10:130339691-130339694, chr12:49740749-49740752, chr17:74497720-74497723, chr2:201983198-201983201,chr2:201983331-201983334, chr3:197183565-197183568, chrll:19735659-19735662, chrll:72387963-72387966, chrll:463092-72463095, chr19:19281042-19281045, chr 1:220921501-220921504, chr4:108745592-108745595, chr20:10652811-10652814.
[0007] The embodiment of the present application also provides a screening method of the methylation marker site, comprising the following steps: collecting brainstem glioma related methylation chip data through a public database; performing standardization processing on the methylation chip data to obtain a beta matrix; performing pairwise comparison on H3K27M, IDH and Double-negative of BSG to obtain differential methylation sites; screening the methylation sites with large differences by setting the condition of Δβ≥0.3 and SD≤0.1; screening important CpGs by using a plurality of feature selection algorithms, and finally obtaining the methylation marker site.
[0008] Further, the plurality of feature selection algorithms include Random Forest, Support Vector Machine and Lasso, the CpGs reserved by at least two algorithms are screened to obtain the important CpGs.
[0009] The embodiment of the present application also provides a methylation model construction method for non-diagnostic and therapeutic purposes, comprising the following steps: sequencing the methylation marker site for a known sample to obtain sequencing data; performing outlier processing on the sequencing data; training a model by using the processed sequencing data, the model is a random forest model, the random forest model outputs prediction probabilities of three subtypes of BSG; verifying the trained random forest model by using a test sample.
[0010] The embodiment of the present application also provides a methylation model, which is constructed by using the construction method.
[0011] The embodiment of the present application also provides a method for classifying brainstem glioma according to the methylation model and gene mutation, the method is used for non-diagnostic and therapeutic purposes, and comprises the following steps: collecting cerebrospinal fluid of a patient; performing gene mutation detection and methylation level detection on the cerebrospinal fluid; calculating prediction probabilities of three subtypes of BSG by using a methylation model based on methylation detection data; and determining the classification of the patient based on the prediction probabilities and mutation information of the gene.
[0012] Further, the genes include IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes.
[0013] Further, the typing of the patient is determined based on the gene mutation information and the methylation information, and the typing includes:
[0014] The typing of the sample is performed by using the IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes, and the mutant type is marked as 1 and the wild type is marked as 0; the prediction probability of three subtypes is obtained based on the methylation model, and the value range is between 0 and 1; the mutation result and the prediction probability are added to obtain the score of the three subtypes, and the subtype with the maximum score is the final subtype of the sample.
[0015] The embodiment of the present application also provides a computer storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to realize the above method.
[0016] The embodiment of the present application also provides a BSG patient prognosis stratification model, the prediction probability of the H3K27M subtype is calculated based on the above methylation model and the cerebrospinal fluid; the survival curve is drawn based on the survival package, and the prediction probability value when the Youden index is maximum is selected as a threshold value for stratification; when the prediction probability is greater than the threshold value, the sample is high risk; and when the prediction probability is not greater than the threshold value, the sample is low risk.
[0017] The technical scheme provided by the embodiment of the present application has the beneficial effects including:
[0018] The methylation marker site provided by the embodiment of the present application is closely related to the typing and prognosis of brain stem glioma. First, the abnormal methylation of DNA is highly related to the occurrence and development of tumors, and there are obvious differences in the DNA methylation characteristics of central nervous system tumors, and the typing based on the DNA methylation characteristics can better reflect the clinical actual situation, and the research finds that the DNA methylation related to BSG shows that the overall methylation characteristics of different subtypes are obviously different, and have good discrimination. But at the same time there are corresponding problems: how to select the corresponding methylation marker site is a technical problem to be solved, the applicant selects the corresponding methylation marker site by using a specific selection method, which can well predict the typing and prognosis of brain glioma, and is more in line with the clinical actual situation; finally, the methylation markers screened by the present application have high specificity in different subtypes of BSG, and the machine learning model trained by using the methylation markers combined with mutation information has very high accuracy and specificity, and the model constructed based on the above methylation markers also has high accuracy and specificity for cerebrospinal fluid samples, which provides the possibility for subsequent tumor dynamic monitoring, and the prediction result of methylation data has guiding significance for patient prognosis stratification. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 The methylation of the methylation marker provided by the embodiment 1 of the present application in the BSG sample;
[0021] Figure 2 The schematic diagram of sample typing determined by the gene mutation and methylation model provided by the embodiment 2 of the present application;
[0022] Figure 3 The survival curve of the methylation model provided by the embodiment 3 of the present application for the H3K27M subtype. DETAILED DESCRIPTION
[0023] The present application will be further described in detail by specific embodiments. However, those skilled in the art will understand that the following embodiments are only used to illustrate the present application, and should not be regarded as limiting the scope of the present application. If the specific technology or condition is not specified in the embodiment, it is carried out according to the technology or condition described in the literature in the art or according to the product instruction. If the reagent or instrument is not specified by the manufacturer, it is a conventional product that can be obtained by market purchase.
[0024] The terms "comprise", "contain", "include", "have" or any other similar forms are intended to cover non-exclusive inclusions. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include a plurality of discussion objects.
[0025] Unless otherwise defined, all technical and scientific terms used in the present application have the same meanings as commonly understood by one of ordinary skill in the art. In addition to the specific methods, devices, materials used in the examples, any methods, devices and materials similar or equivalent to those described in the examples of the present application can be used according to the knowledge of the prior art and the description of the present application to implement the present application.
[0026] Unless otherwise specified, the experimental methods, detection methods, and preparation methods disclosed in the present application all adopt conventional molecular biology, immunology, testing, gene sequencing technology, bioinformatics technology, and related conventional techniques in the field.
[0027] The embodiments of the present application provide a methylation marker site for brain stem glioma typing and prognosis stratification, which is closely related to brain stem glioma typing and prognosis. First, the abnormal methylation of DNA is highly related to the occurrence and development of tumors, and there are obvious differences in the DNA methylation characteristics of central nervous system tumors, and the typing based on the DNA methylation characteristics can better reflect the clinical actual situation. Research has found that DNA methylation related to BSG shows obvious overall methylation characteristics of different subtypes, with good discrimination. However, there are corresponding problems: how to select the corresponding methylation marker site is a technical problem to be solved, the applicant selects the corresponding methylation marker site by using a specific selection method, which can well predict brain glioma typing and prognosis, and is more in line with the clinical actual situation; finally, the methylation markers screened in the present application have high specificity in different subtypes of BSG, and the machine learning model trained by using the methylation markers combined with mutation information has very high accuracy and specificity, and the model constructed based on the above methylation markers also has high accuracy and specificity for cerebrospinal fluid samples, which provides the possibility for subsequent tumor dynamic monitoring, and the prediction result of the methylation data has guiding significance for patient prognosis stratification.
[0028] The screening method of the methylation marker site comprises the following steps:
[0029] S1, collecting brain stem glioma related methylation chip data through a public database.
[0030] 450K / 850K methylation chip data related to brainstem glioma tissue samples were collected through public databases, including GSE90496, GSE109379, GSE161944, GSE50022, GSE64509, EGAS00001004341.
[0031] S2, standardizing the methylation chip data to obtain a beta matrix.
[0032] The raw methylation chip data is processed using the minfi (v1.50.0) package in R language (v4.3.2) to obtain the standardized beta matrix.
[0033] S3, pairwise comparison is performed for H3K27M, IDH and Double-negative of BSG to obtain differential methylation sites.
[0034] S4, methylation sites with large differences are screened by setting the condition of Δβ≥0.3 and SD≤0.1.
[0035] S5, a plurality of feature selection algorithms are used to screen important CpGs to finally obtain the methylation marker site.
[0036] The plurality of feature selection algorithms include Random Forest (RF), Support Vector Machine (SVM) and Lasso, and the CpGs retained by at least two algorithms are screened to obtain the important CpGs. In the above algorithms, SVM and Lasso fit the CpGs and retain the CpGs with non-zero coefficients, and RF retains the CpGs according to the importance.
[0037] The methylation marker site information for brainstem glioma typing and prognosis stratification is shown in Table 1 (with hg19 as the reference genome)
[0038] Table 1 Methylation site information
[0039]
[0040] The embodiment of the application also discloses a methylation model construction method for non-diagnostic and therapeutic purposes, comprising the following steps:
[0041] S10, for known samples, sequencing the above methylation marker sites to obtain sequencing data.
[0042] The methylation detection and data analysis process is as follows:
[0043] (1) Sample DNA extraction;
[0044] Tumor tissue samples were used for DNA extraction and purification using a nucleic acid extraction and purification kit (QIAamp DNA Mini Kit 250, QIAGEN). The concentration of the extracted DNA was detected using a DNA nucleic acid quantification kit (Qubit dsDNA HS Assay Kits, ThermoFisher Scientific). Samples with a total amount of DNA greater than 10 ng were subjected to subsequent library construction procedures, and samples with a total amount of DNA less than 10 ng were re-extracted for sample DNA.
[0045] (2) Library construction and sequencing;
[0046] The library was constructed using the method of patent ZL201910983038.8, and the process is as follows:
[0047] S10, first use methylation-sensitive restriction enzymes (HhaI and HinpI) to cut DNA, then fill in A and linker to form a pre-library, then amplify and purify the pre-library, take 400 ng of pre-library and use the designed primers to amplify the target region, then form a methylation detection library according to the method of patent ZL201910983038.8. Finally, according to the on-machine instructions of the sequencer, use Illumina or Huada sequencer for on-machine sequencing.
[0048] Table 2 Methylation site primer information
[0049]
[0050] Wherein, A represents the primer of the positive strand, B represents the primer of the negative strand, 1A and 2A, 1B and 2B are respectively the primers of two different regions corresponding to the site.
[0051] S20, the sequencing data is subjected to abnormal value processing.
[0052] Abnormal value processing: the target region without read coverage is defined as a missing value, and samples with a missing value proportion higher than 10% are removed; for part of the missing value, the average read value of the site in the sample is used to fill in.
[0053] S30, the processed sequencing data is used for model training, the model is a random forest model, and the random forest model outputs the prediction probability of the three subtypes of BSG.
[0054] The random forest model in caret (v4.3.3) in R language (v4.3.2) is used: model=train (data=data, method='rf', trControl, metric='kappa'), wherein data is training data, method defines the method adopted by the model, trControl defines the way of cross-validation, 10-fold cross-validation is adopted, metric defines the parameter for selecting the optimal model, and kappa is used as the parameter for optimizing the classification model.
[0055] S40, the trained random forest model is verified by using the test sample.
[0056] The model is tested by using the samples of the test set: predict (object, testdata, type = 'prob'), wherein object is the trained model, testdata is the test set data, and type is the output format of the predicted result. Finally, the output result is the prediction probability of the three subtypes of BSG, and the subtype with the maximum prediction probability is used as the final sample subtype prediction.
[0057] The embodiment of the application also provides a methylation model, which is constructed by using the construction method.
[0058] The embodiment of the application also provides a method for classifying brainstem glioma according to the methylation model and gene mutation, which is used for non-diagnostic and therapeutic purposes, and the method comprises the following steps:
[0059] S100, collecting cerebrospinal fluid of a patient. The cerebrospinal fluid sample is subjected to cfDNA extraction by using Apostle MiniMax high efficiency cfDNA isolation kit (Apostle; Pleasanton, CA, USA).
[0060] S200, detecting gene mutation and methylation level of the cerebrospinal fluid.
[0061] The genes include IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes.
[0062] The methylation level detection region includes methylation marker sites.
[0063] The library is constructed by the method of ZL201910983038.8 patent, and the process is as follows: first, the DNA is digested by methylation-sensitive restriction enzymes (HhaI and HinpI), then the pre-library is formed by adding A and adapter, then the pre-library is amplified and purified, 1 μg of pre-library is used to capture the target region using the primary brain tumor 68 gene detection product (Beijing Fenshengzi Medical Laboratory Co., Ltd.), and the mutation detection library is formed, 400 ng of pre-library is used to amplify the target region using the designed primer, and then the methylation detection library is formed according to the method of ZL201910983038.8 patent. Finally, according to the on-machine instruction of the sequencer, Illumina or Huada sequencer is used for on-machine sequencing.
[0064] S300, based on the gene sequencing data, the gene mutation detection result is obtained, and the mutant type is marked as 1 and the wild type is marked as 0.
[0065] S400, based on the methylation sequencing data, the prediction probability of the three subtypes of BSG is calculated by the methylation model, and the value range is between 0 and 1.
[0066] S500, based on the gene mutation information and the methylation model prediction probability, the patient's typing is determined. The gene mutation detection result and the methylation model prediction probability are added to obtain the scores of the three subtypes, and the subtype with the maximum score is the final subtype of the sample.
[0067] It should be pointed out that although molecular pathology is the gold standard for diagnostic typing, it has certain limitations. As described in the background art, molecular pathology needs to be analyzed after invasive sampling, but for brain stem glioma, invasive sampling causes great damage to the base, so the use of ctDNA in cerebrospinal fluid is considered. Because of the obstruction of the blood-brain barrier, the content of ctDNA (circulating tumor DNA) in the blood is limited, and the content of ctDNA in the cerebrospinal fluid, which fills in each brain ventricle, subarachnoid space and central canal of the spinal cord, is relatively high, and it can better represent the true information of the tumor. Therefore, liquid biopsy of cerebrospinal fluid will help to overcome the clinical diagnostic difficulties of BSG and achieve dynamic monitoring. In addition, single gene variation detection is limited by single point and qualitative marker detection, and in the case of low tumor content or limited tumor release, it is easy to miss the phenomenon from the probability; and the characteristic prediction of subtypes by methylation multi-site and quantitative markers can make up for the limitations of single point detection of gene variation through joint detection of multiple sites. Therefore, the combination of methylation multi-site and gene mutation has more practical guiding significance.
[0068] The application further provides a computer storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method described above.
[0069] The application further provides a BSG patient prognosis model, which is based on ctDNA in cerebrospinal fluid, calculates a prediction probability of an H3K27M subtype through the methylation model described above, draws a survival curve based on a survival package, selects a prediction probability value when a Youden index is maximum as a threshold value for stratification, and when the prediction probability is greater than the threshold value, the sample is high-risk, and when the prediction probability is not greater than the threshold value, the sample is low-risk.
[0070] Embodiment 1
[0071] Machine learning model of methylation markers.
[0072] The inventors collected 67 tissue samples of BSG patients, all the enrolled patients signed informed consent forms, and the patients were divided into a training set and a test set according to a certain proportion, the training set was used for construction of the machine learning model described below, and the test set was used for performance test of the model, and the sample amounts of the subtypes in each data set are shown in the following table.
[0073] Table 3 Sample allocation
[0074]
[0075] The methylation detection and data analysis process is as follows:
[0076] (1) Sample DNA extraction;
[0077] The tumor tissue sample is subjected to DNA extraction and purification using a nucleic acid extraction and purification kit (QIAamp DNA Mini Kit 250, QIAGEN). The extracted DNA is subjected to concentration detection using a DNA nucleic acid quantification kit (Qubit dsDNA HS Assay Kits, ThermoFisher Scientific). The samples with a total amount of DNA greater than 10 ng are subjected to subsequent library construction process, and the samples with a total amount of DNA less than 10 ng are subjected to re-extraction of sample DNA.
[0078] (2) Library construction and sequencing;
[0079] The library is constructed by the method of ZL201910983038.8 patent, and the general process is as follows: first, the DNA is digested by methylation-sensitive restriction endonuclease (HhaI and HinpI), then the pre-library is formed by adding A and adapter, then the pre-library is amplified and purified, 400 ng of pre-library is used to amplify the target region using the designed primer, and then the methylation detection library is formed according to the method of ZL201910983038.8 patent. Finally, according to the instrument operation instruction of the sequencer, Illumina or Huada sequencer is used for on-machine sequencing.
[0080] (3) Data analysis after machine;
[0081] After quality control, the data is compared with the human reference genome hg19. According to the method of ZL201910983038.8 patent, the target region methylation information is obtained, and the distribution of target site methylation in different samples is shown in Figure 1 .
[0082] (4) Methylation model construction and verification;
[0083] According to the target region methylation information, the methylation model is constructed. The specific steps are as follows:
[0084] 1. Abnormal value processing: the target region without reads coverage is defined as missing value. Remove samples with missing value proportion higher than 10%; for part of the missing value, use the average reads value in the sample to fill in;
[0085] 2. Methylation model training: use the random forest model in caret (v4.3.3) in R language (v4.3.2): model = train (data = data, method = ’rf’, trControl, metric = ’kappa’), wherein data is training data, method defines the method adopted by the model, trControl defines the way of cross-validation, 10-fold cross-validation is adopted, metric defines the selection of optimal model parameters, and kappa is used as the optimization parameter of classification model.
[0086] 3. Methylation model test: use the predict function to test the model with the samples of the test set: predict (object, testdata, type = ’prob’), wherein object is the trained model, testdata is the test set data, and type is the output format of the prediction result. Finally, the largest subtype of the three prediction probabilities is used as the final sample subtype prediction.
[0087] 4. Methylation model performance index statistics: the accuracy, specificity and AUC of the model on the training set and the test set are respectively counted, the sensitivity, specificity and AUC of the methylation model on the training set are shown in the following table, the sensitivity and specificity of the model on the training set for the three subtypes are all 1, and the AUC is 1.
[0088] Table 4 Performance of methylation model on training set
[0089]
[0090] The accuracy, specificity and AUC of the methylation model on the test set are shown in the following table, the sensitivity of the model on the test set for IDH subtype is 0.5, and the sensitivity for the other two subtypes is 1, combined with the detection result of IDH1 gene mutation (VAF is 0.047) indicating that the tumor cell content in the sample is less, and usually the minimum detection limit of methylation level is 5%, which ultimately leads to sample misclassification.
[0091] Table 5 Performance of methylation model on test set
[0092]
[0093] Example 2
[0094] Detection of BSG cerebrospinal fluid samples.
[0095] The tissue samples and cerebrospinal fluid samples of 29 BSG patients were collected respectively, and the cfDNA extraction was performed on the cerebrospinal fluid samples using Apostle MiniMax high efficiency cfDNA isolation kit (Apostle; Pleasanton, CA, USA). The tissue samples were subjected to molecular typing using the Panomics primary brain tumor 68 gene detection product as a reference standard.
[0096] (1) The methylation marker data of cerebrospinal fluid was used for sample typing.
[0097] In order to further verify the detection performance of the above markers and methylation model on cerebrospinal fluid samples, the methylation detection and data processing of 29 cerebrospinal fluid samples were also carried out according to the method of Example 1. The specific sample information is shown in the following table.
[0098] Table 6 Number of cerebrospinal fluid samples
[0099]
[0100] Sample typing prediction was performed using separate methylation marker data. The methylation typing model trained in Example 1 above was used to predict the typing of the above cerebrospinal fluid samples, and the prediction results are shown in the table below. The sensitivity and specificity of the H3K27M subtype were 0.783 and 0.833, respectively, the sensitivity and specificity of the IDH subtype were 0.000 and 1.000, respectively, and the sensitivity and specificity of the Double negative subtype were 0.750 and 0.720, respectively. Given the large difference between cerebrospinal fluid and tissue samples, firstly, the content of ctDNA in cerebrospinal fluid is small, and secondly, not all methylation markers will enter the cerebrospinal fluid, so that individual sample classification errors occur.
[0101] Table 7 Methylation model detection performance of cerebrospinal fluid samples
[0102]
[0103] (2) Using gene mutation information of cerebrospinal fluid for sample typing.
[0104] The same samples in (1) were used for gene mutation detection and data processing.
[0105] The library construction was performed using the method of ZL201910983038.8 patent, and the process is as follows: first, the DNA was digested using methylation-sensitive restriction endonuclease (HhaI and HinpI), then A-tailing and adapter were added to form a pre-library, then the pre-library was amplified and purified, 1 μg of pre-library was used to capture the target region using the primary brain tumor 68 gene detection product (Beijing Fenshengzi Medical Laboratory Co., Ltd.), and a mutation detection library was formed, and finally according to the instrument operation instruction, Illumina or Huada sequencer was used for instrument sequencing.
[0106] The samples were typed according to the mutation results of IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes, wherein the samples with IDH1 / 2 gene mutation were IDH subtype, the samples with any one of HIST1H3B / C, HIST2H3C and H3F3A gene mutation were H3K27M subtype, and the samples without the above gene mutation were Double negative subtype. The specific typing results are shown in the table below, wherein the sensitivity and specificity of the H3K27M subtype were 0.913 and 1.000, respectively, the sensitivity and specificity of the IDH subtype were 1.000 and 1.000, respectively, and the sensitivity and specificity of the Double negative subtype were 1.000 and 0.920, respectively.
[0107] Table 8 Mutation detection performance of cerebrospinal fluid samples
[0108]
[0109] (3) Using gene mutation and methylation marker data for sample typing.
[0110] First, the samples were typed using the IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes, with the mutant type marked as 1 and the wild type as 0; second, the methylation data were predicted using the methylation model trained in Example 2, and the prediction probabilities of the three subtypes were obtained, with the value ranging from 0 to 1; finally, the mutation results and the methylation model prediction results were added to obtain the scores of the three subtypes, and the subtype with the highest score was the final subtype of the sample. The specific steps are shown in Figure 2 The sensitivity and specificity of the H3K27M subtype were 0.956 and 0.833, respectively, the sensitivity and specificity of the IDH subtype were 1, and the sensitivity and specificity of the Double negative subtype were 0.750 and 0.960, respectively. The combination of mutation and methylation markers significantly improved the accuracy of BSG cerebrospinal fluid typing, although one case of Double negative subtype patient was misclassified, but other clinical information showed that the patient was a GBM (glioblastoma) subtype with poor prognosis, with a survival time of only 250 days, which was more similar to the clinical manifestations of the H3K27M subtype.
[0111] Table 9 Performance of cerebrospinal fluid sample methylation model + mutation detection
[0112]
[0113] Example 3
[0114] Prognostic stratification of BSG patient cerebrospinal fluid.
[0115] The prognosis of BSG patients of different subtypes is quite different, among which the prognosis of the H3K27M subtype is the worst. In order to better realize the prognosis stratification of H3K27M patients, the preoperative or intraoperative cerebrospinal fluid methylation marker data of the patients in the H3K27M subtype in Example 2 were combined with the survival information of the patients for prognosis analysis. After the sample methylation data were processed, the methylation model trained in Example 1 was used to predict the samples, and the prediction probability of the H3K27M subtype of all samples, i.e. the risk score of H3K27M, was obtained. The survival curve was drawn using the survival package, and the risk score at the maximum youden index was selected as the threshold to stratify the samples, and the results are shown in Figure 3When the risk score > 0.44, the sample is high risk, and the survival time is 328 days (95% CI, 166-527 days), and vice versa, the survival time is 769 days (95% CI, 281-notreached), P value is 0.023. The results show that the prediction result of the methylation typing model can be used for prognosis stratification of patients with H3K27M subtype.
[0116] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A screening method for methylation marker loci for brainstem glioma typing and prognosis stratification for methylation detection in cerebrospinal fluid or tissue samples, characterized in that, The method comprises the following steps: collecting brainstem glioma-related methylation chip data through a public database; standardizing the methylation chip data to obtain a β matrix; performing pairwise comparison on H3K27M, IDH and Double-negative of BSG to obtain differential methylation sites; filtering out methylation sites with large differences by setting Δβ≥0.3 and SD≤0.1; screening important CpGs by using multiple feature selection algorithms to obtain the methylation marker sites; the multiple feature selection algorithms include Random Forest, Support Vector Machine and Lasso, and the important CpGs are obtained by screening CpGs commonly retained by at least two algorithms; the methylation marker sites include: chr6:36355479-36355532, chr17:74497236-74497334, chr17:79480435-79480471, chr1: 13910793-13910796, chr5:132082727-132082730, chr8:82193629-82193706, chr10:72200923-72200938, chr10:130339528-130339595, chr12:123380366-123380410, chr2:171568407-171568410, chr3:71631220-71631223, chr3:101568920-101568923, chr3:139258596-139258642, chr5:176170373-176170376, chr9:82188499-82188502, chr12:53614079-53614082, chr17:38334060-38334063, chr19:11450022-11450036, chr19:52207581-52207591, chr1:13910566-13910609, chr1:13910697-13910737, chr3:134032302-134032392, chr5:132082823-132082826, chr1:11752202-11752209, chr1:155043729-155043732, chr4:54965828-54965831, chr11:72463414-72463417, chr12:52445120-52445197, chr19:52207340-52207343, chr9:109623014-109623066, chr19:19281254-19281272, chr1:210466206-210466209, chr1:228783348-228783351, chr10:130339689-130339692, chr10:130339691-130339694, chr12:49740749-49740752, chr17:74497720-74497723, chr2:201983198-201983201, chr2:201983331-201983334, chr3:197183565-197183568, chr11:19735659-19735662,chrll:72387963-72387966, chrll:463092-72463095, chr19:19281042-19281045, chr 1 :220921501-220921504, chr4:108745592-108745595, chr20:10652811-10652814., 2. A method for constructing a methylation model for non-diagnostic and therapeutic purposes, characterized by, The method comprises the following steps: sequencing the methylation marker sites obtained by the screening method of claim 1 for known samples to obtain sequencing data; performing outlier processing on the sequencing data; training a model using the processed sequencing data, wherein the model is a random forest model, and the random forest model outputs prediction probabilities of three subtypes of BSG; verifying the trained random forest model using test samples. 3.A methylation model constructed by the construction method of claim 2. 4.A computer storage medium storing a computer program, wherein the computer program is executed by a processor to implement a method for classifying brainstem glioma based on the methylation model and gene mutation of claim 3, comprising: collecting cerebrospinal fluid of a patient; detecting gene mutation and methylation level of the cerebrospinal fluid; calculating prediction probabilities of three subtypes of BSG based on methylation detection data by using a methylation model; determining the classification of the patient based on the prediction probabilities and the gene mutation information. 5.The computer storage medium of claim 4, wherein the genes include IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes. 6.The computer storage medium of claim 4, wherein determining the classification of the patient based on the gene mutation information and the methylation information comprises: using IDH1 / 2, HIST1H3B / C, HIST2H3C and H3F3A genes to classify samples, and marking mutations as 1 and wild types as 0; obtaining prediction probabilities of three subtypes based on the methylation model, and the values range from 0 to 1; adding the mutation results and the prediction probabilities to obtain scores of three subtypes, and the subtype with the maximum score is the final subtype of the sample. 7.A BSG patient prognosis model for non-diagnostic and therapeutic purposes, wherein the prediction probability of the H3K27M subtype is calculated based on the cerebrospinal fluid by using the methylation model of claim 3; the prediction probability value at which the Youden index is maximum is selected as a threshold for stratification based on the survival package to draw a survival curve. when the predicted probability is greater than the threshold, then the sample is high risk; when the predicted probability is not greater than the threshold, then the sample is low risk.
Citation Information
Patent Citations
Method for detecting variation and methylation of tumor specific genes in ctDNA
CN112176419A