A method, device and program product for screening methylation CpG sites that constrain multi-omics
Through the methylated CPG site screening method integrating multiomics data, combined with regression models and clinical characteristics, the limitations of methylated CpG site screening in the prior art were solved, and higher accuracy site screening and glioma prediction were achieved, especially glioma typing related to H3 K27 mutations.
Patent Information
- Application Number
- CN202411906130.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The existing methylated CpG site screening methods have limitations in cost, coverage and accuracy, and it is difficult to fully understand the mechanisms of gene expression regulation and disease occurrence and development, and lack effective multi-dimensional screening methods.
The methylated CPG site screening method that constrains multiomics is used to integrate transcription data, CNV data and methylated CPG probes, and data integration and screening are used for regression models, combining clinical characteristics and mutation characteristics for glioma prediction. The optimal regularization parameters are found using Lasso regression, ridge regression and other models to screen out key methylated CPG sites.
It improves the accuracy of methylated CPG site screening and the reliability of glioma prediction, and can early warning and intervention to diagnose gliomas of different subtypes, providing a more comprehensive understanding of gene regulation networks.
Smart Images

Figure CN119741966B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent medicine, and particularly relates to a method, device, program product and computer-readable storage medium for screening methylation CpG sites that constrain multiple omics. Background Art
[0002] The screening of methylation CpG sites is an important link in epigenetics research, involving a variety of technologies and methods. Each method has its unique advantages and limitations. For example, through whole-genome bisulfite sequencing (WGBS), the methylation status of each CpG site in the entire genome is detected, providing the most comprehensive data. However, WGBS requires a high sequencing depth, resulting in very high experimental costs and is not suitable for the study of large-scale samples; subsequently, the bisulfite array was proposed, which can detect thousands of specific CpG sites simultaneously and is suitable for the study of large-scale samples. However, it depends on the design of commercial arrays, may miss some important methylation sites, and has high costs. Methylation-specific PCR (MSP) and reduced representation bisulfite sequencing (RRBS) have lower costs, but PCR requires the design of specific primers, has high requirements for primer design, and can only detect a few CpG sites at a time. RRBS has a limited coverage range and mainly detects high CpG density regions, and has weak detection ability for low CpG density regions. Each method has its applicable scenarios and limitations, but none of them have ever considered methylation site screening from multiple dimensions. Summary of the Invention
[0003] The screening of methylated CpG sites by integrating multiple omics is a method that comprehensively combines multiple data types, aiming to comprehensively understand the mechanisms of gene expression regulation and disease occurrence and development from multiple perspectives. Multi-omics data usually includes genomics, transcriptomics, proteomics, metabolomics, etc. By integrating these data, the functions and biological significance of methylated CpG sites can be more accurately identified and verified. Therefore, in view of the above problems, the present invention provides a method for screening methylated CPG sites by constraining multiple omics, specifically including: S1, obtaining gene expression data of a sample to be tested, including transcriptional data, CNV data, and methylated CPG probes; S2, performing differential calculation on the transcriptional data to obtain differentially expressed genes; S3, performing correlation calculation on the transcriptional data to obtain the correlation of transcription factors; S4, when the correlation of the transcription factor is negatively correlated, assigning a negative value to the negatively correlated transcription factor to obtain a transformed transcription factor; S5, assigning a negative value to the methylated CPG probe to obtain a transformed methylated CPG probe; S6, performing regression calculation on the transformed methylated CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed genes to obtain the regression methylated CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed genes; S7, performing gene mutation analysis based on the regression methylated CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed genes to obtain methylated CpG sites.
[0004] The regression calculation is performed through a regression model, including one or more of the following: Lasso regression, ridge regression, Bayesian linear regression, elastic net regression;
[0005] Optionally, the regression calculation obtains the optimal regularization parameter through ridge regression, sorts the regression coefficients based on the optimal regularization parameter to obtain the first-ranked regression coefficient, and the corresponding methylated CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed genes ranked first.
[0006] Optionally, the optimal regularization parameter is obtained by first generating a value range of the regularization parameter and optimizing the regularization parameter within the value range.
[0007] Optionally, the parameter optimization is performed through one or more of the following: grid search method, random search method, particle swarm optimization algorithm, niche algorithm, simulated annealing algorithm, genetic algorithm, ant colony optimization algorithm.
[0008] The correlation calculation of the transcription factor obtains a correlation ranking, and the top N transcription factors are retained, where N is a natural number greater than 2.
[0009] Optionally, the correlation calculation is performed using the Pearson correlation coefficient to obtain a correlation ranking.
[0010] Optionally, the methylated CPG probes are screened to obtain screened methylated CPG probes, and the screened methylated CPG probes are subjected to a negative value conversion to obtain converted methylated CPG probes; wherein the screening is to screen probes with methylation higher than a preset threshold or methylation lower than a preset threshold.
[0011] Optionally, the methylation screening further includes similarity screening. The screened methylated CPG probes are subjected to similarity calculation to obtain similar methylated CPG probes, and the similar methylated CPG probes are subjected to a negative value conversion to obtain converted methylated CPG probes; optionally, the similar methylated CPG probes are the top L similar methylated CPG probes sorted according to similarity, and L is a natural number greater than 2.
[0012] Optionally, the negative value is multiplied by negative one.
[0013] Optionally, the method further includes negative value assignment for differentially expressed genes, calculating the correlation of differentially expressed genes, and performing negative value conversion on negatively correlated differentially expressed genes to obtain converted differentially expressed genes; performing regression calculation on the converted methylated CPG probes, transcription factors, converted transcription factors, CNV data, differentially expressed genes, and converted differentially expressed genes to obtain regression methylated CPG probes, transcription factors, converted transcription factors, CNV data, differentially expressed genes, and converted differentially expressed genes.
[0014] Optionally, the method further includes multi-scale conversion. The converted methylated CPG probes, transcription factors, converted transcription factors, CNV data, and differentially expressed genes are normalized to obtain normalized methylated CPG probes, normalized transcription factors, normalized converted transcription factors, normalized CNV data, and normalized differentially expressed genes, and regression calculation is performed on the normalized methylated CPG probes, normalized transcription factors, normalized converted transcription factors, normalized CNV data, and normalized differentially expressed genes to obtain regression methylated CPG probes, transcription factors, converted transcription factors, CNV data, and differentially expressed genes.
[0015] The method further includes data preprocessing, performing normalization processing on gene expression data to obtain processed gene expression data, and performing the processing of S2-S7 based on the processed gene expression data.
[0016] Optionally, the CNV data includes data in the SEG MEAN format.
[0017] Optionally, the methylated CPG probes include beta values.
[0018] The object of the present invention is to provide a glioma prediction method based on H3 K27 mutation, including: obtaining gene data of a sample to be tested; extracting methylation site features of the gene data; the methylation site features are obtained according to the above-mentioned methylation CPG site screening method for constrained multi-omics; predicting based on the methylation site features to obtain a prediction result that the sample to be tested is pontine-like or thalamic-like.
[0019] The pontine-like means that the tumor location is within the pons, and the thalamic-like means that the tumor location is outside the pons. The features further include mutation features. Extract the mutation features of the gene data, and predict based on the methylation site features and mutation features to obtain a prediction result that the sample to be tested is pontine-like or thalamic-like.
[0020] Optionally, the mutation features include one or more of the following: TP53 mutation, NF1 mutation.
[0021] Optionally, the features further include clinical features. Obtain clinical data, extract the clinical features in the clinical data, and predict based on the methylation site features and clinical features to obtain a prediction result that the sample to be tested is pontine-like or thalamic-like.
[0022] Optionally, the clinical features include one or more of the following: age, MKI67 IHC index.
[0023] Optionally, the features include clinical features and mutation features. Predict based on the methylation site features, mutation features, and clinical features to obtain a prediction result that the sample to be tested is pontine-like or thalamic-like.
[0024] Optionally, the method further includes data processing. Perform data dimensionality reduction on the methylation features to obtain reduced-dimensional data, then sort the reduced-dimensional data to obtain the top S reduced-dimensional methylation features, where S is a natural number greater than 3. Predict based on the reduced-dimensional methylation site features to obtain a prediction result that the sample to be tested is pontine-like or thalamic-like.
[0025] Optionally, perform feature selection based on the edge to screen the mutation features or clinical features to obtain the screening.
[0026] The prediction is performed through one or more of the following models: Naive Bayes, Support Vector Machine, Random Forest, Decision Tree, Convolutional Neural Network, Transformer.
[0027] Optionally, the prediction is performed by inputting the features into their respective Naive Bayes to obtain their respective prediction probabilities, and voting on the respective prediction probabilities to obtain the final prediction result.
[0028] The object of the present invention is to provide a computer program product, which includes a computer program or instruction, and the computer program or instruction is executed by a processor to implement the above-mentioned method for screening methylated CpG sites of constrained multi-omics, or to implement the above-mentioned method for predicting glioma based on H3 K27 mutation.
[0029] The object of the present invention is to provide a computer device, which includes a memory, a processor, and a computer program or instruction stored on the memory, and the computer program or instruction is executed by the processor to implement the above-mentioned method for screening methylated CpG sites of constrained multi-omics, or to implement the above-mentioned method for predicting glioma based on H3 K27 mutation.
[0030] The object of the present invention is to provide a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is executed by a processor to implement the above-mentioned method for screening methylated CpG sites of constrained multi-omics, or to implement the above-mentioned method for predicting glioma based on H3 K27 mutation.
[0031] Advantages of the present invention:
[0032] 1. The present invention proposes a method for screening methylated CpG sites of constrained multi-omics, which combines transcriptional data, gene data, copy number variation data, and methylation data. The methylated CpG sites screened by the feature screening of constrained multi-omics are more suitable for the working environment of the gene regulatory network than the single methylation data screening method, thereby improving the accuracy of site screening.
[0033] 2. In the methylated CpG screening sites of constrained multi-omics, the multi-omics data corresponding to the optimal regularization parameter is obtained by integrating the multi-omics data through a regression model, and then the methylated CpG sites are obtained by analyzing the multi-omics data screened based on this. The optimal regularization parameter is obtained by optimizing the parameter.
[0034] 3. Diffuse midline glioma (DMG) with H3 K27 alteration is a highly lethal tumor of the central nervous system, but there is a lack of effective treatment for glioma. Through research, it is found that only H3 K27 alteration is not sufficient to cause tumorigenesis, and tumor transformation requires at least one companion alteration. Animal models show that H3 K27-altered tumors with different companion alterations exhibit different growth rates, spread, and invasion phenotypes. Therefore, based on the data analysis of glioma with H3 K27 alteration, the present invention proposes new classifications, namely pontine-like and thalamic-like, and predicts different subtypes of glioma through the methylated CpG sites screened by constrained multi-omics data, which is helpful for early prediction, early warning, and intervention diagnosis and treatment of the disease.
[0035] 4. The typing prediction of gliomas also includes mutation characteristics and clinical features, and prediction is carried out from multiple dimensions to further improve the accuracy and reliability of prediction. Description of the Drawings
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 Schematic flowchart of a method for screening methylation CpG sites by constrained multi-omics provided by an embodiment of the present invention;
[0038] Figure 2 Schematic diagram of a system for screening methylation CpG sites by constrained multi-omics provided by an embodiment of the present invention;
[0039] Figure 3 Schematic diagram of a device for screening methylation CpG sites by constrained multi-omics provided by an embodiment of the present invention;
[0040] Figure 4 Flowchart of CM3R provided by an embodiment of the present invention;
[0041] Figure 5 Data analysis results of gliomas provided by an embodiment of the present invention, (B) Clinical information, data availability and somatic mutation status of H3-DSGs included in this study. C / T / L represent the cervical, thoracic and lumbar segments of the human spinal cord. Summarize the driver genes of diffuse midline gliomas identified by MutSigCV and the reported key mutations. (C) Show the cSNF network of two multi-omics subtypes of H3-DSGs. The TP53 and NF1 mutation statuses are represented by the edge colors. The edge width represents the strength of evidence of patient-patient similarity supported by multi-omics data. The lymph node size is positively correlated with the age of the patients at the time of diagnosis (7 - 61 years old). The node colors represent the H3-DSG subtypes defined by the cSNF method. (D) Mutation frequencies of key driver genes in different diseases. Brain-M-Pons and brain-M-Medulla are from Chen et al. (PMID: 32555164). (E) Diagnostic ages of patients in different disease groups. (F) Kaplan-Meier survival curves of tumor patients in different groups;
[0042] Figure 6Functional annotation of differentially expressed genes between pontine-like and thalamic-like H3-DSGs provided by embodiments of the present invention: (A) Volcano plot showing differentially expressed genes between two tumor subtypes in bulk RNA-seq; (B) Comparison of tumor purity calculated based on WES data between the two subtypes; (C and D) Comparison of lymphocyte infiltration and myeloid infiltration between the pontine-like and thalamic-like subgroups through RNA-seq data.
[0043] Figure 7 Proportion of cycling cells in (A) pontine-like and thalamic-like H3-DSGs provided by embodiments of the present invention; (B) Immunohistochemistry of Ki-67 (left panel) and Ki-67 index in pontine-like and thalamic-like H3-DSGs.
[0044] Figure 8 CNVs / LOH events calculated by comparing tumor WES data of two subgroups provided by embodiments of the present invention.
[0045] Figure 9 Expression of FANCI protein in tumors of two groups provided by embodiments of the present invention, scale bar: 20 μm. Detailed implementation manners
[0046] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0047] In some processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0048] Figure 1 Schematic diagram of a method for screening methylation CpG sites by constraining multi-omics provided by embodiments of the present invention, specifically including:
[0049] S1: Obtain gene expression data of a sample to be tested, including transcription data, CNV data, and methylation CpG probes.
[0050] In one embodiment, the method further includes data preprocessing, normalizing gene expression data to obtain processed gene expression data, and performing the processing of S2 - S7 based on the processed gene expression data.
[0051] In one embodiment, the CNV data includes data in SEG MEAN format.
[0052] In one embodiment, the methylation CPG probe includes beta values.
[0053] In a specific embodiment, a new method of CM3R (Constrained Multi - omics Integration Regression - based Regulatory Factor Ranking Algorithm), which can identify copy number variations, transcription factors, and DNA methylation CpG sites highly correlated with subtype - differential gene expression. The method includes three steps: (a) data preparation, (b) TF extraction and CpG probe pre - selection, and (c) multi - scale normalization and regression, as specifically shown in Figure 4 shown. In the data preparation of the CM3R method, gene expression data has to go through a series of normalization procedures before being used as input, including quantile normalization and min - max normalization. Meanwhile, CNV data in'segmean' format and the beta values of CpG probes are required.
[0054] S2: Calculate differential expression genes by performing differential calculation on transcription data;
[0055] In one embodiment, the method further includes assigning negative values to differential expression genes, calculating the correlation of differential expression genes, and performing negative - value conversion on negatively - correlated differential expression genes to obtain converted differential expression genes; positively - correlated differential expression genes remain unchanged. Perform regression calculation on the converted methylation CPG probes, transcription factors, converted transcription factors, CNV data, (positively - correlated) differential expression genes, and converted differential expression genes to obtain regression results of methylation CPG probes, transcription factors, converted transcription factors, CNV data, (positively - correlated) differential expression genes, and converted differential expression genes.
[0056] S3: Calculate the correlation of transcription factors by performing correlation calculation on transcription data;
[0057] In one embodiment, perform correlation calculation on the transcription factors to obtain a correlation ranking, and retain the top N transcription factors, where N is a natural number greater than 2.
[0058] In one embodiment, the correlation calculation is performed using the Pearson correlation coefficient to obtain a correlation ranking.
[0059] In a specific embodiment, in (b) transcription factor extraction and CpG probe pre-selection, the transcription factors (TFs) of each gene are first extracted from the RegNetwork database and further selected based on the Pearson correlation coefficient index. In the CM3R method, up to 3 transcription factors are considered for each gene.
[0060] S4: When the correlation of the transcription factor is negative, assign a negative value to the negatively correlated transcription factor to obtain a transformed transcription factor;
[0061] In a specific embodiment, two additional transformations are performed on the beta values of these selected CpG probes and the gene expression of the transcription factors. Multiply the beta values of all selected CpG probes by (-1). For the gene expression of the transcription factor showing a negative correlation with the current differentially expressed (DE) gene, multiply by (-1) is also applied to reverse their values.
[0062] S5: Assign a negative value to the methylated CpG probe to obtain a transformed methylated CpG probe;
[0063] In one embodiment, the methylated CpG probe is screened to obtain a screened methylated CpG probe, and the screened methylated CpG probe is subjected to a negative value conversion to obtain a transformed methylated CpG probe; wherein the screening is to screen the probes with methylation higher than a preset threshold or methylation lower than a preset threshold.
[0064] In one embodiment, the methylation screening further includes similarity screening. The similarity of the screened methylated CpG probe is calculated to obtain a similar methylated CpG probe, and the similar methylated CpG probe is subjected to a negative value conversion to obtain a transformed methylated CpG probe.
[0065] Optionally, the similar methylated CpG probes are the top L similar methylated CpG probes sorted according to the similarity, and L is a natural number greater than 2.
[0066] Optionally, the assignment of the negative value is to multiply by negative one.
[0067] In a specific embodiment, for the selection of methylated CpG probes, first select the CpG probes showing high methylation or low methylation patterns in the entire gene body and 2000b upstream of the transcription start site (TSS) region, and further adopt a selection procedure similar to that of the TF.
[0068] In a specific embodiment, a multi-scale normalization technique is adopted, and the sigmoid transformation is used to appropriately scale the gene expression data (including positively / negatively correlated transcription factors and differentially expressed genes), CNV data, and CpG beta values.
[0069] S6: Perform regression calculations on the converted methylated CpG probes, transcription factors, converted transcription factors, CNV data, and differentially expressed genes to obtain the regressed methylated CpG probes, transcription factors, converted transcription factors, CNV data, and differentially expressed genes.
[0070] In one embodiment, the regression calculation is performed through a regression model, including one or more of the following: Lasso regression, ridge regression, Bayesian linear regression, elastic net regression. In one embodiment, the regression calculation obtains the optimal regularization parameter through ridge regression, sorts the regression coefficients based on the optimal regularization parameter to obtain the first-ranked regression coefficient, and the methylated CpG probes, transcription factors, converted transcription factors, CNV data, and differentially expressed genes corresponding to the first rank.
[0071] In one embodiment, the optimal regularization parameter is to first generate the value range of the regularization parameter, and optimize the regularization parameter within the value range to obtain the optimal regularization parameter;
[0072] Optionally, the parameter optimization is performed through one or more of the following: grid search method, random search method, particle swarm optimization algorithm, niche algorithm, simulated annealing algorithm, genetic algorithm, ant colony optimization algorithm.
[0073] In one embodiment, the method further includes multi-scale conversion. Normalize the converted methylated CpG probes, transcription factors, converted transcription factors, CNV data, and differentially expressed genes to obtain normalized methylated CpG probes, normalized transcription factors, normalized converted transcription factors, normalized CNV data, and normalized differentially expressed genes. Perform regression calculations on the normalized methylated CpG probes, normalized transcription factors, normalized converted transcription factors, normalized CNV data, and normalized differentially expressed genes to obtain the regressed methylated CpG probes, transcription factors, converted transcription factors, CNV data, and differentially expressed genes.
[0074] In one embodiment, the method further includes multi-scale transformation, normalizing the transformed methylation CPG probes, transcription factors, transformed transcription factors, CNV data, differentially expressed genes, and transformed differentially expressed genes to obtain normalized methylation CPG probes, normalized transcription factors, normalized transformed transcription factors, normalized CNV data, normalized differentially expressed genes, and normalized transformed differentially expressed genes, performing regression calculation on the normalized methylation CPG probes, normalized transcription factors, normalized transformed transcription factors, normalized CNV data, normalized transformed differentially expressed genes, and normalized differentially expressed genes to obtain regression methylation CPG probes, transcription factors, transformed transcription factors, CNV data, transformed differentially expressed genes, and differentially expressed genes, and performing gene difference analysis based on the regression methylation CPG probes, transcription factors, transformed transcription factors, CNV data, transformed differentially expressed genes, and differentially expressed genes to obtain methylation CPG sites.
[0075] In a specific embodiment, a restricted ridge regression model is used to model the relationship between the gene expression of current differentially expressed (DE) genes and independent variables, including CNV data, selected CpG probes, and transcription factors, which forces the coefficients of the regulators to be non-negative.
[0076] In the regression stage of the CM3R method, a range of values for the regularization parameter (denoted as α) is systematically adopted. These values are varied to optimize the performance of the model and achieve the best fit for the given data. By exploring different values of α, the CM3R method can find the optimal regularization parameter that balances model complexity and prediction accuracy. Once the optimal α value is determined, the CM3R method assigns the regulatory factor type ranked first to the current DE gene according to the regression coefficient ranking, that is, determines the transcription factor (TF) ranked highest, the CpGs ranked highest, and the CNVs ranked highest related to the DE gene.
[0077] S7: Performing gene variation analysis based on the regression methylation CPG probes, transcription factors, transformed transcription factors, CNV data, and differentially expressed genes to obtain methylation CPG sites.
[0078] In one embodiment, gene variable analysis is performed by GSVA to obtain the analysis result.
[0079] The disclosed embodiments of the present invention also provide a computer program product or system, including a computer program, which when executed by a processor, implements the steps of the above-mentioned method for screening methylation CPG sites of constrained multi-omics.
[0080] Figure 2Schematic diagram of a system for screening methylation CpG sites by constraining multi-omics provided by an embodiment of the present invention, specifically including: Acquisition unit: acquiring gene expression data of a sample to be tested, including transcription data, CNV data, and methylation CpG probes;
[0081] Differential unit: performing differential calculation on the transcription data to obtain differentially expressed genes;
[0082] Transcription unit: performing correlation calculation on the transcription data to obtain the correlation of transcription factors;
[0083] Conversion unit: when the correlation of the transcription factor is negatively correlated, assigning a negative value to the negatively correlated transcription factor to obtain a converted transcription factor;
[0084] Methylation unit: assigning a negative value to the methylation CpG probe to obtain a converted methylation CpG probe;
[0085] Regression unit: performing regression calculation on the converted methylation CpG probe, transcription factor, converted transcription factor, CNV data, and differentially expressed genes to obtain the regression methylation CpG probe, transcription factor, converted transcription factor, CNV data, and differentially expressed genes;
[0086] Analysis unit: performing gene mutation analysis based on the regression methylation CpG probe, transcription factor, converted transcription factor, CNV data, and differentially expressed genes to obtain methylation CpG sites.
[0087] Figure 3 Schematic diagram of a device for screening methylation CpG sites by constraining multi-omics provided by an embodiment of the present invention, specifically including:
[0088] A memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, any one of the above-mentioned methods for screening methylation CpG sites by constraining multi-omics.
[0089] The disclosed embodiment of the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any one of the above-mentioned methods for screening methylation CpG sites by constraining multi-omics.
[0090] An embodiment of the present invention provides a method for predicting glioma based on H3 K27 mutation, including: acquiring gene data of a sample to be tested; extracting methylation site features of the gene data; the methylation site features are obtained by the above-mentioned method for screening methylation CpG sites by constraining multi-omics; predicting based on the methylation site features to obtain a prediction result that the sample to be tested is pontine-like or thalamic-like.
[0091] In one embodiment, the pontine-like tumor is located within the pons, and the thalamic-like tumor is located outside the pons.
[0092] In one embodiment, the features further include mutation features. The mutation features of gene data are extracted, and based on the methylation site features and mutation features, a prediction result of whether the sample to be tested is pontine-like or thalamic-like is obtained.
[0093] In one embodiment, the mutation features include one or more of the following: TP53 mutation, NF1 mutation.
[0094] In one embodiment, the features further include clinical features. Clinical data is obtained, the clinical features in the clinical data are extracted, and based on the methylation site features and clinical features, a prediction result of whether the sample to be tested is pontine-like or thalamic-like is obtained.
[0095] In one embodiment, the clinical features include one or more of the following: age, MKI67 IHC index.
[0096] In one embodiment, the features include clinical features and mutation features. Based on the methylation site features, mutation features, and clinical features, a prediction result of whether the sample to be tested is pontine-like or thalamic-like is obtained.
[0097] In one embodiment, the method further includes data processing. The methylation features are dimensionally reduced to obtain reduced-dimensional data, and then the reduced-dimensional data is sorted to obtain the top S reduced-dimensional methylation features, where S is a natural number greater than 3. Based on the reduced-dimensional methylation site features, a prediction result of whether the sample to be tested is pontine-like or thalamic-like is obtained.
[0098] In one embodiment, feature selection based on the margin is used to screen the mutation features or clinical features to obtain the screening results.
[0099] In one embodiment, the prediction is performed through one or more of the following models: Naive Bayes, Support Vector Machine, Random Forest, Decision Tree, Convolutional Neural Network, Transformer;
[0100] Optionally, the prediction is obtained by inputting the features into their respective Naive Bayes to obtain their respective prediction probabilities, and then voting on the respective prediction probabilities to obtain the final prediction result.
[0101] In one embodiment, five Naive Bayes are respectively constructed for the methylation features, mutation features (TP53 mutation, NF1 mutation), and clinical features (age, MKI67 IHC index) and trained in parallel to obtain the prediction models for the corresponding features.
[0102] In a specific embodiment, the present invention analyzed multi-omics and spatially resolved single-cell data from 36 H3-DSGs and identified two distinct subgroups: pontine-like and thalamic-like. Thalamic-like tumors were characterized by NF1 mutations and exhibited more immune infiltration. They contained a large number of pro-inflammatory myeloid cells and CD8+ cytotoxic T cells, with an older patient age and relatively better prognosis. In contrast, pontine-like tumors showed abundant TP53 mutations, high proliferation niches, genomic instability, an older patient age, and relatively poor prognosis.
[0103] To clarify the subtypes of H3 K27M-mutant spinal gliomas, multi-omics clustering and subtype identification were performed on 36 cases with DNA methylation, WES, and RNA sequencing data ( Figure 5 B). There were 22 males and 14 females; the age ranged from 7 to 61 years, with a median age of 33 years. The tumors were widely resected from the first cervical segment (C1) to the first lumbar segment (L1). Two cases were multifocal tumors.
[0104] In another specific embodiment, to better characterize the tumor features of H3-DSG subtype 1 and subtype 2, the present invention compared the somatic mutation spectra, age at diagnosis, and overall survival outcomes of the H3-DSG subgroups with those of previously reported brain DMG subgroups in an independent cohort consisting of m-pons and m-medulla oblongata DMG17. The m-pons DMG subgroup tumors were mainly located within the pons, while the m-medulla oblongata DMG subgroup tumors were mostly located outside the pons. Notably, the newly discovered H3-DSG subtype 1 exhibited somatic features similar to those of the m-pons brain DMG subgroup, while the somatic features of H3-DSG subtype 2 were similar to those of the m-medulla oblongata brain DMG subgroup, with the only exceptions being SF3B1 and FGFR1 mutations, and the incidence of SF3B1 and FGFR1 mutations being higher in m-medulla oblongata brain DMG than in H3-DSG subtype 2 ( Figure 5 D). The age distributions at diagnosis of the H3-DSG subtype 1 and M-Pons brain DMG groups were similar and were significantly lower than those of H3-DSG subtype 2 (P = 4.5e-05) and the m-medulla oblongata brain DMG group (P = 6.4e-06) respectively ( Figure 5 E). The survival rates of the H3-DSG 1 group and the m-pons DMG group were significantly lower than those of the H3-DSG 2 group (P = 8.7e-04) and the m-medulla oblongata brain DMG group (P = 1.4e-04) respectively ( Figure 5 F). These results indicate that H3-DSG subtype 1 and H3-DSG subtype 2 have similar molecular and clinical features to m-pons and m-medulla oblongata type DMGs respectively. Based on these findings, the present invention named H3-DSG subtype 1 pontine-like and H3-DSG subtype 2 thalamic-like respectively.
[0105] In a specific embodiment, by studying the differentially expressed genes between pontine-like and thalamic-like H3-DSGs ( Figure 6 A), we found that the genes significantly upregulated in pontine-like tumors were mainly related to cell division and cell cycle-related functions, such as MEX3A, KIF2C, FANCI, etc. ( Figure 6 A). Conversely, the genes upregulated in thalamic-like tumors included immune and inflammatory-related activities. Therefore, compared with pontine-like H3-DSGs, thalamic-like tumors had lower tumor purity (P = 0.006), higher lymphocyte (P = 0.0044) and myeloid scores (P = 0.0006) ( Figure 6 B-D). In fact, scRNA-seq data showed that pontine-like and thalamic-like H3-DSGs showed different cell type compositions and tumor cell states. Malignant cells accounted for 80% - 90% of pontine-like tumors, and the remaining components were mostly myeloid cells with few T cells. In contrast, in thalamic-like tumors, malignant cells accounted for only about 50%, and the remaining 50% included myeloid cells, T cells, NK cells, neutrophils, endothelial cells, and oligodendrocytes. Consistent with this, immunohistochemistry (IHC) of 28 H3-DSGs showed that there were more CD14-positive myeloid cells in thalamic-like H3-DSGs. Multicolor staining also showed that there were more myeloid and lymphoid cells in thalamic-like H3-DSGs. Correspondingly, CD14 (myeloid), CD3e (T cells), MPO (neutrophils), CD44, HLA-A, and HLA-DR (antigen presentation) were all more abundant in thalamic-like H3-DSGs than in pontine-like H3-DSGs.
[0106] In a specific embodiment, the results of single-cell analysis also showed that the proportion of CD8-positive cytotoxic T cells in T cells was significantly increased in thalamic-like H3-DSGs. The results of immunohistochemistry showed that the proportions of CD3e-positive cells and GZMB-positive cells in thalamic-like H3-DSGs were also significantly higher than those of the corresponding cells in pontine-like H3-DSGs. This suggested an increase in the proportion of cytotoxic T cells in thalamic-like H3-DSGs.
[0107] In a specific embodiment, single-cell analysis found that more cells in pontine-like H3-DSGs were in the cell proliferation cycle ( Figure 7 A), and the index of the cell proliferation marker protein Ki-67 was significantly increased in pontine-like H3-DSGs ( Figure 7 B). Subsequently, we found that the occurrence rates of copy number variations and loss of heterozygosity in pontine-like H3-DSGs were significantly higher than those in thalamic-like H3-DSGs ( Figure 8), suggesting that its genome is more unstable. And the expression level of the protein FANCI related to DNA damage repair is significantly increased in pontine-like H3-DSG ( Figure 9 ).
[0108] In a specific embodiment, the prediction model:
[0109] The machine learning model was constructed using the naïve Bayes algorithm in the naivebayes (v0.9.7) R package. The original feature pool consists of several parts: the top 5 individual cell driver mutations identified by the MutsigCV pipeline, the top 659 CpGs identified by the CMR3 process, and clinical features such as age and KI67 IHC index.
[0110] For methylation features, principal component analysis (PCA) was first used for dimensionality reduction and data integration, and the first 4 principal components (PCs) were retained. For other features, edge-based feature selection was adopted. Each feature was input into the Naïve Bayes model, and leave-one-out cross-validation was performed. Finally, age, MKI67 IHC, TP53, and NF1 were selected.
[0111] Then, a multi-omics feature soft voting technique based on the Naïve Bayes algorithm was used to assign groups to different samples, and leave-one-out cross-validation was used for training. Different voting models were used to assign subgroup labels based on available data sources.
[0112] The disclosed embodiment of the present invention also provides a computer program product or system, including a computer program, which when executed by a processor implements the steps of the above-mentioned glioma prediction method based on H3 K27 mutation.
[0113] The glioma prediction device based on H3 K27 mutation provided by the embodiment of the present invention specifically includes:
[0114] A memory and a processor; the memory is used to store program instructions; the processor is used to call program instructions, and when the program instructions are executed, any one of the above-mentioned glioma prediction methods based on H3 K27 mutation.
[0115] The disclosed embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, any one of the above-mentioned glioma prediction methods based on H3 K27 mutation.
[0116] The verification results of this verification embodiment show that allocating fixed weights for the indication can improve the performance of this method compared to the default settings. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here. In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, or optical disk, etc.
[0117] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned medium storage can be read-only memory, magnetic disk, or optical disk, etc.
[0118] The above has introduced in detail a computer device provided by the present invention. For those of ordinary skill in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A screening method for methylation CpG sites that constraints multi-omics, characterized in that, Including: S1. Obtain the gene expression data of the sample to be tested, including transcription data, CNV data, and methylation CPG probes; S2. Perform differential calculation on the transcription data to obtain differentially expressed genes; S3. Perform correlation calculation on the transcription data to obtain the correlation of transcription factors; S4. When the correlation of the transcription factor is negatively correlated, assign a negative value to the negatively correlated transcription factor to obtain a transformed transcription factor; S5. Assign a negative value to the methylation CPG probe to obtain a transformed methylation CPG probe; S6. Perform regression calculation on the transformed methylation CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed genes to obtain the regression methylation CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed genes; S7. Perform gene mutation analysis based on the regression methylation CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed genes to obtain methylation CPG sites; The correlation of the transcription factor is calculated to obtain a correlation ranking, and the first N transcription factors are retained, where N is a natural number greater than 2; The methylation CPG probe is screened to obtain a screened methylation CPG probe, and the screened methylation CPG probe is transformed by assigning a negative value to obtain a transformed methylation CPG probe; Wherein the screening is to screen probes with methylation higher than a preset threshold or methylation lower than a preset threshold; The methylation screening further includes similarity screening. The similarity of the screened methylation CPG probe is calculated to obtain a similar methylation CPG probe, and the similar methylation CPG probe is transformed by assigning a negative value to obtain a transformed methylation CPG probe; The similar methylation CPG probe is the first L similar methylation CPG probes sorted according to the similarity, where L is a natural number greater than 2; The assignment of a negative value is to multiply by negative one; The method further includes assigning a negative value to the differentially expressed gene, calculating the correlation of the differentially expressed gene, and performing a negative value transformation on the negatively correlated differentially expressed gene to obtain a transformed differentially expressed gene; performing regression calculation on the transformed methylation CPG probe, transcription factor, transformed transcription factor, CNV data, differentially expressed gene, and transformed differentially expressed gene to obtain the regression methylation CPG probe, transcription factor, transformed transcription factor, CNV data, differentially expressed gene, and transformed differentially expressed gene; The method further includes multi-scale transformation. The transformed methylation CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed gene are normalized to obtain a normalized methylation CPG probe, normalized transcription factor, normalized transformed transcription factor, normalized CNV data, and normalized differentially expressed gene. Perform regression calculation on the normalized methylation CPG probe, normalized transcription factor, normalized transformed transcription factor, normalized CNV data, and normalized differentially expressed gene to obtain the regression methylation CPG probe, transcription factor, transformed transcription factor, CNV data, and differentially expressed gene.
2. The method for screening methylation CpG sites that constrain multi-omics according to claim 1, wherein, The regression calculation is performed through a regression model, including one or more of the following: Lasso regression, ridge regression, Bayesian linear regression, elastic net regression.
3. The method for screening methylation CpG sites for constrained multi-omics according to claim 2, wherein The regression calculation obtains the optimal regularization parameter through ridge regression, ranks the regression coefficients based on the optimal regularization parameter to obtain the regression coefficient ranked first, and the corresponding methylation CpG probe, transcription factor, conversion transcription factor, CNV data, and differentially expressed gene ranked first.
4. The method for screening methylation CpG sites for constraining multi-omics according to claim 3, wherein The optimal regularization parameter is obtained by first generating a value range of the regularization parameter and then optimizing the regularization parameter within the value range.
5. The method for screening methylation CpG sites for constrained multi-omics according to claim 4, wherein The parameter optimization is performed by one or more of the following: grid search method, random search method, particle swarm optimization algorithm, niche algorithm, simulated annealing algorithm, genetic algorithm, ant colony optimization algorithm.
6. The method for screening methylation CpG sites for constrained multi-omics according to claim 1, wherein The correlation calculation is performed using the Pearson correlation coefficient to obtain the correlation ranking.
7. The methylation CpG site screening method for constraining multi-omics according to claim 1, wherein The method further includes data preprocessing, normalizing the gene expression data to obtain the processed gene expression data, and performing the processing of S2 - S7 based on the processed gene expression data.
8. The method for screening methylation CpG sites for constrained multi-omics according to claim 1, wherein The CNV data includes data in the SEG MEAN format.
9. The method for screening methylation CpG sites for constraining multi-omics according to claim 1, wherein The methylation CpG probe includes beta values.
10. A glioma prediction method based on H3 K27 mutation, characterized in that, It includes: Obtaining gene data of a test sample; Extracting the methylation site features of the gene data; The methylation site features are obtained according to the methylation CpG site screening method of the constrained multi - omics described in claims 1 - 9; Based on the methylation site features, predicting to obtain a prediction result that the test sample is pontine - like or thalamic - like.
11. The glioma prediction method based on H3 K27 mutation according to claim 10, wherein The pontine - like means that the tumor location is within the pons, and the thalamic - like means that the tumor location is outside the pons.
12. The glioma prediction method based on H3 K27 mutation according to claim 10, wherein The features further include mutation features. Extracting the mutation features of the gene data, and based on the methylation site features and mutation features, predicting to obtain a prediction result that the test sample is pontine - like or thalamic - like.
13. The glioma prediction method based on H3 K27 mutation according to claim 12, wherein The mutation features include one or more of the following: TP53 mutation, NF1 mutation.
14. The glioma prediction method based on H3 K27 mutation according to claim 10, wherein The features further include clinical features. Obtaining clinical data, extracting the clinical features from the clinical data, and based on the methylation site features and clinical features, predicting to obtain a prediction result that the test sample is pontine - like or thalamic - like.
15. The glioma prediction method based on H3 K27 mutation according to claim 14, wherein, The clinical features include one or more of the following: age, MKI67 IHC index.
16. The glioma prediction method based on H3 K27 mutation according to claim 10, wherein The features include clinical features and mutation features. Based on the methylation site features, mutation features, and clinical features, predicting to obtain a prediction result that the test sample is pontine - like or thalamic - like.
17. The glioma prediction method based on H3 K27 mutation according to claim 10, wherein The method further includes data processing. Reducing the dimension of the methylation features to obtain the reduced - dimension data, then sorting the reduced - dimension data to obtain the top S reduced - dimension methylation features, where S is a natural number greater than 3, and based on the reduced - dimension methylation site features, predicting to obtain a prediction result that the test sample is pontine - like or thalamic - like.
18. The glioma prediction method based on H3 K27 mutation according to claim 16, wherein Screening the mutation features or clinical features through edge - based feature selection to obtain the screening.
19. The glioma prediction method based on H3 K27 mutation according to claim 10, wherein The prediction is performed through one or more of the following models: Naive Bayes, Support Vector Machine, Random Forest, Decision Tree, Convolutional Neural Network, Transformer.
20. The glioma prediction method based on H3 K27 mutation according to claim 19, wherein, The prediction is obtained by inputting features into respective Naive Bayes to obtain respective prediction probabilities, and voting on the respective prediction probabilities to obtain a final prediction result.
21. A computer program product, which includes a computer program or instructions thereon, characterized in that, The computer program or instruction is executed by a processor to implement the method for screening methylated CpG sites of constrained multi-omics according to any one of claims 1-9, or to implement the method for predicting glioma based on H3 K27 mutation according to any one of claims 10-20.
22. A computer device, comprising a memory, a processor, and a computer program or instruction stored on the memory, characterized in that, The computer program or instruction is executed by a processor to implement the method for screening methylated CpG sites of constrained multi-omics according to any one of claims 1-9, or to implement the method for predicting glioma based on H3K27 mutation according to any one of claims 10-20.
23. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction is executed by a processor to implement the method for screening methylated CpG sites of constrained multi-omics according to any one of claims 1-9, or to implement the method for predicting glioma based on H3 K27 mutation according to any one of claims 10-20.
Citation Information
Patent Citations
Notch signal channel-based pan-cancer multi-omics molecular typing and prognosis model construction method
CN114927166A
Data-driven immune checkpoint blockade therapy response prediction
WO2024181928A1