Method and system for predicting cancer type based on methylation data

Through the hierarchical processing of genome-wide methylation data and the two-way attention deep learning model, the problem of inability to effectively capture the complex regulatory mode of DNA methylation in traditional methods is solved, and accurate prediction of cancer types is achieved.

CN120279984AInactive Publication Date: 2025-07-08SHENZHEN RAPHA BIOTECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510766857.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional cancer type prediction methods based on methylation data cannot effectively capture the complex regulatory patterns of DNA methylation in the epigenomic environment, ignore the special biological significance of the epigenomic functional region where the methylation site is located, and it is difficult to distinguish different cell types in tumor tissue.

Method used

By obtaining genome-wide methylation data, stratified processing of epigenome functional regions is carried out, local methylation topology patterns are constructed, regional methylation stability index is calculated, and a pre-trained two-way attention deep learning model is used for classification learning, including region encoder, chromosomal attention module and epigenetic attention module.

Benefits of technology

Accurately divide functional areas, comprehensively characterize methylation characteristics, improve the accuracy and reliability of cancer type prediction, and can deeply explore the potential relationship between various functional areas and cancer, and achieve accurate cancer type prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279984A_ABST
    Figure CN120279984A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical care informatics, in particular to a cancer type prediction method and system based on methylation data. The method comprises the following steps: acquiring whole genome methylation data of a sample; performing layering processing based on an epigenome functional region on the whole genome methylation data to obtain functional region layering data; according to the layered data of the functional regions, constructing a local methylation topological mode in each functional region to obtain a local methylation topological representation, the local methylation topological mode comprising a methylation space gradient, a methylation curvature feature and a region methylation entropy; mapping of methylation mode variable coefficients among different cell types is carried out according to local methylation topological representation, and a regional methylation stability index is obtained through calculation. According to the method, functional region layering and local topology construction are performed on methylation data, so that the regulation state of each functional region in the genome can be reflected more finely.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of healthcare informatics, and in particular, to a method and system for predicting cancer types based on methylation data. Background Art

[0002] Cancer is a general term for a group of diseases, characterized by the loss of normal regulatory mechanisms in body cells, abnormal growth and proliferation, and the ability to invade adjacent tissues and metastasize to distant sites. Common classification features of cancer include tissue type, pathological type, and site of occurrence. The occurrence and progression of cancer are not only closely related to gene mutations but also to epigenetic regulation, among which abnormal DNA methylation has been proven to play a key role in various cancers. Whole-genome methylation data obtained through high-throughput sequencing technology has become an important information source for cancer research.

[0003] Traditional methods for predicting cancer types based on methylation data often have the following problems: simply analyzing the methylation level of a single CpG site cannot capture the complex regulatory patterns of DNA methylation in the epigenomic environment; the special biological significance of epigenomic functional regions where methylation sites are located, such as promoter, enhancer, etc., is ignored; tumor tissues usually contain multiple cell types, and it is difficult for traditional methods to distinguish methylation signals from different cell types. Summary of the Invention

[0004] Based on this, it is necessary for the present invention to provide a method and system for predicting cancer types based on methylation data to solve at least one of the above technical problems.

[0005] To achieve the above object, a method for predicting cancer types based on methylation data includes the following steps: Step S1: Obtain the whole-genome methylation data of a sample; perform hierarchical processing on the whole-genome methylation data based on epigenomic functional regions to obtain functionally region-stratified data; Step S2: Construct local methylation topological patterns within each functional region according to the functionally region-stratified data to obtain a local methylation topological representation, where the local methylation topological patterns include methylation spatial gradient, methylation curvature feature, and regional methylation entropy; Step S3: Map the coefficient of variation of methylation patterns between different cell types according to the local methylation topological representation, calculate the regional methylation stability index; screen specific methylation regions based on the regional methylation stability index according to the significance p-value of the difference between cancer and normal samples to obtain specific methylation region data; Step S4: Use the pre-trained bidirectional attention deep learning model to perform classification learning on the specific methylation region data, so as to output the cancer type prediction result, where the bidirectional attention deep learning model includes a region encoder, a chromosome attention module, and an epigenotype attention module.

[0006] In the present invention, by obtaining the sample whole-genome methylation data and performing hierarchical processing based on the epigenomic functional regions, different functional regions can be accurately divided, laying a foundation for subsequent analysis, making the data more targeted and organized, and helping to deeply explore the potential associations between each functional region and cancer. Construct a local methylation topological pattern according to the hierarchical data of the functional regions, and obtain a local methylation topological representation including methylation spatial gradient, curvature features, and regional methylation entropy, comprehensively characterizing the methylation features from multiple dimensions, more precisely reflecting the genomic methylation status, and providing rich and accurate feature information for cancer type prediction. Perform mapping of the methylation pattern variation coefficient between different cell types based on the local methylation topological representation, calculate the regional methylation stability index, and screen out specific methylation regions, effectively removing irrelevant or interfering data, focusing on the regions most valuable for cancer prediction, and improving the accuracy and reliability of the prediction. Use the pre-trained bidirectional attention deep learning model to perform classification learning on the specific methylation region data. This model includes a region encoder, a chromosome attention module, and an epigenotype attention module, and can fully explore the deep features of the methylation region data and the relationships between different chromosomes and epigenomic environment types, realizing accurate prediction of cancer types.

[0007] Preferably, step S1 includes the following steps: Step S11: Obtain the whole-genome methylation data of the sample, including the methylation values of each CpG site; Step S12: Obtain the functional region annotation data in the human epigenome atlas database, and classify the whole-genome methylation sites according to the functional region annotation information, so as to obtain five types of functional region data, where the five types of functional regions include active promoter regions, active enhancer regions, transcriptional repression regions, insulator regions, and heterochromatin regions; Step S13: Perform density calculation processing on the CpG sites within each functional region according to the five types of functional region data, and perform statistical analysis on the position distribution information of CpG islands, CpG shores, and CpG shelves, so as to obtain the distribution characteristic data of the CpG sites; Step S14: Process the distribution characteristic data to generate functional region hierarchical data, where the functional region hierarchical data includes the boundary coordinates of each region, the region type label, and the methylation value matrix of the CpG sites within the region.

[0008] The present invention obtains the whole-genome methylation data of a sample, determines the methylation values of each CpG site, provides comprehensive basic data for subsequent analysis, and ensures that no potential key site information is missed. Using the annotation data of the Human Epigenome Atlas database to classify the methylation sites, five types of functional region data are obtained, making the subsequent analysis more targeted. Different functional regions play different roles in gene regulation and other aspects. After classification, the association between each region and cancer can be better explored. Calculating the CpG site density and statistical information on the position distribution within each functional region to obtain the distribution characteristic data, which helps to understand the layout rule of CpG sites in the genome and provides a basis for analyzing the relationship between them and the methylation status. Generating hierarchical data for functional regions based on the distribution characteristic data, including region boundary coordinates, type labels, and methylation beta value matrices, structuring and streamlining the complex data, facilitating in-depth mining and accurate analysis of the methylation characteristics of each functional region in subsequent steps, and laying a solid data foundation for cancer type prediction.

[0009] Preferably, step S12 includes the following steps: Step S121: Obtain histone modification information, DNA accessibility, and transcription factor binding site information in the Human Epigenome Atlas database to obtain functional region annotation data; Step S122: Perform active region recognition processing on the whole-genome region based on the histone modification information in the functional region annotation data to obtain active region data; Step S123: Perform promoter-enhancer classification processing on the active regions according to the active region data and the transcription factor binding site information in the functional region annotation data to obtain active promoter region data and active enhancer region data, where the promoter region includes the H3K4me3 enrichment peak within 2 kb upstream and downstream of the transcription start site, and the enhancer region includes the combined markers of distal H3K27ac and H3K4me1; Step S124: Determine the remaining genome region data according to the active promoter region data and the active enhancer region data; Step S125: Perform functional state classification processing on the remaining genome region data according to the DNA accessibility information and histone modification information in the functional region annotation data to obtain five types of functional region data, where the five types of functional regions include active promoter regions, active enhancer regions, transcriptionally repressed regions, insulator regions, and heterochromatin regions.

[0010] The present invention obtains histone modification information, DNA accessibility, and transcription factor binding site information from the human epigenomic atlas database, integrates them to obtain functional region annotation data, provides a key basis for subsequent accurate classification, and ensures the scientificity and accuracy of classification. Based on the histone modification information in the annotation data, active region identification is performed on the whole genome region to obtain active region data, effectively locking in the regions in the genome with the potential for active transcriptional regulation, and focusing on key regions for subsequent analysis. Combining the active region data and transcription factor binding site information, promoter-enhancer classification is performed on the active regions. The promoter region is defined as the H3K4me3 enrichment peak within 2 kb upstream and downstream of the transcription start site, and the enhancer region is defined as the distal region marked by the combination of H3K27ac and H3K4me1, further refining the functional categories of the active regions, which helps to deeply explore the unique roles of different types of regions in gene regulation and cancer development. According to the identified active promoter and enhancer region data, the remaining genome region data is determined, laying a foundation for subsequent comprehensive classification processing and ensuring full coverage of the genome regions. Using DNA accessibility information and histone modification information, functional state classification is performed on the remaining genome region data, and finally five types of functional region data are obtained, completely constructing a genome functional region classification system, providing a clear framework for subsequent methylation data analysis based on functional regions, making the analysis more targeted and systematic, and helping to deeply explore the association between methylation characteristics of different functional regions and cancer.

[0011] Preferably, step S2 includes the following steps: Step S21: Perform a sliding window process on the hierarchical data of functional regions, select adjacent intervals as local region units in units of 500 bp, so as to obtain local region unit data; Step S22: Perform basic feature extraction processing on the methylation value distribution data of CpG sites within each local region unit, so as to obtain basic feature data; Step S23: Perform a first derivative calculation process on the methylation value difference between adjacent CpG sites according to the local region unit data, so as to obtain methylation spatial gradient data; Step S24: Perform statistical feature calculation of gradient values including the mean, median, maximum, minimum, and interquartile range of all valid gradient values within the local region on the methylation spatial gradient data, so as to obtain gradient feature vector data; Step S25: Perform a weighted calculation on the CpG site density within the region according to the gradient feature vector data, so as to obtain region-normalized gradient feature data; Step S26: Perform a second derivative feature vector calculation of the methylation pattern within the local region unit data, so as to obtain methylation curvature feature data; Step S27: Calculate the information entropy of the methylation pattern within the region based on the local region unit data, so as to obtain the regional methylation entropy data characterizing the complexity of the methylation state within the region; Step S28: Perform feature vector combination processing on the basic feature data, regional normalized gradient feature data, methylation curvature feature data, and regional methylation entropy data, so as to obtain the local methylation topological feature vector; perform principal component analysis processing on the local methylation topological feature vector, and retain the principal components that explain 90% of the variance, so as to obtain the dimension-reduced local methylation topological representation.

[0012] The present invention performs a sliding window process on the hierarchical data of the functional region, selects adjacent intervals as local region units in units of 500bp, divides the genome into finer units, makes subsequent analysis more targeted and accurate, can focus on the methylation feature changes within a smaller range, and lays a foundation for deeply mining the local methylation pattern. Extract basic features from the methylation value distribution data of CpG sites within each local region unit to obtain basic feature data, such as the average methylation level and methylation variability, and initially grasp the methylation state of the local region as a whole, providing a basic reference index for subsequent complex feature analysis. Calculate the first-order derivative of the methylation value difference between adjacent CpG sites based on the local region unit data to obtain the methylation spatial gradient data, which reflects the change trend of the methylation value in space, helps to identify the increasing and decreasing regions of the methylation value, and provides key information for analyzing the dynamic changes of the methylation pattern. Calculate the statistical features of the methylation spatial gradient data to obtain the gradient feature vector data, and comprehensively describe the distribution of the gradient values through statistics such as the mean and median, more finely depict the characteristics of the methylation spatial gradient, and provide rich information for subsequent feature fusion. Perform weighted calculation on the CpG site density within the region based on the gradient feature vector data to obtain the regional normalized gradient feature data, eliminate the influence of the CpG site density on the gradient feature, make the gradient feature more comparable and representative, and ensure that the gradient features between different regions can accurately reflect the essence of the methylation change. Calculate the second-order derivative feature vector of the local region unit data to obtain the methylation curvature feature data, which reflects the acceleration of the methylation value change, further reveals the complex change law of the methylation pattern, and helps to identify the inflection points and mutation regions of the methylation pattern. Calculate the information entropy of the methylation pattern within the region based on the local region unit data to obtain the regional methylation entropy data, quantify the complexity of the methylation state within the region, and provide an important index for evaluating the diversity and uncertainty of the methylation pattern. Combine multiple feature data into a feature vector and perform dimension reduction by principal component analysis to obtain the dimension-reduced local methylation topological representation, reduce the data dimension while retaining the key information, simplify the subsequent analysis and calculation, make the features more stable and distinguishable, and provide high-quality feature input for cancer type prediction.

[0013] The present invention also provides a prediction system for cancer types based on methylation data, which is used to execute the prediction method for cancer types based on methylation data as described above. The prediction system for cancer types based on methylation data includes: A functional region stratification module, which is used to obtain the whole-genome methylation data of a sample; perform stratification processing on the whole-genome methylation data based on epigenomic functional regions to obtain functional region stratification data; A local topology construction module, which is used to construct local methylation topology patterns within each functional region according to the functional region stratification data to obtain a local methylation topology representation, where the local methylation topology pattern includes a methylation spatial gradient, a methylation curvature feature, and a regional methylation entropy; A stability mapping and screening module, which is used to perform mapping of the coefficient of variation of methylation patterns between different cell types according to the local methylation topology representation, and calculate a regional methylation stability index; perform screening of specific methylation regions based on the difference significance p-value between cancer and normal samples according to the regional methylation stability index to obtain specific methylation region data; A bidirectional attention classification module, which is used to perform classification learning on the specific methylation region data by using a pre-trained bidirectional attention deep learning model, so as to output a cancer type prediction result, where the bidirectional attention deep learning model includes a region encoder, a chromosome attention module, and an epigenomic type attention module.

[0014] By obtaining the whole-genome methylation data of a sample and performing stratification processing based on epigenomic functional regions, the present invention provides a clear framework and basic data for subsequent analysis, ensuring the pertinence and systematicness of the analysis. The local topology construction module constructs local methylation topology patterns within each functional region according to the functional region stratification data, including a methylation spatial gradient, a methylation curvature feature, and a regional methylation entropy. These features can comprehensively characterize the complexity and dynamic changes of the methylation pattern, providing rich feature information for cancer type prediction. The stability mapping and screening module calculates the coefficient of variation of methylation patterns between different cell types to obtain a regional methylation stability index, and screens out specific methylation regions based on the difference significance p-value between cancer and normal samples. This step can effectively eliminate irrelevant or interfering data, focus on the regions most valuable for cancer prediction, and improve the accuracy and reliability of the prediction. The bidirectional attention classification module performs classification learning on the specific methylation region data by using a pre-trained bidirectional attention deep learning model. The region encoder, chromosome attention module, and epigenomic type attention module in the model can fully mine the deep features of the methylation region data and the relationships between different chromosomes and epigenomic environment types, and finally output a cancer type prediction result through a softmax classification layer. Description of the Drawings

[0015] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non - limiting embodiments read in conjunction with the accompanying drawings: Figure 1 It is a schematic flowchart of the steps of a method for predicting cancer types based on methylation data of the present invention; Figure 2 is Figure 1 a detailed flowchart of step S1 in Figure 3 is Figure 1 a detailed flowchart of step S2 in Specific embodiments

[0016] The technical method of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.

[0017] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0018] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be called the second unit, and similarly the second unit may be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0019] To achieve the above - mentioned purpose, please refer to Figures 1 to 3 , the present invention provides a method for predicting cancer types based on methylation data, and the method includes the following steps: Step S1: Obtain the whole - genome methylation data of a sample; perform hierarchical processing on the whole - genome methylation data based on epigenomic functional regions to obtain functionally - region - stratified data; In the embodiments of the present invention, multiple tissue or cell samples containing cancer samples and normal samples are selected, and whole-genome methylation sequencing technologies (such as WGBS or RRBS) are used to obtain their whole-genome methylation data. The obtained data includes the methylation levels of each CpG site on the genome, usually represented by β values (i.e., the proportion of methylation). After obtaining the methylation data, hierarchical processing based on epigenomic functional regions is performed, that is, classification is carried out according to different functional regions such as CpG islands, promoter regions, enhancer regions, gene body regions, exons, and intron regions. Specifically, first, based on existing functional region annotation databases (such as ENCODE, ROADMAP, etc.), the methylation data is mapped to the corresponding functional regions according to the genomic coordinates, and then the average methylation level within each functional region is calculated. At the same time, local window sliding averaging is performed for the CpG island boundary regions to improve the fineness of region division, and finally, the hierarchical methylation data of each functional region is obtained.

[0020] Step S2: Construct local methylation topological patterns within each functional region based on the hierarchical data of functional regions to obtain a local methylation topological representation, where the local methylation topological patterns include methylation spatial gradients, methylation curvature features, and regional methylation entropy; In the embodiments of the present invention, the hierarchical data of functional regions obtained in step S1 is used to construct local methylation topological patterns within each functional region. First, the methylation spatial gradient is calculated within each functional region, that is, the change rate of the methylation levels of adjacent CpG sites is calculated according to the genomic coordinates, and the moving window method is used to smooth the gradient change to make the change trend of the methylation gradient more stable. Secondly, the methylation curvature feature is calculated. The specific method is to calculate the local methylation curvature through the third-order difference method, that is, the change rate of the methylation levels is calculated within multiple windows before and after a certain CpG site, and its second derivative is further calculated to reflect the non-linear feature of the methylation level changing with the genomic coordinates. Finally, the regional methylation entropy is calculated. This entropy value is calculated through the Shannon entropy formula, that is, based on the probability distribution of the methylation levels of CpG sites, the uncertainty level of the methylation state within the region is calculated. The larger the entropy value, the more unstable the methylation state of the region. Finally, the above three characteristic data are combined to construct a local methylation topological representation, forming a complete description of the methylation pattern within the functional region.

[0021] Step S3: Map the coefficient of variation of methylation patterns between different cell types according to the local methylation topological representation, and calculate the regional methylation stability index; screen specific methylation regions based on the regional methylation stability index according to the significance p-value of the difference between cancer and normal samples to obtain specific methylation region data; In the embodiments of the present invention, the local methylation topological representation obtained in step S2 is used to map the coefficient of variation of methylation patterns between different cell types. Specifically, for each functional region, the mean and standard deviation of the methylation gradient, curvature feature, and methylation entropy of multiple cancer samples and normal samples in this region are calculated, and based on this, the coefficient of variation is calculated, that is, the ratio of the standard deviation to the mean, to measure the stability of the methylation pattern in this region. Subsequently, the regional methylation stability index is calculated according to the coefficient of variation, and this index is used to measure the stability degree of the methylation pattern in this region between different cell types. Further, for all regional methylation stability indices, the significant p-value between cancer samples and normal samples is calculated, and specific methylation regions with significant methylation differences in cancer samples are screened out. The screening criteria can be based on the FDR multiple hypothesis test correction, and usually a threshold such as p < 0.05 is set to ensure that the screened specific methylation regions are statistically significant.

[0022] Step S4: Use the pre-trained bidirectional attention deep learning model to perform classification learning on the data of specific methylation regions, so as to output the cancer type prediction result, where the bidirectional attention deep learning model includes a region encoder, a chromosome attention module, and an epigenetic type attention module.

[0023] In the embodiments of the present invention, the pre-trained bidirectional attention deep learning model is used to perform classification learning on the data of specific methylation regions screened in step S3, so as to realize the prediction of cancer types. This deep learning model includes a region encoder, a chromosome attention module, and an epigenetic type attention module. First, the region encoder extracts features from the input data of specific methylation regions, uses a multi-layer convolutional neural network (CNN) to extract the spatial topological information in the region, and uses a bidirectional long short-term memory network (BiLSTM) to capture the long-range dependence relationship between CpG sites in the region. Subsequently, the chromosome attention module, based on the self-attention mechanism (Self-Attention), calculates the correlation between specific methylation regions on different chromosomes to ensure that the methylation features across chromosomes can influence each other. Finally, the epigenetic type attention module further optimizes the classification ability, enhances the separability of cancer types through the multi-head attention mechanism, and ensures that the model can identify the characteristic patterns of different cancer types. During the training process, the cross-entropy loss function is adopted, and the gradient is updated through the Adam optimizer. Finally, the prediction accuracy of the model is evaluated on the test data, and the prediction result of the cancer type is output.

[0024] By obtaining the whole-genome methylation data of samples and performing hierarchical processing based on epigenomic functional regions, the present invention can accurately divide different functional regions, laying a foundation for subsequent analysis, making the data more targeted and organized, and helping to deeply explore the potential associations between each functional region and cancer. According to the hierarchical data of functional regions, a local methylation topological pattern is constructed to obtain a local methylation topological representation including methylation spatial gradient, curvature features, and regional methylation entropy, comprehensively characterizing methylation features from multiple dimensions, more precisely reflecting the genomic methylation status, and providing rich and accurate feature information for cancer type prediction. Based on the local methylation topological representation, the coefficient of variation mapping of methylation patterns between different cell types is performed, the regional methylation stability index is calculated, and specific methylation regions are screened out, effectively removing irrelevant or interfering data, focusing on the regions most valuable for cancer prediction, and improving the accuracy and reliability of prediction. A pre-trained bidirectional attention deep learning model is used to classify and learn the data of specific methylation regions. The model includes a regional encoder, a chromosomal attention module, and an epigenetic type attention module, which can fully explore the deep features of methylation region data and the relationships between different chromosomes and epigenetic environment types, and achieve accurate prediction of cancer types.

[0025] Preferably, step S1 includes the following steps: Step S11: Obtain the whole-genome methylation data of the sample, including the methylation values of each CpG site; In the embodiment of the present invention, multiple cancer tissue samples and normal tissue samples are selected, and whole-genome bisulfite sequencing (WGBS) or restriction enzyme digestion methylation sequencing (RRBS) technology is used to perform whole-genome methylation sequencing on the samples. Data processing first passes through quality control steps such as adapter removal and low-quality base removal to ensure the accuracy of the sequencing data. Subsequently, the high-quality sequencing data is aligned to the human reference genome (such as GRCh38), and special alignment tools such as Bismark are used to identify the methylation status of CpG sites, and the methylation β value of each site is calculated, that is, the methylation sequencing coverage divided by the total coverage. The range of the β value is between 0 and 1, representing the completely unmethylated and completely methylated states respectively. Finally, high-confidence CpG sites with a coverage ≥10X are screened, and the whole-genome methylation data is formatted and output, including information such as the genomic coordinates, methylation β value, and sequencing depth of the CpG sites.

[0026] Step S12: Obtain the functional region annotation data in the human epigenomic atlas database, and classify the whole-genome methylation sites according to the functional region annotation information to obtain five types of functional region data, where the five types of functional regions include active promoter regions, active enhancer regions, transcriptional repression regions, insulator regions, and heterochromatin regions; In an embodiment of the present invention, functional region annotation data is obtained from a human epigenomic atlas database (such as ENCODE, ROADMAP Epigenomics, etc.). This data is derived from ChIP-seq experiments on multiple cell types and provides annotation information on epigenomic functional regions in the genome. First, according to the functional region coordinates provided by these databases, the whole-genome methylation data is aligned by genomic location. Then, based on the characteristics of the regions where CpG sites are located, they are classified into five types of functional regions, including active promoter regions (i.e., regions enriched in H3K4me3, usually near the transcription start site), active enhancer regions (regions marked by H3K27ac and H3K4me1, usually located distally to genes and regulating gene expression), transcriptional repression regions (regions enriched in H3K27me3, related to gene silencing), insulator regions (CTCF binding sites, preventing interactions between different regulatory regions), and heterochromatin regions (regions enriched in H3K9me3, usually located in highly condensed chromatin regions). During the classification process, to improve the classification accuracy, the window sliding method is used to process the boundary regions. That is, for a certain CpG site, if it falls into two different functional regions, its final category is determined according to the attribution of its center point. Finally, the epigenomic environment classification results of each CpG site are output and formatted for storage for subsequent analysis.

[0027] Step S13: Calculate the density of CpG sites within each functional region based on the data of the five types of functional regions, and statistically analyze the position distribution information of CpG islands, CpG shores, and CpG shelves, so as to obtain the distribution characteristic data of CpG sites; In the embodiment of the present invention, five types of functional region data obtained in step S12 are used to calculate the density of CpG sites in each functional region, and the position distribution information of CpG islands, CpG shores, and CpG shelves is statistically analyzed. First, for each functional region, the number of CpG sites per unit length is calculated according to a fixed window (such as 200bp or 500bp) to obtain CpG density data, and the mean and standard deviation of the CpG methylation β value in the region are further calculated to evaluate the methylation level uniformity in the region. Secondly, for CpG islands (regions with a high CpG density and a GC content > 50%), CpG shores (regions within ±2kb from the edge of the CpG island), and CpG shelves (regions located in the range of 2kb - 10kb distal to the CpG island), alignment is performed according to a reference database (such as the CpG island annotation provided by the UCSC Genome Browser), and the proportion of different types of CpG distributions in each functional region is statistically analyzed. For example, in the promoter region, CpG islands are usually abundant, while in the enhancer region, the proportion of CpG shores and CpG shelves is relatively high. Finally, the distribution characteristic data of CpG sites are sorted and output, including information such as the CpG density, CpG island / shore / shelf ratio, mean methylation level, and variance of each functional region for subsequent analysis.

[0028] Step S14: Generate functional region hierarchical data from the distribution characteristic data, where the functional region hierarchical data includes the boundary coordinates of each region, the region type label, and the methylation value matrix of CpG sites within the region.

[0029] In the embodiment of the present invention, the distribution characteristic data generated in step S13 are used to construct functional region hierarchical data. First, for the five types of functional regions, the boundary information of the regions is extracted according to the genomic coordinates, that is, the start and end positions of each functional region are recorded, and the region type label is marked. Secondly, for each functional region, a CpG site methylation β value matrix is established. The matrix uses rows to represent different samples and columns to represent CpG sites within the region, and the values in the matrix are the methylation β values of the corresponding CpG sites. To ensure data integrity, for CpG sites with low coverage or missing values, the nearest neighbor interpolation or local regression method is used to fill in the missing values. In addition, to improve the readability of CpG site data within the region, the CpG sites are arranged in the order of genomic coordinates, and the mean and standard deviation of the methylation β value in each region are calculated to provide descriptive information about the overall methylation level. Finally, the functional region hierarchical data includes the functional region boundary coordinates, the region type label, and the CpG site methylation β value matrix, and is stored in a standard format (such as BED format or HDF5 format) for subsequent modeling analysis.

[0030] The present invention obtains the whole-genome methylation data of a sample to clarify the methylation of each CpG site Values provide comprehensive basic data for subsequent analysis, ensuring that no potential key site information is missed. Using the annotation data of the human epigenomic atlas database to classify methylation sites, five types of functional region data are obtained, making the subsequent analysis more targeted. Different functional regions play different roles in gene regulation and other aspects. After classification, the association between each region and cancer can be better explored. Calculating the CpG site density and statistical information on the position distribution within each functional region to obtain distribution characteristic data, which helps to understand the layout rules of CpG sites in the genome and provides a basis for analyzing their relationship with the methylation status. Generating functional region stratification data based on the distribution characteristic data, including regional boundary coordinates, type labels, and methylation beta value matrices, structuring and streamlining complex data, facilitating in-depth exploration and precise analysis of the methylation characteristics of each functional region in subsequent steps, and laying a solid data foundation for cancer type prediction.

[0031] Preferably, step S12 includes the following steps: Step S121: Obtain histone modification information, DNA accessibility, and transcription factor binding site information in the human epigenomic atlas database, so as to obtain functional region annotation data; In the embodiment of the present invention, histone modification information, DNA accessibility data, and transcription factor binding site information are obtained from the human epigenomic atlas database (such as ENCODE, ROADMAP Epigenomics, BLUEPRINT, etc.). First, download ChIP-seq data of multiple cell types, where the histone modification information includes H3K4me3, H3K27ac, H3K4me1, H3K27me3, H3K9me3, etc.; the DNA accessibility data is derived from ATAC-seq or DNase-seq experiments, reflecting the chromatin open state; the transcription factor binding site data is derived from ChIP-seq binding site annotation. Data preprocessing includes removing low-quality data, aligning to the human reference genome (such as GRCh38), and identifying significant signal regions through peak calling (such as the MACS2 tool). Subsequently, integrate different data sources and uniformly format and output functional region annotation data, including the genomic coordinates, modified peak signal intensity, DNA accessibility score, and transcription factor binding information of each functional region.

[0032] Step S122: Perform active region identification processing on the whole genome region based on the histone modification information in the functional region annotation data, so as to obtain active region data; Based on the functional region annotation data obtained in step S121, the embodiments of the present invention perform active region identification processing on the entire genome region. First, epigenetic activity modification signals are screened, including H3K4me3 (a promoter activity marker), H3K27ac (an enhancer activity marker), and H3K4me1 (a potential enhancer activity marker). By calculating the coverage of these modifications in different genomic segments (such as a 200bp sliding window) and setting a peak threshold (such as a signal intensity three times higher than the background noise), significant active regions are screened out. Secondly, in combination with DNA accessibility data (such as ATAC-seq or DNase-seq signals), chromatin open regions are screened to exclude non-functional regions caused by histone modification false positives. Finally, active region data is output, including genomic coordinates, epigenetic activity signal intensity, and DNA accessibility scores.

[0033] Step S123: According to the active region data and the transcription factor binding site information in the functional region annotation data, perform promoter-enhancer classification processing on the active regions, so as to obtain active promoter region data and active enhancer region data, where the promoter region includes H3K4me3 enrichment peaks within 2kb upstream and downstream of the transcription start site, and the enhancer region includes a combined marker of distal H3K27ac and H3K4me1; The embodiments of the present invention use the active region data obtained in step S122 and combine the transcription factor binding site information to perform promoter-enhancer classification on the active regions. First, the H3K4me3 enrichment region is screened to determine the promoter region, and according to the transcription start site (TSS) information, the H3K4me3 peaks within 2kb upstream and downstream of the TSS are classified as active promoter regions. Secondly, the double-labeled region of H3K27ac and H3K4me1 is screened, and the part overlapping with the promoter region is removed, and the distal H3K27ac-H3K4me1 co-enrichment region is classified as an active enhancer region. In addition, the transcription factor binding data (such as ChIP-seq experimental data) is used to further optimize the classification results. For example, if a region has both promoter characteristics and enhancer characteristics, but the transcription factor strongly binds near the TSS, then this region is more likely to be a promoter region. Finally, active promoter region data and active enhancer region data are output, including region boundaries, histone modification signals, and transcription factor binding conditions.

[0034] Step S124: Determine the remaining genome region data according to the active promoter region data and the active enhancer region data; After obtaining the data of active promoter regions and active enhancer regions in the embodiments of the present invention, the remaining genomic regions are determined. First, the whole genome is segmented, and the regions that have been classified as active promoters and active enhancers are removed from it to obtain unclassified regions. Subsequently, these unclassified regions are further screened according to chromatin states. For example, genomic regions without information (such as satellite DNA and regions rich in repetitive sequences) are removed to ensure that only sequences with potential regulatory functions are retained. Finally, the data of the remaining genomic regions are output, including genomic coordinates and preliminary classification labels.

[0035] Step S125: Perform functional state classification processing on the data of the remaining genomic regions according to the DNA accessibility information and histone modification information in the functional region annotation data, so as to obtain five types of functional region data, where the five types of functional regions include active promoter regions, active enhancer regions, transcription repression regions, insulator regions, and heterochromatin regions.

[0036] Based on the data of the remaining genomic regions obtained in step S124, the embodiments of the present invention perform functional state classification on them by combining the DNA accessibility information and histone modification information in the functional region annotation data. First, regions enriched in H3K27me3 are screened and defined as transcription repression regions in combination with low DNA accessibility. Secondly, regions enriched in H3K9me3 are screened and defined as heterochromatin regions in combination with the characteristics of highly condensed chromatin. Subsequently, CTCF binding sites are identified and defined as insulator regions in combination with ATAC-seq signals. Finally, in combination with the active promoter regions and active enhancer regions identified in step S123, five types of functional region data are comprehensively generated. Finally, the complete functional region annotation data including genomic coordinates, region type labels, epigenetic modification signals, and DNA accessibility scores are output.

[0037] The present invention obtains histone modification information, DNA accessibility, and transcription factor binding site information from the human epigenomic atlas database, integrates them to obtain functional region annotation data, provides a key basis for subsequent accurate classification, and ensures the scientificity and accuracy of classification. Based on the histone modification information in the annotation data, active region recognition is performed on the whole genome region to obtain active region data, effectively locking in the regions in the genome with the potential for active transcriptional regulation, and focusing on key regions for subsequent analysis. Combining the active region data and transcription factor binding site information, promoter-enhancer classification is performed on the active regions. The promoter region is defined as the H3K4me3 enrichment peak within 2 kb upstream and downstream of the transcription start site, and the enhancer region is defined as the distal combined marker region of H3K27ac and H3K4me1, further refining the functional categories of the active regions, which helps to deeply explore the unique roles of different types of regions in gene regulation and cancer occurrence. According to the identified active promoter and enhancer region data, the remaining genome region data is determined, laying a foundation for subsequent comprehensive classification processing and ensuring full coverage of the genome regions. Using the DNA accessibility information and histone modification information, functional state classification is performed on the remaining genome region data, and finally five types of functional region data are obtained, completely constructing a genome functional region classification system, providing a clear framework for subsequent methylation data analysis based on functional regions, making the analysis more targeted and systematic, and helping to deeply explore the association between methylation characteristics of different functional regions and cancer.

[0038] Preferably, step S2 includes the following steps: Step S21: Perform a sliding window process on the stratified data of functional regions. Select adjacent intervals as local region units in units of 500 bp, so as to obtain local region unit data; In the embodiment of the present invention, a sliding window process is performed on the stratified data of functional regions. The window size is selected as 500 bp, and the window sliding step size is set as 250 bp to ensure a 50% overlap between adjacent windows to avoid the influence of boundary effects. In actual operation, first traverse the stratified data of functional regions of the entire genome, and extract the corresponding base intervals in windows of 500 bp. For each window, record the type of functional region within this interval (such as active promoter, enhancer, inhibitory region, etc.), and number each window to form local region unit data. This processing method can ensure that the epigenetic characteristics of the entire genome are evenly covered, while reducing the influence of signal fluctuations within a single region on subsequent analysis. Finally, output the local region unit data, including the genomic coordinates of each window, the corresponding functional region label, and window number information.

[0039] Step S22: Perform basic feature extraction processing on the methylation value distribution data of CpG sites within each local region unit, so as to obtain basic feature data; Based on the local region unit data obtained in step S21, the embodiments of the present invention extract the methylation value distribution data of CpG sites in each region and calculate the basic features. The specific operations include: First, screen all CpG sites in the local region unit and read their corresponding methylation values (usually taking decimal values between 0 and 1, representing the methylation ratio). Then, calculate the methylation mean, standard deviation, median, maximum value, and minimum value of all CpG sites in the region to characterize the overall methylation level and variability. In addition, count the proportions of hypermethylated (greater than 0.8), moderately methylated (0.2 to 0.8), and hypomethylated (less than 0.2) sites in the region to construct the methylation status distribution feature. This method can effectively summarize the methylation features in the region and provide basic data for subsequent gradient and curvature analysis. Finally, output the basic feature data, including information such as the methylation mean, standard deviation, median, extreme values, and methylation status ratio of the local region unit.

[0040] Step S23: Perform a first derivative calculation process on the methylation value differences between adjacent CpG sites according to the local region unit data, so as to obtain methylation spatial gradient data; Based on the local region unit data obtained in step S21, the embodiments of the present invention calculate the methylation value differences between adjacent CpG sites to obtain methylation spatial gradient data. The specific operations include: First, arrange all CpG sites in the local region in ascending order of genomic coordinates, and calculate the difference between the methylation values of adjacent two CpG sites to obtain the first derivative value, which can represent the local change rate of the methylation level. If the distance between CpG sites is greater than a set threshold (such as 200bp), then this comparison is ignored to avoid calculation errors caused by too large a distance. In addition, in order to reduce the influence of sequencing noise on gradient calculation, all gradient values are smoothed, for example, using the moving average method (window size set to 3 CpG sites) for signal smoothing. Finally, output the methylation spatial gradient data, including the genomic coordinates of each CpG site in the local region unit, the methylation difference of adjacent CpG pairs, and the smoothed gradient values.

[0041] Step S24: Calculate the statistical features of the gradient values including the mean, median, maximum value, minimum value, and interquartile range of all valid gradient values within the local region for the methylation spatial gradient data, so as to obtain gradient feature vector data; Based on the methylation spatial gradient data obtained in step S23, the embodiments of the present invention calculate the gradient statistical features in the local region. The specific operations include: for each local region unit, extract all valid gradient values and calculate their mean value to characterize the overall methylation change trend; calculate the median of the gradient values to reduce the influence of extreme values on the results; calculate the maximum and minimum values to depict the extreme cases of methylation change within the region; calculate the interquartile range of the gradient values to measure the variability of the gradient data. In addition, in order to reduce the statistical bias caused by low coverage, a minimum CpG site number threshold (such as 10 CpG sites) is set, and regions below this threshold are not subject to statistical analysis. Finally, output the gradient feature vector data, including five gradient statistical features such as mean value, median, maximum value, minimum value and interquartile range.

[0042] Step S25: Calculate the weighted value of the CpG site density within the region according to the gradient feature vector data, so as to obtain the region-normalized gradient feature data; Based on the gradient feature vector data obtained in step S24, the embodiments of the present invention calculate the region-normalized gradient features in combination with the CpG site density. The specific operations include: first, calculate the CpG site density within each local region unit, that is, the total number of CpG sites divided by the region length (500bp); then, use this density value as the weight to perform weighted adjustment on the mean value, median, maximum value and minimum value in the gradient feature vector data to compensate for the influence of uneven CpG distribution on the gradient calculation. For example, for regions with low CpG density, if the gradient changes violently, it may be due to insufficient sample coverage, so normalization processing is required to make its gradient features more comparable. Finally, output the region-normalized gradient feature data, including weighted gradient mean value, median, extreme value and other features.

[0043] Step S26: Calculate the second derivative feature vector of the methylation pattern within the region for the local region unit data, so as to obtain the methylation curvature feature data; Based on the local region unit data obtained in step S21, the embodiments of the present invention calculate the second derivative feature vector of the methylation pattern within the region to depict the methylation curvature feature. The specific operations include: first, use the methylation spatial gradient data calculated in step S23 to perform difference calculation on adjacent gradient values to obtain the change rate of the gradient, that is, the second derivative; then, statistically calculate the mean value and standard deviation of all second derivative values within the region to measure the smoothness of methylation change; in addition, calculate the maximum and minimum values of the second derivative within the region to depict the extreme curvature change cases. High curvature usually means that there may be epigenetic regulation boundaries in this region, such as promoter-enhancer conversion regions or chromatin remodeling regions, so this feature is of great significance in the classification of regulatory regions. Finally, output the methylation curvature feature data, including the second derivative mean value, standard deviation and extreme values.

[0044] Step S27: Calculate the information entropy of the methylation pattern within the region based on the local region unit data, so as to obtain the regional methylation entropy data characterizing the complexity of the methylation state within the region; In the embodiment of the present invention, based on the local region unit data obtained in step S21, the information entropy of the methylation pattern within the region is calculated to characterize the complexity of the methylation state. The specific operations include: First, count the distribution of methylation values of CpG sites within the local region unit and construct a probability distribution of methylation states. For example, divide it into three states according to the methylation value range (0.0 - 0.2, 0.2 - 0.8, 0.8 - 1.0) and calculate the frequency of each state; Then, use the Shannon entropy formula to calculate the methylation information entropy. A higher entropy value indicates that the methylation state within the region is more random, and a lower entropy value indicates that the methylation pattern is more stable or single. The information entropy can be used to identify functional genomic regions. For example, high-entropy regions usually correspond to gene regulatory switch regions, while low-entropy regions may be highly stable silent chromatin regions. Finally, output the regional methylation entropy data, including the entropy value and the methylation state probability distribution information.

[0045] Step S28: Perform feature vector combination processing on the basic feature data, regional normalized gradient feature data, methylation curvature feature data, and regional methylation entropy data, so as to obtain a local methylation topological feature vector; perform principal component analysis processing on the local methylation topological feature vector, and retain the principal components that explain 90% of the variance, so as to obtain the dimension-reduced local methylation topological representation.

[0046] In the embodiment of the present invention, the feature data obtained in steps S22, S25, S26, and S27 are combined into feature vectors to construct a local methylation topological feature vector, and dimensionality reduction is performed using principal component analysis (PCA). The specific operations include: First, standardize the data of all local region units to make the numerical ranges of different features consistent; Then, combine all features into a high-dimensional feature vector, which includes basic methylation features, normalized gradient features, curvature features, and information entropy features; Next, apply the PCA method to perform feature dimensionality reduction on the feature matrix and select the principal components that can explain 90% of the variance to reduce the data dimension and remove redundant information. Finally, output the dimension-reduced local methylation topological representation, which contains several optimally selected feature components for subsequent machine learning models or bioinformatics analysis.

[0047] The present invention performs a sliding window process on the hierarchical data of functional regions, selects adjacent intervals as local region units in units of 500bp, divides the genome into finer units, makes subsequent analysis more targeted and accurate, can focus on methylation feature changes in a smaller range, and lays a foundation for deeply mining local methylation patterns. Basic feature extraction is performed on the methylation value distribution data of CpG sites within each local region unit to obtain basic feature data, such as the average methylation level and methylation variability, to initially grasp the methylation status of the local region as a whole and provide basic reference indicators for subsequent complex feature analysis. First-order derivative calculation is performed on the methylation value differences between adjacent CpG sites according to the local region unit data to obtain methylation spatial gradient data, which reflects the change trend of methylation values in space, helps identify regions of increasing and decreasing methylation values, and provides key information for analyzing the dynamic changes of methylation patterns. Statistical feature calculation is performed on the methylation spatial gradient data to obtain gradient feature vector data, and the distribution of gradient values is comprehensively described through statistics such as the mean and median, more finely depicting the characteristics of the methylation spatial gradient and providing rich information for subsequent feature fusion. Weighted calculation of the CpG site density within the region is performed according to the gradient feature vector data to obtain region-normalized gradient feature data, eliminating the influence of CpG site density on the gradient features, making the gradient features more comparable and representative, and ensuring that the gradient features between different regions can accurately reflect the essence of methylation changes. Second-order derivative feature vector calculation is performed on the local region unit data to obtain methylation curvature feature data, which reflects the acceleration of methylation value changes, further reveals the complex change rules of methylation patterns, and helps identify the inflection points and mutation regions of methylation patterns. The information entropy of the methylation pattern within the region is calculated according to the local region unit data to obtain region methylation entropy data, quantifying the complexity of the methylation status within the region and providing an important indicator for evaluating the diversity and uncertainty of methylation patterns. Multiple feature data are combined into feature vectors and principal component analysis is performed for dimensionality reduction to obtain a dimensionally reduced local methylation topological representation, reducing the data dimension while retaining key information, simplifying subsequent analysis and calculation, making the features more stable and discriminative, and providing high-quality feature inputs for cancer type prediction.

[0048] Preferably, step S22 includes: Perform basic feature extraction processing on the methylation value distribution data of CpG sites within each local region unit, thereby obtaining basic feature data, where the basic feature data includes the average methylation level and methylation variability, and the average methylation level is the arithmetic mean of the values of all CpG sites within the region values, and the methylation variability is the standard deviation of the values of CpG sites within the region values.

[0049] In the embodiments of the present invention, the data of each local region unit is screened to ensure that there are at least two or more valid CpG sites in the region to ensure the rationality of the calculation. Then, the methylation β values of all CpG sites in the local region are extracted from the high-throughput sequencing data or methylation microarray data. Specifically, the β value refers to the methylation level of each CpG site, and its value range is between 0 and 1, where 0 indicates complete unmethylation and 1 indicates complete methylation. Next, the arithmetic mean of the β values of all CpG sites in the region is calculated, that is, all β values are added and then divided by the number of CpG sites in the region, so as to obtain the average methylation level of the local region. Subsequently, in order to characterize the dispersion degree of the methylation level of CpG sites in the region, its methylation variability is calculated, that is, the standard deviation of the β values is calculated. Specifically, first calculate the deviation of the β values of all CpG sites from the average methylation level, sum their squares, then divide by the number of CpG sites in the region minus one, and finally take the square root of this value to obtain the methylation variability. In a specific application scenario, for example, when analyzing the promoter region of cancer-related genes in the human genome, if the β values of CpG sites in a certain local region unit are 0.85, 0.78, 0.91, and 0.80 respectively, the average methylation level of this region is (0.85 + 0.78 + 0.91 + 0.80) / 4 = 0.835, and the methylation variability reflects the dispersion degree of these β values. For example, when the standard deviation calculation result is 0.05, it indicates that the methylation level of this region is relatively stable, while if the standard deviation is 0.15, it indicates that there are large fluctuations in the methylation level of CpG sites in this region. In this way, by performing the above calculation and processing on each local region unit, basic feature data can be obtained, providing data support for subsequent methylation pattern analysis and topological feature construction.

[0050] The present invention extracts the basic feature data from the methylation value distribution data of CpG sites in each local region unit, and this step is crucial. The basic feature data includes the average methylation level and the methylation variability. The average methylation level, as the arithmetic mean of the β values of all CpG sites in the region, can intuitively reflect the overall methylation degree of the local region, providing a basic quantitative index for subsequent analysis and helping researchers quickly understand whether the methylation status of the region is high or low. The methylation variability is the standard deviation of the β values of CpG sites in the region, which reveals the fluctuation of methylation values in the region and reflects the stability of the methylation status in the region. A high variability means that there are significant differences in the methylation status in the region, which may be related to the complexity of gene regulation. Through these two basic feature data, the methylation status of the local region can be effectively characterized in the preliminary stage, providing key starting information for subsequent more complex methylation topological pattern analysis, helping to more accurately mine methylation features related to cancer, laying a solid data foundation for the prediction of cancer types, and improving the accuracy and reliability of the prediction.

[0051] Preferably, step S23 includes the following steps: Step S231: Obtain the genomic coordinates of CpG sites in the local region unit according to the local region unit data, and perform sorting processing on adjacent CpG sites to obtain ordered CpG site data; The embodiment of the present invention extracts the genomic coordinate information of all CpG sites in the local region unit, and these coordinates are usually represented by the chromosome number and the starting position on the chromosome. For example, in human genome data, a CpG site may be located at "chr1:3456789", indicating that it is located at the 3456789th base position on chromosome 1. After extracting the coordinates of all CpG sites, perform ascending sorting on them to ensure that the arrangement order of CpG sites is consistent with their physical positions on the genome, thereby obtaining ordered CpG site data. In a specific application scenario, for example, in the promoter region of a certain gene, assuming that the initially extracted CpG site coordinates are "chr1:3456789, chr1:3456795, chr1:3456823, chr1:3456901", then it remains unchanged after sorting because it is already arranged in ascending order according to the genomic coordinates.

[0052] Step S232: Identify adjacent CpG site pairs according to the ordered CpG site data, and calculate the ratio of the methylation value difference to the position difference of each pair of adjacent CpG sites to obtain the original gradient value data; In the embodiments of the present invention, all adjacent CpG site pairs are identified according to the ordered CpG site data, and the ratio of the methylation value difference to the position difference of each pair of CpG sites is calculated, that is, the original gradient value data. Specifically, for each pair of adjacent CpG sites, their methylation β values are extracted, the difference between the β values of the two CpG sites is calculated, and at the same time, the difference in the genomic coordinate positions of the two is calculated. Then, the difference in β values is divided by the difference in positions to obtain the original gradient value of the CpG site pair. For example, assuming that the coordinates of a CpG site pair are "chr1:3456789" and "chr1:3456795", and their β values are 0.85 and 0.80 respectively, the methylation value difference is 0.85 - 0.80 = 0.05, and the position difference is 3456795 - 3456789 = 6bp. Therefore, the original gradient value of this CpG site pair is 0.05 / 6 ≈ 0.0083.

[0053] Step S233: Perform site distance screening on the original gradient value data, and set the gradient values of CpG site pairs with a distance exceeding 200bp to 0, so as to obtain effective gradient value data; In the embodiments of the present invention, site distance screening is performed on the original gradient value data to remove CpG site pairs with too far genomic positions, so as to avoid the influence of gradient calculation of distant CpG sites on the analysis of local methylation patterns. Specifically, a maximum distance threshold (such as 200bp) is set, and the difference in genomic coordinate positions of CpG site pairs is screened. If the distance of a pair of CpG sites exceeds 200bp, its gradient value is set to 0. For example, assuming that the coordinates of a pair of CpG sites are "chr1:3456789" and "chr1:3457020" respectively, and their position difference is 231bp, which exceeds the 200bp threshold. Therefore, the gradient value of this pair of CpG sites is set to 0, while other CpG site pairs that meet the distance requirements retain their original gradient values.

[0054] Step S234: Perform positive / negative sign direction analysis on the effective gradient value data and perform amplitude analysis, so as to obtain methylation spatial gradient data.

[0055] In the embodiments of the present invention, positive / negative sign direction analysis is performed on the effective gradient value data, and amplitude analysis is carried out to obtain methylation spatial gradient data. Specifically, the positive / negative sign direction analysis is used to judge the trend of methylation level change. If the β values of adjacent CpG sites change from small to large, the gradient value is positive, indicating that the methylation level increases in this local region; if the β values change from large to small, the gradient value is negative, indicating that the methylation level decreases. For example, if the β values of a pair of CpG sites are 0.70 and 0.75 respectively, the gradient value is positive, indicating methylation increase; if the β values are 0.85 and 0.78 respectively, the gradient value is negative, indicating methylation decrease. At the same time, amplitude analysis is performed to evaluate the severity of methylation level change, that is, statistical features such as the mean, maximum, minimum, and standard deviation of all effective gradient values are calculated, so as to characterize the methylation change pattern in this local region. For example, in cancer samples, if the absolute values of the gradient values in a certain gene promoter region are generally large, it may indicate abnormal changes in the methylation pattern in this region, thereby affecting gene expression regulation.

[0056] The present invention obtains the genomic coordinates of CpG sites based on the local region unit data and sorts them to obtain ordered data, laying a foundation for subsequent accurate analysis of the relationship between adjacent sites, ensuring the accuracy and orderliness of the calculation, and making the calculation of methylation value differences more logical and targeted. By identifying adjacent CpG site pairs from the ordered data and calculating the ratio of methylation value difference to position difference, the original gradient value data is obtained, which intuitively reflects the change rate of methylation values in the genomic space, helps to identify regions where methylation values change sharply, and provides clues for discovering potential methylation regulatory regions. Site distance screening is performed on the original gradient value data, and the gradient values of site pairs with a distance exceeding 200 bp are set to 0, effectively filtering out site pairs with a relatively large distance and weak correlation, making the gradient value data more focused on closely related sites within the local region, improving the correlation and reliability of the data, and reducing noise interference. Positive / negative sign direction and amplitude analysis are performed on the effective gradient value data to obtain methylation spatial gradient data, which not only clarifies the direction of methylation value change, but also quantifies the change amplitude, more comprehensively depicts the dynamic change characteristics of methylation values in space, provides key information for subsequent in-depth analysis of methylation patterns, helps to accurately identify subtle differences in methylation patterns, and improves the accuracy of cancer type prediction.

[0057] Preferably, the mapping of the methylation pattern coefficient of variation between different cell types in step S3 specifically includes: Obtain the reference methylation information of the main tissue types of the human body, including the methylation data of epithelial cells, mesenchymal cells, hematopoietic cells, nerve cells, and endocrine cells, so as to obtain the reference cell type data; In the embodiments of the present invention, the genome-wide methylation data of epithelial cells, mesenchymal cells, hematopoietic cells, nerve cells, and endocrine cells are extracted from publicly available methylation databases (such as The Cancer Genome Atlas, TCGA, or Epigenome Roadmap) or experimental sequencing data. The specific operations include: 1) screening samples with high-quality sequencing data, 2) extracting the β values (0 represents unmethylated, 1 represents fully methylated) of each cell type at CpG sites, 3) organizing the methylation data of CpG sites according to chromosomal positions, and 4) generating genome-wide methylation reference data for each cell type. For example, in practical applications, the methylation data of mammary epithelial cells, hematopoietic stem cells derived from bone marrow, neurons in brain tissue, etc. from the TCGA database may be extracted and stored in a matrix format, where the rows represent CpG sites and the columns represent samples or cell types.

[0058] Perform local methylation topological pattern construction processing on each cell type according to the reference cell type data, so as to obtain the methylation topological representation of the reference cell type; In the embodiments of the invention, according to the reference cell type data, functional regions in the genome (such as promoters, enhancers, CpG islands, etc.) are selected, and the local methylation topological representations of these regions are constructed. The specific methods include: 1) within each cell type, connecting adjacent CpG sites according to genomic coordinates to form a local topological structure, 2) calculating the methylation level correlation between CpG sites, and using the Pearson correlation coefficient or distance metric method to evaluate the co-methylation pattern of CpG sites, 3) using network modeling methods (such as adjacency matrix or graph convolutional network) to represent the local methylation topology. For example, in mammary epithelial cells, assuming that the β values of three CpG sites in a promoter region are 0.85, 0.87, and 0.90 respectively, when constructing its topological representation, the methylation correlation between adjacent CpG sites will be calculated and stored in the form of a matrix or graph to form the methylation topological pattern of this region.

[0059] Calculate the standard deviation and mean of the methylation topological feature vectors of each functional region among different cell types according to the local methylation topological representation and the methylation topological representation, so as to obtain regional variation statistical data; In the invention embodiment, based on the local methylation topology representations of different cell types, the standard deviation and mean of the methylation topology features of each functional region in different cell types are calculated to quantify the variability of methylation patterns between different tissue types. The specific method includes: 1) For each functional region, extract its topology feature vector (such as the methylation correlation matrix or topological connectivity degree between CpG sites) from all reference cell types; 2) Calculate the mean of the topology features of this functional region in all cell types; 3) Calculate the standard deviation of the topology features of this functional region among cell types. For example, assume that the mean values of the topology features of a certain promoter region in five cell types are 0.75, 0.80, 0.78, 0.82, and 0.79 respectively, then its mean value is (0.75 + 0.80 + 0.78 + 0.82 + 0.79) / 5 = 0.788, and the standard deviation reflects the degree of methylation topology differences between different cell types.

[0060] Calculate the coefficient of variation of the methylation pattern of each functional region according to the regional variation statistical data, thereby obtaining the coefficient of variation data, where the calculation formula of the coefficient of variation of the methylation pattern is: Where, is the serial number of the functional region, is the coefficient of variation of the methylation pattern of the functional region , is the standard deviation of the local methylation topology representation constructed by the functional region in all reference cell types, is the arithmetic mean of the local methylation topology representation constructed by the functional region in all reference cell types; In the invention embodiment, using the standard deviation and mean calculated in the previous step, calculate the coefficient of variation of the methylation pattern of each functional region to evaluate the stability or variability of the regional methylation pattern. The specific method includes: 1) For each functional region, extract the standard deviation and mean of its topology features in all reference cell types; 2) Calculate the coefficient of variation of this region, that is, divide the standard deviation by the mean to normalize the variability. In the application scenario, for example, in cancer research, the promoter regions of some cancer-related genes may have a relatively high coefficient of variation, indicating that the methylation patterns of this region in different cell types vary greatly, while the promoter regions of some housekeeping genes may have a relatively low coefficient of variation, indicating that the methylation patterns of this region are relatively stable in different cell types.

[0061] Perform hyperbolic tangent function mapping processing on the coefficient of variation data and perform sorting processing on the whole genome region, thereby obtaining the regional methylation stability index.

[0062] The inventive embodiments perform a non-linear transformation on the calculated coefficient of variation data to enhance the contrast between high-variation regions and low-variation regions, and sort all genome-wide functional regions according to the coefficient of variation to obtain a regional methylation stability index. The specific method includes: 1) applying a hyperbolic tangent function mapping to the coefficient of variation data of all functional regions to smooth the influence of extreme values; 2) sorting all functional regions from high to low according to the coefficient of variation; 3) assigning a stability index according to the sorting result, where a higher index indicates a more stable regional methylation pattern. For example, if the coefficient of variation of some functional regions is between 0.1 and 0.3, the mapped value may fall between -1 and 1, while regions with a larger coefficient of variation (such as 0.8) may be mapped to a higher stability index range. This process helps to reveal which gene regions have more consistent methylation patterns among different cell types.

[0063] The present invention obtains reference methylation information of major human tissue types, covering multiple cell types such as epithelial cells and mesenchymal cells, to obtain reference cell type data, providing comprehensive methylation background data for subsequent analysis and ensuring the comprehensiveness and representativeness of the analysis. Based on the reference cell type data, local methylation topological patterns of each cell type are constructed to obtain a methylation topological representation, visually showing the methylation characteristics of different cell types and facilitating the comparison and analysis of methylation differences between cells. By calculating the standard deviation and mean of the methylation topological feature vectors of each functional region among different cell types through the local methylation topological representation, regional variation statistical data is obtained, quantifying the variation of methylation characteristics among different cell types and providing key data support for evaluating regional methylation stability. Using the regional variation statistical data, the coefficient of variation of the methylation pattern of each functional region is calculated, using the formula , to accurately measure the relative variation degree of the methylation pattern of each functional region, further clarifying the stability differences of regional methylation. The coefficient of variation data is processed by a hyperbolic tangent function mapping, and the whole genome region is sorted to obtain a regional methylation stability index, converting the coefficient of variation into a more comparable and sortable stability index, providing an important basis for subsequent screening of specific methylation regions and helping to accurately locate methylation abnormal regions related to cancer.

[0064] Preferably, the screening of specific methylation regions described in step S3 specifically includes: Obtaining sample data with known cancer type labels and normal sample data, thereby obtaining training sample set data; In the embodiments of the present invention, methylation data containing cancer samples and normal tissue samples are extracted from public databases (such as TCGA, GEO databases) or clinical sequencing data, and the quality and consistency of the data sources are ensured. The specific operations include: 1) Selecting a target cancer type, such as liver cancer (HCC) or lung adenocarcinoma (LUAD); 2) Screening samples that meet the criteria, such as 200 cases of cancer tissue samples with pathological confirmation and 100 cases of matched normal tissue samples in TCGA; 3) Extracting the whole-genome methylation data of each sample to ensure consistent data formats, such as using Infinium 450K chips or WGBS (whole-genome bisulfite sequencing) data; 4) Performing standardization processing on the data, such as β-value (range of 0 - 1) or M-value conversion (logit conversion), and storing it in matrix format, where rows represent CpG sites or functional regions and columns represent samples, finally obtaining structured training sample set data.

[0065] According to the regional methylation stability index and the training sample set data, a significance analysis of the differences between cancer and normal samples is performed for each functional region, thereby obtaining regional difference significance p-value data, where the significance analysis of the differences includes comparative analysis, hypothesis testing, and analysis of the mean difference between groups. In the embodiments of the present invention, according to the regional methylation stability index and the training sample set data, a statistical analysis is performed on each functional region to identify methylation regions with significant differences between cancer and normal tissues. The specific methods include: 1) Comparative analysis, that is, calculating the average methylation level of each functional region in cancer samples and normal samples; 2) Hypothesis testing, such as using a t-test or Mann-Whitney U test to calculate the p-value between the two groups; 3) Analysis of the mean difference between groups, that is, calculating the mean difference in methylation levels between the cancer group and the normal group. For example, for the promoter region, the average methylation level of the cancer sample group is 0.85, while that of the normal group is 0.40, so the mean difference is 0.45, and its statistical significance is calculated. The finally output p-value data is used to evaluate the degree of methylation difference of each functional region between cancer and normal tissues.

[0066] Perform a p-value to logarithm conversion process based on the regional difference significance p-value data, and perform a product calculation process based on the regional methylation stability index, thereby obtaining cancer-specific score data. In the embodiments of the present invention, in order to enhance the ability to distinguish cancer-specific regions, the negative logarithm transformation is performed on the significance p-value, and the weighted calculation is combined with the regional methylation stability index. The specific method includes: 1) performing the -log10 transformation on the p-value of each functional region. For example, the region with a p-value of 0.0001 is transformed to 4. 2) Performing the product calculation based on the regional methylation stability index to adjust the influence of the p-value. For example, if the stability index of a certain region is 1.2, then its cancer-specific score is calculated as 4×1.2 = 4.8. 3) For all functional regions, the calculated cancer-specific score data is obtained and stored in a matrix format, where each row represents a functional region and each column represents the cancer-specific score of that region. This process can amplify the regions with high stability and significant differences, making them more representative in the subsequent screening process.

[0067] All regions are sorted in descending order according to the cancer-specific score data, and the top 5% of the regions are selected to obtain the candidate specific region data; In the embodiments of the present invention, all functional regions are sorted according to the cancer-specific score, and the top 5% of the regions are selected as candidate specific regions. The specific method includes: 1) sorting all functional regions in descending order according to the cancer-specific score. For example, if there are 200,000 functional regions in the whole genome, then the top 10,000 regions are selected. 2) Storing the screened candidate region data, including information such as gene names, genomic coordinates, and specific scores. For example, assuming that some CpG islands or promoter regions rank high after sorting, they are more likely to be specific in cancer diagnosis. 3) Formatting the candidate region data into a list for subsequent modeling analysis. This step ensures that the selected regions have the most significant methylation differences between cancer and normal tissues and are relatively stable among different cell types.

[0068] Perform logistic regression analysis based on L1 regularization on the candidate specific region data to obtain the regional discriminative power score data; In the embodiments of the present invention, machine learning modeling is performed on candidate specific regions to further screen out functional regions with high discriminative ability. The specific methods include: 1) constructing a logistic regression model, with the input feature being the methylation level of the candidate specific region and the label being cancer or normal; 2) applying L1 regularization (Lasso regression) to reduce the influence of redundant regions by increasing sparsity, so that only the most discriminative functional regions are retained. For example, training is performed using LogisticRegression(solver='liblinear',penalty='l1') in sklearn; 3) calculating the L1 regularization weight value of each region to evaluate its discriminative ability. For example, if the regression coefficient of a certain region is 0.8, it indicates that it makes a greater contribution to the classification of cancer, while regions with coefficients close to 0 may be excluded; 4) storing the final regional discriminative score data for subsequent screening of high-discriminative regions. This process can effectively remove redundant information and enable the finally screened methylation regions to have stronger classification ability.

[0069] High-discriminative region screening processing is performed according to the regional discriminative score data, and 1000 regions with high discriminative ability are selected to obtain specific methylation region data.

[0070] In the embodiments of the present invention, according to the regional discriminative score data, the top 1000 regions are selected to finally determine specific methylation regions. The specific methods include: 1) sorting all candidate regions in descending order according to the L1 regularization score; 2) selecting the top 1000 regions to ensure that they have the highest information content in the cancer classification task. For example, the promoter regions of some cancer-related genes (such as TP53, BRCA1) may be selected; 3) storing the finally screened specific methylation region data, including information such as gene names, genomic coordinates, and methylation levels. The finally obtained specific methylation region data can be used for early cancer detection, development of methylation markers, and personalized medical research.

[0071] The present invention obtains sample data of known cancer type labels and normal sample data to form a training sample set, providing basic data for subsequent analysis and ensuring the accuracy and reliability of the analysis. Through the regional methylation stability index and the data of the training sample set, a significance analysis of the differences between cancer and normal samples is performed for each functional region, obtaining the p-value data of regional significance, which can quantify the degree of difference between each functional region in cancer and normal samples, providing key information for screening regions related to cancer. The p-value is converted into a logarithm and multiplied by the regional methylation stability index to obtain cancer-specific score data. This conversion and calculation process can more intuitively reflect the association strength between each region and cancer, providing a clearer basis for subsequent screening. According to the cancer-specific score data, all regions are sorted in descending order, and the top 5% of the regions are selected as candidate specific region data, which can effectively narrow the screening range, focus on the regions most likely related to cancer, and improve the analysis efficiency. A logistic regression analysis based on L1 regularization is performed on the candidate specific region data to obtain region discriminative score data, which can further evaluate the discriminative ability of each region for cancer types, providing a more accurate indicator for the final screening. According to the region discriminative score data, 1000 regions with high discriminative ability are selected as specific methylation region data. These region data can more accurately reflect the methylation characteristics related to cancer, providing strong support for cancer type prediction and improving the accuracy and credibility of the prediction.

[0072] Preferably, step S4 includes the following steps: Step S41: Perform feature vector extraction processing on the specific methylation region data, and input the feature vector into the region encoder of the pre-trained bidirectional attention deep learning model for feature encoding processing, thereby obtaining region latent representation data; In the embodiments of the present invention, feature vector extraction processing is performed on specific methylation region data so as to be input into the region encoder of a pre-trained bidirectional attention deep learning model for feature encoding. The specific operations include: extracting multi-dimensional features for each specific methylation region, such as the mean methylation level, variance, methylation state transition rate, distribution density of CpG sites within the region, relative position to the gene promoter, etc., to ensure that the feature vector can comprehensively represent the methylation state; using a dimensionality reduction method (such as PCA or t-SNE) to optimize the feature space to reduce redundant information and improve feature discrimination, and finally converting each specific methylation region into a high-dimensional feature vector, for example, constructing a vector with 128 feature dimensions for each region; inputting the feature vector into the region encoder, which uses a stacked bidirectional LSTM network to simultaneously capture the local patterns of the region and the context information across regions. Specifically, each feature vector is processed by forward and backward LSTM cells and combined with an attention mechanism to enhance the attention to key methylation patterns; after encoding, the output region latent representation data is stored in a tensor format (such as shape [number of samples, number of regions, 128]) for subsequent attention calculation.

[0073] Step S42: Calculate the correlation weight matrix between specific regions on different chromosomes through the chromosome attention module in the bidirectional attention deep learning model based on the region latent representation data, so as to obtain the chromosome attention output data; In the embodiments of the present invention, based on the region latent representation data, the correlation weight matrix between specific regions on different chromosomes is calculated through the chromosome attention module in the bidirectional attention deep learning model to obtain the chromosome attention output data. The specific operations include: grouping the region latent representation data according to chromosomes based on the chromosome position information of each specific methylation region to ensure that the feature correlation within the chromosome can be reflected during attention calculation; using a self-attention mechanism to calculate the weighted correlation between specific regions on the same chromosome. For example, using the dot product attention method, calculating the similarity between the latent representation of each region and other regions, and normalizing the weights through Softmax, so that regions with high correlation contribute more to the final feature representation; through normalized weighted summation, aggregating the representations of all specific methylation regions to generate the overall methylation feature of each chromosome, and finally forming the chromosome attention output data, which is stored in a tensor form, such as shape [number of samples, number of chromosomes, 128], to ensure that it can be further used for feature modeling across epigenetic environments.

[0074] Step S43: Aggregate the region latent representation data based on the epigenomic environment type through the epigenotype attention module, and calculate the correlation weight matrix between different epigenomic environment types, so as to obtain the epigenotype attention output data; In the embodiment of the present invention, based on the epigenotype attention module, the region latent representation data is aggregated based on the epigenomic environment type, and the correlation weight matrix between different epigenomic environment types is calculated to obtain the epigenotype attention output data. The specific operations include: grouping the region latent representation data according to the epigenomic environment type where the specific methylation region is located (such as promoter region, enhancer region, suppressor region, CpG island, etc.), to ensure that the commonality of the same type of epigenomic environment can be reflected during attention calculation; using the attention mechanism to calculate the regional correlation weights within the same epigenotype, and weighted summing the region latent representation data of this epigenotype based on the calculated weights to form the overall representation of this epigenotype. For example, for all methylation sites in the promoter region, its weighted summation feature representation can better reflect the epigenetic regulation state of the promoter region; further calculate the correlation between different epigenotypes, such as measuring whether the methylation characteristics of the enhancer region affect the adjacent promoter region, and integrate the characteristics of different epigenotypes through weighted summation. Finally, the epigenotype attention output data is obtained. The storage format of this data is usually in tensor format, such as shape [number of samples, number of epigenotypes, 128], for subsequent global feature fusion.

[0075] Step S44: Perform weighted fusion processing on the chromosome attention output data and the epigenotype attention output data, so as to obtain the global epigenomic feature representation data; In the embodiment of the present invention, weighted fusion processing is performed on the chromosome attention output data and the epigenotype attention output data to obtain the global epigenomic feature representation data. The specific operations include: respectively assigning learnable weight parameters to the two attention output data, so as to automatically adjust the contribution ratio of the two according to the characteristics of the data during the training process; adopting a fusion strategy, such as direct weighted summation fusion, that is, adding the chromosome attention output data and the epigenotype attention output data in proportion, or using a gating mechanism (such as GatedAttention) to dynamically adjust the fusion method of different source information, to ensure that the fused features can more comprehensively represent the methylation characteristics of cancer samples; the finally obtained global epigenomic feature representation data retains the chromosome structure information and epigenomic environment information of the methylation region and is used for subsequent cancer classification tasks. The storage format of this data is usually in tensor format, such as shape [number of samples, 1, 128], representing the global epigenomic features of each sample.

[0076] Step S45: Use the softmax classification layer to perform multi-classification on the global epigenomic feature representation data, thereby outputting the cancer type prediction result.

[0077] In the embodiment of the present invention, the softmax classification layer is used to perform multi-classification on the global epigenomic feature representation data to output the cancer type prediction result. The specific operations include: inputting the global epigenomic feature representation data into a fully connected layer, which is used to further extract features related to the classification task and map the feature dimension to an output dimension that matches the number of cancer types. For example, if 10 different cancer types need to be predicted, the output dimension of the fully connected layer is 10; applying the softmax activation function to convert the output into a probability distribution, ensuring that the sum of the prediction probabilities of all classes is 1, so as to intuitively express the possibility of the sample belonging to each cancer type. For example, the softmax output of a certain sample is [0.1, 0.05, 0.6, 0.15, 0.1], indicating that the sample has the highest probability of belonging to the third cancer type; finally, according to the maximum probability value of the softmax output, determine the cancer prediction category of the sample and output the classification result, such as "This sample is predicted to be lung adenocarcinoma (LUAD)"; during the model training process, the cross-entropy loss function is used to optimize the classification effect, and the model parameters are updated through stochastic gradient descent (SGD) or the Adam optimizer to improve the accuracy of cancer classification. Finally, this classification result can be used in the fields of early cancer detection and precision medicine.

[0078] The present invention extracts eigenvectors from specific methylation region data and inputs them into the region encoder of a pre-trained bidirectional attention deep learning model to obtain region latent representation data. This step can effectively extract and encode the key features of methylation regions, providing a basis for subsequent analysis. By means of a chromosome attention module, the correlation weight matrix between specific regions on different chromosomes is calculated to obtain chromosome attention output data, which helps to reveal the mutual relationship of methylation features between different chromosomes and enhances the model's ability to capture complex relationships between chromosomes. Through an epigenomic type attention module, aggregation processing based on the epigenomic environment type is performed, and the correlation weight matrix between different epigenomic environment types is calculated to obtain epigenomic type attention output data. This step can integrate information of different epigenomic environment types and improve the model's sensitivity to epigenomic environment changes. The chromosome attention output data and the epigenomic type attention output data are weighted and fused to obtain global epigenomic feature representation data. This step integrates feature information at different levels to form a more comprehensive feature representation, providing richer information for the final cancer type prediction. The softmax classification layer is used to perform multi-classification processing on the global epigenomic feature representation data to output cancer type prediction results. Through multi-classification processing, this step can accurately classify samples into different cancer types.

[0079] The present invention also provides a prediction system for cancer types based on methylation data, which is used to execute the prediction method for cancer types based on methylation data described above. The prediction system for cancer types based on methylation data includes: A functional region stratification module, which is used to obtain the whole-genome methylation data of a sample; perform stratification processing on the whole-genome methylation data based on epigenomic functional regions to obtain functional region stratification data; A local topology construction module, which is used to construct local methylation topological patterns within each functional region according to the functional region stratification data to obtain local methylation topological representations, where the local methylation topological patterns include methylation spatial gradients, methylation curvature features, and regional methylation entropy; A stability mapping and screening module, which is used to map the coefficient of variation of methylation patterns between different cell types according to the local methylation topological representation, calculate the regional methylation stability index; perform screening of specific methylation regions based on the difference significance p-value between cancer and normal samples according to the regional methylation stability index to obtain specific methylation region data; A bidirectional attention classification module, which is used to perform classification learning on the specific methylation region data by using a pre-trained bidirectional attention deep learning model, so as to output cancer type prediction results, where the bidirectional attention deep learning model includes a region encoder, a chromosome attention module, and an epigenomic type attention module.

[0080] The present invention provides a clear framework and basic data for subsequent analysis by obtaining the whole-genome methylation data of samples and performing hierarchical processing based on epigenomic functional regions, ensuring the pertinence and systematicness of the analysis. The local topology construction module constructs the local methylation topology patterns within each functional region according to the hierarchical data of functional regions, including methylation spatial gradient, methylation curvature features, and regional methylation entropy. These features can comprehensively characterize the complexity and dynamic changes of methylation patterns, providing rich feature information for cancer type prediction. The stability mapping and screening module calculates the coefficient of variation of methylation patterns between different cell types to obtain the regional methylation stability index, and screens out specific methylation regions based on the differential significance p-value between cancer and normal samples. This step can effectively eliminate irrelevant or interfering data, focus on the regions most valuable for cancer prediction, and improve the accuracy and reliability of prediction. The bidirectional attention classification module uses a pre-trained bidirectional attention deep learning model to classify and learn the data of specific methylation regions. The regional encoder, chromosome attention module, and epigenomic type attention module in the model can fully explore the deep features of methylation region data and the relationships between different chromosomes and epigenomic environment types, and finally output the cancer type prediction results through the softmax classification layer.

[0081] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed by the present invention.

[0082] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A prediction method for cancer types based on methylation data, characterized in that It includes the following steps: Step S1: Obtain the whole-genome methylation data of the sample; perform hierarchical processing on the whole-genome methylation data based on epigenomic functional regions to obtain functional region hierarchical data; Step S2: Construct local methylation topological patterns within each functional region according to the functional region hierarchical data to obtain local methylation topological representations, where the local methylation topological patterns include methylation spatial gradients, methylation curvature features, and regional methylation entropy; Step S3: Perform mapping of the coefficient of variation of methylation patterns between different cell types according to the local methylation topological representations, and calculate the regional methylation stability index; screen specific methylation regions based on the regional methylation stability index according to the significance p-value of the difference between cancer and normal samples to obtain specific methylation region data; Step S4: Use a pre-trained bidirectional attention deep learning model to perform classification learning on the specific methylation region data, and thus output cancer type prediction results, where the bidirectional attention deep learning model includes a region encoder, a chromosome attention module, and an epigenotype attention module.

2. The prediction method of cancer type based on methylation data according to claim 1, wherein Step S1 includes the following steps: Step S11: Obtain the whole-genome methylation data of the sample, including the methylation values of each CpG site; Step S12: Obtain functional region annotation data from the human epigenome atlas database, and classify the whole-genome methylation sites according to the functional region annotation information to obtain five types of functional region data, where the five types of functional regions include active promoter regions, active enhancer regions, transcription repression regions, insulator regions, and heterochromatin regions; Step S13: Perform density calculation processing on CpG sites within each functional region according to the five types of functional region data, and perform statistical analysis on the position distribution information of CpG islands, CpG shores, and CpG shelves, so as to obtain the distribution characteristic data of CpG sites; Step S14: Process the distribution feature data to generate functional region hierarchical data, where the functional region hierarchical data includes the boundary coordinates of each region, the region type label, and the methylation value matrix of CpG sites within the region. Value matrix.

3. The prediction method of cancer type based on methylation data according to claim 2, characterized in that, Step S12 includes the following steps: Step S121: Obtain histone modification information, DNA accessibility, and transcription factor binding site information from the human epigenome atlas database to obtain functional region annotation data; Step S122: Perform active region recognition processing on the whole-genome region based on the histone modification information in the functional region annotation data to obtain active region data; Step S123: Perform promoter-enhancer classification processing on the active regions according to the active region data and the transcription factor binding site information in the functional region annotation data to obtain active promoter region data and active enhancer region data, where the promoter region includes the H3K4me3 enrichment peak within 2 kb upstream and downstream of the transcription start site, and the enhancer region includes the combined markers of distal H3K27ac and H3K4me1; Step S124: Determine the remaining genome region data according to the active promoter region data and the active enhancer region data; Step S125: Classify the remaining genomic region data according to the DNA accessibility information and histone modification information in the functional region annotation data, so as to obtain five types of functional region data, where the five types of functional regions include active promoter regions, active enhancer regions, transcriptional repression regions, insulator regions, and heterochromatin regions.

4. The prediction method of cancer type based on methylation data according to claim 3, wherein Step S2 includes the following steps: Step S21: Perform a sliding window process on the functional region stratified data, select adjacent intervals as local region units in units of 500 bp, so as to obtain local region unit data; Step S22: Extract basic features from the methylation value distribution data of CpG sites within each local region unit, so as to obtain basic feature data; Step S23: Calculate the first derivative of the methylation value difference between adjacent CpG sites according to the local region unit data, so as to obtain methylation spatial gradient data; Step S24: Calculate the statistical features of the gradient values including the mean, median, maximum, minimum, and interquartile range of all valid gradient values within the local region for the methylation spatial gradient data, so as to obtain gradient feature vector data; Step S25: Perform a weighted calculation on the CpG site density within the region according to the gradient feature vector data, so as to obtain region-normalized gradient feature data; Step S26: Calculate the second derivative feature vector of the methylation pattern within the local region unit data, so as to obtain methylation curvature feature data; Step S27: Calculate the information entropy of the methylation pattern within the region according to the local region unit data, so as to obtain region methylation entropy data characterizing the complexity of the methylation state within the region; Step S28: Combine the basic feature data, region-normalized gradient feature data, methylation curvature feature data, and region methylation entropy data to form a feature vector, so as to obtain a local methylation topological feature vector; perform principal component analysis on the local methylation topological feature vector, and retain the principal components that explain 90% of the variance, so as to obtain the dimension-reduced local methylation topological representation.

5. The prediction method of cancer type based on methylation data according to claim 4, wherein Step S22 includes: Extract the basic feature data from the methylation value distribution data of CpG sites in each local region unit, so as to obtain the basic feature data, where the basic feature data includes the average methylation level and the methylation variability, and the average methylation level is all CpG sites in the region The arithmetic mean of the values, and the methylation variability is the CpG sites in the region The standard deviation of the values.

6. The prediction method of cancer type based on methylation data according to claim 5, wherein Step S23 includes the following steps: Step S231: Obtain the genomic coordinates of CpG sites within the local region unit according to the local region unit data, and sort the adjacent CpG sites, so as to obtain ordered CpG site data; Step S232: Identify adjacent CpG site pairs according to the ordered CpG site data, and calculate the ratio of the methylation value difference to the position difference of each pair of adjacent CpG sites, so as to obtain raw gradient value data; Step S233: Perform site distance screening on the raw gradient value data, and set the gradient values of CpG site pairs with a distance exceeding 200 bp to 0, so as to obtain valid gradient value data; Step S234: Analyze the positive and negative sign directions of the valid gradient value data, and perform amplitude analysis, so as to obtain methylation spatial gradient data.

7. The prediction method of cancer type based on methylation data according to claim 6, wherein The mapping of the methylation pattern coefficient of variation between different cell types described in Step S3 specifically includes: Obtain reference methylation information of major human tissue types, including methylation data of epithelial cells, mesenchymal cells, hematopoietic cells, nerve cells, and endocrine cells, so as to obtain reference cell type data; Perform local methylation topological pattern construction processing on each cell type according to the reference cell type data, so as to obtain the methylation topological representation of the reference cell type; Calculate the standard deviation and mean of the methylation topological feature vectors of each functional region among different cell types according to the local methylation topological representation and the methylation topological representation, so as to obtain regional variation statistical data; Calculate the methylation pattern coefficient of variation of each functional region according to the regional variation statistical data, so as to obtain coefficient of variation data, where the calculation formula of the methylation pattern coefficient of variation is: Among them, is the serial number of the functional region, is the methylation pattern coefficient of variation of the functional region , is the standard deviation of the local methylation topological representation constructed for the functional region in all reference cell types, is the arithmetic mean of the local methylation topological representation constructed for the functional region in all reference cell types; Perform hyperbolic tangent function mapping processing on the coefficient of variation data and perform sorting processing on the whole genome region, so as to obtain the regional methylation stability index.

8. The prediction method of cancer type based on methylation data according to claim 7, wherein The specific screening of the specific methylation region described in step S3 specifically includes: Obtain sample data with known cancer type labels and normal sample data, so as to obtain training sample set data; According to the regional methylation stability index and the training sample set data, perform differential significance analysis between cancer and normal samples for each functional region, so as to obtain regional differential significance p-value data, where the differential significance analysis includes comparative analysis, hypothesis testing processing, and between-group mean difference analysis; Perform p-value to logarithm conversion processing according to the regional differential significance p-value data and perform product calculation processing based on the regional methylation stability index, so as to obtain cancer-specific score data; Perform descending order arrangement processing on all regions according to the cancer-specific score data, and select the top 5% regions, so as to obtain candidate specific region data; Perform logistic regression analysis based on L1 regularization on the candidate specific region data, so as to obtain regional discriminative power score data; Perform high discriminative power region screening processing according to the regional discriminative power score data, and select 1000 regions with high discriminative power, so as to obtain specific methylation region data.

9. The prediction method of cancer type based on methylation data according to claim 8, wherein, Step S4 includes the following steps: Step S41: Perform feature vector extraction processing on the specific methylation region data, and input the feature vector into the region encoder of the pre-trained bidirectional attention deep learning model for feature encoding processing, so as to obtain region latent representation data; Step S42: Calculate the correlation weight matrix between specific regions on different chromosomes through the chromosome attention module in the bidirectional attention deep learning model according to the region latent representation data, so as to obtain chromosome attention output data; Step S43: Perform aggregation processing based on the epigenomic environment type according to the region latent representation data through the epigenotype attention module, and calculate the correlation weight matrix between different epigenomic environment types, so as to obtain epigenotype attention output data; Step S44: Perform weighted fusion processing on the chromosome attention output data and the epigenotype attention output data, so as to obtain global epigenomic feature representation data; Step S45: Use the softmax classification layer to perform multi-classification processing on the global epigenomic feature representation data, so as to output the cancer type prediction result.

10. A prediction system for cancer types based on methylation data, characterized in that, A prediction system for a cancer type based on methylation data for performing the prediction method of a cancer type based on methylation data as described in claim 1, the prediction system for a cancer type based on methylation data comprising: A functional region stratification module, configured to obtain the whole-genome methylation data of a sample; perform stratification processing on the whole-genome methylation data based on epigenomic functional regions to obtain functional region stratification data; A local topology construction module, configured to construct a local methylation topology pattern within each functional region according to the functional region stratification data to obtain a local methylation topology representation, wherein the local methylation topology pattern includes a methylation spatial gradient, a methylation curvature feature, and a regional methylation entropy; A stability mapping and screening module, configured to map the coefficient of variation of methylation patterns between different cell types according to the local methylation topology representation, and calculate a regional methylation stability index; perform specific methylation region screening based on the difference significance p-value between cancer and normal samples according to the regional methylation stability index to obtain specific methylation region data; A bidirectional attention classification module, configured to perform classification learning on the specific methylation region data by using a pre-trained bidirectional attention deep learning model, so as to output a cancer type prediction result, wherein the bidirectional attention deep learning model includes a region encoder, a chromosome attention module, and an epigenetic type attention module.

Citation Information

Cited By

  • Methylation signal detection method, device and equipment based on free DNA sequencing fragment in plasma and storage medium

    CN121354681A