Methylation data processing and analysis method and platform, storage medium, and program product
By providing a methylation data processing and analysis method and platform, and integrating task generation, data processing, visualization, and report generation modules, the system solves the problems of low automation in methylation data analysis and insufficient support for differential analysis of multiple groups of samples in existing technologies, and achieves efficient and convenient in-depth data analysis and results presentation.
Patent Information
- Application Number
- PCT/CN2024/082205
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-25
AI Technical Summary
Existing methylation data processing and analysis methods lack a one-stop platform, making it difficult to conduct in-depth data analysis and result presentation efficiently and automatically, especially when operated by non-professionals, and lack support for differential analysis between multiple groups of samples.
A methylation data processing and analysis method and platform is provided, including task generation, data processing, visualization and report generation modules. It generates task scripts by receiving sample information, performs sequencing data quality control, alignment and analysis, generates interactive reports, supports sample methylation, inter-group methylation and differential methylation analysis, and provides multi-dimensional data display.
It achieves highly automated methylation data processing and analysis, can generate in-depth analysis reports in a short time, supports analysis of differences between multiple groups, and outputs interactive reports for easy understanding, making it suitable for non-professional operations.
Smart Images

Figure CN2024082205_25092025_PF_FP_ABST
Abstract
Description
Methylation data processing and analysis method and platform, storage medium and program product Technical Field
[0001] The embodiments of the present disclosure relate to, but are not limited to, the field of biotechnology, and in particular to a methylation data processing and analysis method and platform, a storage medium, and a program product. Background Art
[0002] DNA methylation is a type of genomic epigenetic modification. There is increasing evidence that methylation is involved in the development and progression of diseases and the regulation of related biological pathways. Methylation sequencing is a type of high-throughput sequencing and is the original source of information explaining the involvement of DNA methylation in biological functions. In order to effectively and easily perform in-depth data analysis and display results on methylation, the role of a one-stop platform that integrates methylation data processing, analysis, and visualization is particularly important.
[0003] Summary of the Invention
[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0005] The present disclosure provides a method for processing and analyzing methylation data, including:
[0006] Receive multiple sample information and generate a task script based on the received sample information, wherein the multiple sample information includes: sample queue information, queue comparison information, and sample methylation sequencing data, wherein the sample queue information includes samples and queues to which the samples belong, and the queue comparison information includes one or more queues to be compared, wherein one queue includes two queues;
[0007] According to the generated task script, data processing and data analysis are performed. The data processing includes sequencing data quality control, sequencing data alignment, and alignment result statistics. The data analysis includes sample methylation analysis, inter-group methylation analysis, and differential methylation analysis. The sample methylation analysis is used to analyze the methylation level of each sample. The inter-group methylation analysis is used to compare the methylation levels between samples of each different cohort. The differential methylation analysis is used to identify differentially methylated regions between each different cohort and perform multi-dimensional analysis on them.
[0008] generating graphs and / or tables based on the results of the data processing and data analysis;
[0009] Output interactive reports based on generated graphs and / or tables.
[0010] An embodiment of the present disclosure also provides a methylation data processing and analysis platform, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the methylation data processing and analysis method described in any embodiment of the present disclosure based on the instructions stored in the memory.
[0011] The embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methylation data processing and analysis method described in any embodiment of the present disclosure.
[0012] An embodiment of the present disclosure further provides a program product, comprising instructions. When the computer program product is executed by a computer, the instructions execute the methylation data processing and analysis method as described in any embodiment of the present disclosure.
[0013] The present disclosure also provides a methylation data processing and analysis platform, including: a task generation module, a data processing and analysis module, a visualization module, and a report generation module, wherein:
[0014] The task generation module is configured to receive a plurality of sample information and generate a task script based on the received sample information, wherein the plurality of sample information includes: sample queue information, queue comparison information, and sample methylation sequencing data, wherein the sample queue information includes samples and queues to which the samples belong, and the queue comparison information includes one or more queues to be compared, wherein a queue includes two queues;
[0015] The data processing and analysis module is configured to perform data processing and data analysis according to the generated task script. The data processing includes sequencing data quality control, sequencing data alignment, and alignment result statistics performed in sequence. The data analysis includes sample methylation analysis, inter-group methylation analysis, and differential methylation analysis. The sample methylation analysis is used to analyze the methylation level of each sample. The inter-group methylation analysis is used to compare the methylation levels between samples of each different cohort. The differential methylation analysis is used to identify differentially methylated regions between each different cohort and perform multi-dimensional analysis on them.
[0016] The visualization module is configured to generate graphics and / or tables based on the results of data processing and data analysis by the data analysis module;
[0017] The report generation module is configured to output an interactive report based on the graphs and / or tables generated by the visualization module.
[0018] Other aspects will become apparent upon reading and understanding the drawings and detailed description.
[0019] Summary of the Figures
[0020] The accompanying drawings are intended to provide a further understanding of the technical solutions of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation of the technical solutions of the present disclosure. The shapes and sizes of the components in the drawings do not reflect the actual scale and are intended only to illustrate the contents of the present disclosure.
[0021] FIG1 is a schematic diagram of a process of a methylation data processing and analysis method provided by an exemplary embodiment of the present disclosure;
[0022] FIG2 is a schematic flow chart of another methylation data processing and analysis method provided by an exemplary embodiment of the present disclosure;
[0023] FIG3 is a schematic structural diagram of a methylation data processing and analysis platform provided by an exemplary embodiment of the present disclosure;
[0024] FIG4 is a schematic diagram of an input of a task generation module provided by an exemplary embodiment of the present disclosure;
[0025] FIG5 is a schematic diagram of a processing flow included in a data processing and analysis module provided by an exemplary embodiment of the present disclosure;
[0026] FIG6 is a schematic diagram of a processing flow included in a visualization module provided by an exemplary embodiment of the present disclosure;
[0027] FIG7 is a schematic structural diagram of a report generation module provided by an exemplary embodiment of the present disclosure;
[0028] FIG8 is a schematic structural diagram of another methylation data processing and analysis platform provided by an exemplary embodiment of the present disclosure.
[0029] Details
[0030] To make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other in any manner.
[0031] Unless otherwise defined, the technical or scientific terms used in the embodiments of the present disclosure should have the ordinary meaning understood by people with ordinary skills in the field to which the present disclosure belongs. The words "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprising" mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0032] As shown in FIG1 , an embodiment of the present disclosure provides a methylation data processing and analysis method, comprising:
[0033] Step 101: Receive multiple sample information and generate a task script based on the received sample information. The multiple sample information includes: sample queue information, queue comparison information, and sample methylation sequencing data. The sample queue information includes the sample and the queue to which the sample belongs. The queue comparison information includes one or more queues to be compared, and one queue includes two queues.
[0034] Step 102: Perform data processing and data analysis according to the generated task script. Data processing includes sequencing data quality control, sequencing data alignment, and alignment result statistics. Data analysis includes sample methylation analysis, inter-group methylation analysis, and differential methylation analysis. Sample methylation analysis is used to analyze the methylation level of each sample. Inter-group methylation analysis is used to compare the methylation levels between samples in different cohorts. Differential methylation analysis is used to identify differentially methylated regions between different cohorts and perform multi-dimensional analysis on them.
[0035] Step 103: Generate graphs and / or tables based on the results of data processing and data analysis;
[0036] Step 104: Output an interactive report based on the generated graph and / or table.
[0037] The methylation data processing and analysis method provided by the present disclosure integrates methylation data processing, analysis, and visualization (i.e., graphical and / or tabular display). By simply inputting multiple sample information, a complete interactive report for interpreting the methylation sequencing data analysis results can be produced in one stop. It has a high degree of automation, is convenient for non-professionals to operate, and can produce an in-depth analysis report for sample methylation sequencing data in a short period of time. It supports multi-group difference analysis based on the provided group information. In addition, the interactive report output by the embodiment of the present disclosure has a strong sample size capacity, including graphical and tabular interactions, which is easy for readers to view and understand, and supports inserting and removing content for expansion and deletion.
[0038] In the disclosed embodiments, sample cohort information is used to determine the cohort to which a sample belongs, and cohort comparison information is used to provide any one or more cohorts for which a difference comparison is required. Each cohort contains two cohorts, and each cohort contains multiple samples. For each sample, there will be corresponding sample methylation sequencing data, which is used as a source for downstream data processing and data analysis. When the interactive report is subsequently output, the sample number, cohort number, and comparison scheme number (i.e., the two cohort numbers being compared) will be used as identifiers.
[0039] In some exemplary embodiments, the one or more groups of queues to be compared include at least one of the following:
[0040] Complete response (CR) cohort and partial response (PR) cohort;
[0041] progressive disease (PD) cohort and stable disease (SD) cohort;
[0042] Disease remission cohort and disease non-remission cohort, among which the disease remission cohort includes complete remission cohort and partial remission cohort, and the disease non-remission cohort includes stable disease cohort and disease progression cohort;
[0043] Patient cohort and normal cohort.
[0044] This disclosure takes the methylation data processing and analysis of the prognosis of clinical tumor patients as an example. According to the changes in the size of the target lesions after treatment, the patients can be divided into four types: CR (complete remission), PR (partial remission), SD (stable), and PD (progression). Assuming that the number of patients of each type is 2, each patient can be numbered as "type-01" and "type-02". For example, patient No. 1 of the CR type is "CR-1". If there are more patients, then by analogy, for example, if there are more than 100 patients or 1000 patients, the numbering starts at 001 or 0001. The type of the patient is the cohort information of the patient. For example, the cohort information of patient "CR-01" and patient "CR-02" are both "CR". Each patient has corresponding sample methylation sequencing data. When receiving the sample methylation sequencing data, the user enters the sample methylation sequencing data path.
[0045] The present disclosure can customize the mining of potential biomarkers between cohorts with obvious or vague clinical characteristics through the cohort comparison setting of inter-group differences. The user can set the corresponding cohort comparison information according to the differences between the two cohorts that he wants to observe, for example, set two comparison scheme numbers: "CR_vs_PR" and "SD_vs_PD", and the setting method of the comparison scheme number is "cohort number 1_vs_cohort number 2". More comparison schemes can also be set according to the specific situation. The designation of the comparison scheme is usually set for the differences in clinical phenotypes, traits, etc., but when the biological information is vague (for example, it is unclear whether a certain feature has an impact on the development of the disease), a comparison scheme can also be set to determine whether there are differences between the two cohorts to mine new potential biomarkers.
[0046] In other exemplary embodiments, the one or more groups of queues to be compared may further include:
[0047] A first patient cohort and a second patient cohort, wherein all samples in the first patient cohort have the first feature, and all samples in the second patient cohort do not have the first feature.
[0048] In the disclosed embodiment, the first feature can be set as needed. For example, if the user wants to study the effect of smoking on the efficacy of anti-tumor treatment, then the first feature can be set to smoking, that is, the samples in the first patient cohort are all smokers, and the samples in the second patient cohort are all non-smokers.
[0049] In the embodiment of the present disclosure, for processing and analyzing methylation data of different species, it is necessary to receive corresponding background information while receiving multiple sample information.
[0050] In some exemplary embodiments, the method further includes: receiving background information, wherein the background information includes at least one of the following: file dependency information, software dependency information, custom script information, and parameter information, the file dependency information includes information about dependent files required for data processing and data analysis, the software dependency information includes information about software required for data processing and data analysis, the custom script information includes intermediate data connection scripts and data statistics scripts required for data processing and data analysis, and the parameter information includes parameters required for data processing and data analysis.
[0051] In the embodiment of the present disclosure, as shown in Figure 2, background information includes file dependency information, software dependency information, custom script information and parameter information, wherein the file dependency information includes but is not limited to the reference genome fasta file, the genome chromosome length file, the genome annotation file gff3, the genome gene interval bed file, the methylation CpG island file, the gene transcription start site file, and the result output path; the software dependency information mainly includes but is not limited to sequencing data quality control software, methylation sequencing data comparison software, differential methylation region identification software, Python and R programming software; the custom script information includes but is not limited to intermediate data connection scripts and data statistics scripts; the parameter information includes various parameters in the customizable data analysis process, including but not limited to the number of process running cores, the difference identification threshold, etc.
[0052] Taking human clinical samples as an example, the hg19 or hg38 version of the fasta file is used as the reference genome. The reference genome file can be downloaded from websites such as UCSC, NCBI, and genecode. The fasta file is indexed using samtools, and the reference genome chromosome length file can be generated based on the constructed fasta.fai file. The human genome annotation file gff3, the human genome gene interval bed file, the human methylation CpG island file, and the human gene transcription start site file can also be downloaded from one or more of the above websites.
[0053] Software dependency information can be dynamically adjusted according to different application scenarios. In this embodiment, quality detection can be performed using fastqc software, data cleaning can be performed using trim_galore software, alignment software can use bismark software, alignment and multiple index statistics can be performed using software such as picard, samtools, seqkit and bedtools, and difference identification software can use metilene software.
[0054] Custom script information is mainly used for the format connection of input and output of software analysis results, as well as data statistics of various indicators. Custom scripts can be dynamically added according to specific needs.
[0055] Parameter information is used to dynamically adjust relevant parameters during the analysis process, without requiring modification from the original command end. In some exemplary embodiments, the parameter information may include: the number of server cores, the differential DMR identification threshold, the minimum number of CpG sites constituting a DMR region, and the minimum average methylation difference. For example, the number of server cores can be defined as 5, the differential DMR identification threshold as 0.05, the minimum number of CpG sites constituting a DMR region as 4, and the minimum average methylation difference as 10.
[0056] The technical solution disclosed in the present invention is applicable to the data processing and analysis of the vast majority of species for which reference genomes exist. Exemplarily, when the present invention processes and analyzes methylation data for different species, it is only necessary to provide multiple sample information and file dependency information in the background information (including reference genome, genome annotation file gff3, genome gene interval bed file, methylation CpG island file, gene transcription start site file, etc.), and one-click data processing, analysis and visualization can be performed. When analyzing the same species, since the file dependency information content in the background information used is the same, the file dependency information of a specific species only needs to be introduced once.
[0057] In the disclosed embodiment, as shown in Figure 2, the data processing and analysis process includes: sequencing data quality control process, sequencing data alignment process, alignment result statistics process, sample methylation analysis process, intergroup methylation analysis process, and differential methylation analysis process. The logical relationship between sequencing data quality control, sequencing data alignment, and alignment result statistics is sequential and progressive, and sample methylation analysis, intergroup methylation analysis, and differential methylation analysis can be performed in series or in parallel.
[0058] Specifically, the original sample methylation sequencing data is input into the sequencing data quality control process to obtain the cleaned sample methylation sequencing data, the cleaned sample methylation sequencing data is input into the sequencing data comparison process to obtain the comparison results, and the comparison results are statistically analyzed to obtain the methylation level of the sample.
[0059] In some exemplary embodiments, sequencing data quality control includes: performing quality detection and data cleaning on sample methylation sequencing data, and performing quality detection again on the cleaned data; wherein, quality detection includes: sequence base quality detection, sequence repetition level detection, and sequence GC content detection, and data cleaning includes quality trimming, 3' end trimming, adapter sequence removal, polymorphic nucleotide removal, and short sequence filtering.
[0060] In the disclosed embodiments, sequencing data quality control includes quality inspection and data cleaning of sample methylation sequencing data, and the contents of quality inspection include sequence base quality, sequence repetition level, sequence GC content, total number of reads, total number of bases, Q20, Q30 content, etc. The contents of data cleaning include quality trimming (trimming off bases with lower quality), 3' end trimming (removing low-quality regions at the 3' end of sequencing reads), adapter sequence removal (removing adapter sequences in the data), short sequence filtering (filtering out shorter sequencing sequences), etc. After data cleaning, the cleaned data will be quality inspected again, and the content of quality inspection is the same as above. When the results of the re-quality inspection do not meet the requirements in multiple dimensions, it is recommended to re-experiment or re-sequence the sample to ensure the reliability and stability of the sequence entering the subsequent data analysis.
[0061] In this disclosed embodiment, sequencing data alignment utilizes cleaned sample methylation sequencing data aligned to the human reference genome, and sequence deduplication is performed by detecting fragment start and length information. Sequence duplication is caused by the characteristics of PCR amplification or sequencing technology. Removing sequence duplication improves the accuracy of methylation level identification. Sequencing data alignment and deduplication in this embodiment were both performed using Bismark software.
[0062] In some exemplary embodiments, the comparison result statistical process includes comparison data statistics, methylation statistics, data deduplication statistics, methylation statistics after deduplication, M-deviation statistics, depth and uniformity statistics, etc.
[0063] In the disclosed embodiment, the comparison data statistics include: the total number of sequence pairs analyzed, the number of paired-end alignments with unique optimal alignments, the number of paired sequences that were not aligned under any conditions, and the number of sequences for which genomic sequence information could not be extracted;
[0064] The methylation statistics include: the total number of cytosine Cs analyzed, the number of methylated Cs in CpG, the number of methylated Cs in CHG, the number of methylated Cs in CHH, the number of methylated Cs that could not be determined, the number of unmethylated Cs in CpG, the number of unmethylated Cs in CHG, the number of unmethylated Cs in CHH, the number of unmethylated Cs that could not be determined, the percentage of methylated Cs in CpG, the percentage of methylated Cs in CHG, the percentage of methylated Cs in CHH, and the percentage of methylated Cs that could not be determined;
[0065] The data deduplication statistics include: the number of sequences considered for alignment during the analysis, the number of duplicate aligned sequences removed, the number of sequences with duplicate positions (coordinates) in the alignment results, the number of unique aligned sequences in the alignment results, the proportion of sequences remaining after deduplication, and the proportion of sequences deleted;
[0066] The content of methylation statistics after deduplication is the same as that of methylation statistics;
[0067] The M-deviation statistic is used to display the deviation of methylation levels of R1 and R2 by sequence position in the sequencing results;
[0068] Depth and uniformity statistics include: 0.2x uniformity, the total number of sequencing data bases that fall within the target region (the target region of probe design), the percentage of bases that fall within the target region, the average number of times each base is read during sequencing, the number of all reads (sequencing fragments), the number of times the sequencing data coverage is less than 80% of the expected coverage, and the percentage of bases in the target region with at least 1x coverage.
[0069] In the disclosed embodiment, the comparison result statistics include comparison data statistics, methylation statistics, data deduplication statistics, methylation statistics (after deduplication), M-deviation statistics, depth and uniformity statistics, etc., wherein the comparison data statistics, methylation statistics, data deduplication statistics, and methylation statistics after deduplication are in a serial relationship, and M-deviation statistics, depth and uniformity statistics can be in a serial or parallel relationship. Among them, the data deduplication statistics part can be personalized according to whether the sequencing data contains UMI, thereby further improving the accuracy of methylation level detection. For example, the platform can provide an option for the user to select whether the sequencing data contains UMI. When the user selects that the sequencing data does not contain UMI, the data deduplication statistics process performs deduplication by detecting the fragment start and length information; when the user selects that the sequencing data contains UMI, the data deduplication statistics process performs deduplication by detecting the fragment start, length information, and UMI sequence. The deduplication conditions are more stringent to avoid data loss and reduced analysis accuracy caused by excessive deduplication. M-deviation statistics, depth and uniformity statistics further improve the multi-dimensional control of the comparison result file, which can assist the process to automatically determine whether the sample needs to be included in the subsequent differential methylation analysis process, and provide more stable analysis results. For example, it can be determined whether the M-deviation, depth and uniformity reach the preset threshold to determine whether the subsequent differential methylation analysis process is carried out. For example, it can be set that the 0.2x uniformity must be greater than 95%, the average number of times each base is read in sequencing must be greater than 20X, the multiple of the sequencing data coverage lower than the expected coverage of 80% must be less than 2.0, and the percentage of bases with at least 1x coverage in the target region must be greater than 95%. Only when these conditions are met can the subsequent differential methylation analysis process be carried out.
[0070] In some exemplary embodiments, sample methylation analysis, inter-group methylation analysis, and differential methylation analysis are performed in parallel to speed up methylation data processing and analysis.
[0071] In some exemplary embodiments, the sample methylation analysis includes at least one of the following: genomic chromosome methylation level statistics, genomic functional region methylation level statistics, and gene body methylation level statistics, wherein the genomic chromosome methylation level statistics use preset windows to divide the genome and count the genomic GC content, gene content, and methylation level under each window; the genomic functional region methylation level statistics are performed for different genomic elements according to CpGi, enhancers, exons, introns, promoters, 3'UTRs, and 5'UTRs; the gene body methylation level statistics are performed for the preset regions upstream of the transcription start sites of all genes, exons, and introns, and each region of each gene is divided into equal parts according to length, and the CpG methylation level of each equally divided region is calculated.
[0072] In the disclosed embodiment, the sample methylation analysis process includes genome chromosome methylation level statistics, genome functional region methylation level statistics, and gene body methylation level statistics. Wherein, genome chromosome methylation level statistics are represented by circos diagram, using a 500kb window to divide the genome, and the genome GC content, gene content, and methylation level (CpG) under each window are counted. The Circos diagram is sequentially represented by chromosome number and length, GC content histogram, gene content histogram, and methylation level heat map from outside to inside, wherein the methylation level heat map color indicates that the methylation level is from low to high from light to dark. Genome functional region methylation level statistics are represented by violin diagram, and for different genome elements: CpG methylation level statistics are performed according to CpGi, enhancer, exon, intron, promoter, 3'UTR, and 5'UTR respectively. Gene body methylation levels are presented using a scatter plot. CpG methylation levels are calculated for all genes within 2 kbp upstream of the transcription start site, in exons, and introns. Each region of each gene is divided into 20 bins by length, and the CpG methylation level is calculated for each bin of each region of each gene.
[0073] In some exemplary embodiments, the inter-group methylation analysis includes at least one of the following: inter-group methylation level comparison of genomic functional regions, genomic functional region methylation level correlation analysis, and genomic functional region methylation principal component cluster analysis, wherein the inter-group methylation level comparison of genomic functional regions performs CpG methylation level statistics on all enhancers, preset regions upstream of gene transcription start sites, promoters, 5'UTRs, exons, introns, 3'UTRs and CpGi regions of different groups, and divides each region of each gene into equal lengths, and calculates the CpG methylation level of each equally divided region; the genomic functional region methylation level correlation analysis draws a correlation heat map for functional elements of different samples; and the genomic functional region methylation principal component cluster analysis performs principal component analysis on functional elements of different samples.
[0074] In the disclosed embodiment, the inter-group methylation analysis process includes inter-group genomic functional region methylation level comparison, genomic functional region methylation level correlation analysis, and genomic functional region methylation principal component cluster analysis. Wherein, the inter-group genomic functional region methylation level comparison is represented by a scatter plot, and CpG methylation level statistics are performed on all enhancers (Enhancer), 2Kbp (up2k) upstream of the gene transcription start site, promoters (Promoter), 5'UTR, exons (Exon), introns (Intron), 3'UTR, and CpGi regions of different groups, and each region of each gene is divided into 20 bins according to length, and the CpG methylation level of each bin in each region of each gene is calculated. The genomic functional region methylation level correlation analysis is represented by a cluster heat map, and a correlation heat map is drawn for the functional elements of different samples, wherein the darker the block color, the higher the correlation. Principal component cluster analysis of methylation in functional genomic regions is represented by a two-dimensional PCA scatter plot. The first principal component (PC1) is the direction with the largest variance, and the second principal component (PC2) is perpendicular to PC1 and has the second largest variance. Points closer to PC1 and PC2 represent samples with higher correlation.
[0075] In some exemplary embodiments, the differential methylation analysis includes at least one of the following: identification of differentially methylated regions (DMRs), significant distribution of identification of hypermethylated and hypomethylated regions, distribution of DMR lengths and levels, clustering of DMR regional methylation levels, DMR annotation, distribution of DMR coverage regions, and enrichment analysis of DMR-associated genes.
[0076] The differential methylation analysis process of the embodiment of the present disclosure mainly performs a comprehensive analysis of DMRs for the set differential comparison scheme. After the analysis process is completed, a table file of DMRs and their enrichment analysis can be provided for download. Among them, the differential methylation region (DMR) identification statistics the number of DMRs under each cohort comparison scheme, the average length of DMRs, the average number of CpG sites covered by DMRs, the number of high-methylated DMRs in the latter relative to the former in the comparison scheme number, the number of low-methylated DMRs in the latter relative to the former in the comparison scheme number, and the chromosome number of the DMR under each comparison scheme, DMR start site, DMR end site, P value, differential methylation level, number of CpG sites included, average methylation rate of the latter in the comparison scheme number, and average methylation rate of the former in the comparison scheme number. The genome was divided into 500 kb windows. Within each window, the significance of Hyper DMRs, genomic GC content, gene content, and Hypo DMRs was calculated and plotted as circos diagrams. The circles in the circos diagram represent, from the outermost to the innermost circles, chromosome number and length, Hyper DMR significance (the more outward, the stronger the significance), genomic GC content heatmap, gene content heatmap, and Hypo DMR significance (the more inward, the stronger the significance). Heatmap colors range from light to dark, indicating the correlation value from low to high. DMR length distribution was presented using a bar chart, with the horizontal axis representing DMR length and the vertical axis representing the number of DMRs of a specific length. Methylation levels of DMR regions across different groups were visualized using violin plots, categorized as hypermethylated or hypomethylated DMRs. Methylation levels of DMR regions were clustered using heatmaps based on the average methylation levels between groups, demonstrating differential methylation between groups. DMR annotation software was used to obtain information such as the location of the gene closest to the DMR, its positive and negative strands, gene length, gene ID, distance from the DMR to the transcription start site (TSS), and gene name. A circular plot was used to calculate the percentage of DMRs covering different genomic elements. DMR-associated gene enrichment analysis included GO and KEGG enrichment analysis, with separate GO and KEGG analyses performed for promoter regions with clear biological significance for methylation. Results included enriched pathway identifiers, pathway descriptions, the proportion of genes with pathway function within the entire background gene set, P-values, and the number of DMR-associated genes enriched in that pathway.
[0077] The disclosed embodiments use a differential methylation analysis process to identify differentially methylated regions between each group of different cohorts. The identification method is cohort-to-cohort and many-to-many. There is no need to repeatedly set up differential comparison schemes. Differential analysis can be specified between different cohorts with less obvious clinical significance, and potential biomarkers can be customized to be mined.
[0078] In the disclosed embodiment, as shown in Figure 2, visualization includes graph generation and table generation processes to generate graphs and tables for the data generated during the data processing and analysis processes. Graph generation includes, but is not limited to, box plots, line graphs, curve graphs, pie charts, donut charts, bar charts, circos charts, violin plots, cluster heat maps, and scatter plots. Table generation is based on the provided sample cohort information and cohort comparison information, and is output sequentially numbered by sample or cohort name, facilitating cross-comparison and understanding of samples and cohorts.
[0079] In some exemplary embodiments, outputting an interactive report based on the generated graph and / or table includes:
[0080] Establishing a browsing directory index, which corresponds to the process of data processing and data analysis;
[0081] Generate text content based on the results of data processing and data analysis;
[0082] Output the text content and corresponding graphics and / or tables to the corresponding browsing directory index.
[0083] The interactive report of the embodiment of the present disclosure includes a multi-level browsing directory index. When the user clicks on the multi-level title, he can quickly jump to the corresponding content; the text content can be dynamically updated according to the analysis results, and can also be updated in real time according to user input; the user can perform interactive operations such as dragging, zooming, selecting, searching, sorting, and downloading on graphics and tables.
[0084] In some exemplary embodiments, the method further comprises:
[0085] In response to the user's interactive operation, at least one of the following operations is performed on the graph and / or table: dragging, scaling, searching, sorting, selecting, unselecting, and downloading.
[0086] In the disclosed embodiment, in order to avoid the inability to effectively understand the results due to excessive content displayed in some graphics, zoom functions, frame selection and magnification functions, selection and undo functions, etc. are implemented; the user can interactively adjust the number of rows displayed per page according to the number of rows in the table, and set a search window to quickly locate the results. When the number of rows and columns of results is large, the button can be clicked to download the table in csv or xlsx format for external viewing; in addition, in order to save time in manually changing each part of the text content each time, the text content is passed as a parameter, and the text content of each report is dynamically modified.
[0087] In some exemplary embodiments, the method further comprises:
[0088] According to the generated task script, the classification model is constructed and multiple sample information is split into training sets and test sets;
[0089] Use the training set to train the constructed classification model. During the training process, traverse multiple classification algorithms and test the accuracy of different classification algorithms on the test set. Select the classification model corresponding to the classification algorithm with the highest accuracy on the test set and save and display it.
[0090] In the disclosed embodiment, the sample data of two different queues can be split into training sets and test sets in a hierarchical splitting manner to ensure that the label column ratio of the data set before and after the split is consistent, and the training set is used to traverse most of the classification algorithms to automatically fit the model, and the model with the highest accuracy on the test set after modeling is selected for saving. The classification algorithm can include logistic regression, multi-layer perceptron (MLP Classifier), K nearest neighbor (K neighbors Classifier), support vector machine (SVC), Gaussian process (Gaussian Process Classifier), decision tree (Decision Tree Classifier), Gaussian naive Bayes (Gaussian NB), random forest (Random Forest Classifier), discriminant analysis (Discriminant Analysis), multi-layer perceptron (MLP Classifier), etc. After the training is completed, the ROC curve graph and PR curve graph of the model can be displayed, as well as the specificity, sensitivity, accuracy and AUC result table of the model on the training set and test set.
[0091] Because the amount of data determines the lower limit of machine learning model performance, when building a classification model, the more samples from the two cohorts to be compared, the better. The number of samples in each cohort is not recommended to be less than 20.
[0092] In some exemplary embodiments, when inputting cohort comparison information, depending on the two cohorts for which classification models need to be constructed, the provided comparison scheme number can be appended with "_ML" after the string of the "XX_vs_XX" numbering method, i.e., "XX_vs_XX_ML", for example, "CR_vs_PR_ML" and "SD_vs_PD_ML", and classification models will be constructed between CR and PR, and SD and PD, respectively.
[0093] As shown in FIG3 , the embodiment of the present disclosure further provides a methylation data processing and analysis platform, including a task generation module 301 , a data processing and analysis module 302 , a visualization module 303 , and a report generation module 304 , wherein:
[0094] The task generation module 301 is configured to receive multiple sample information and generate a task script based on the received sample information, wherein the multiple sample information includes: sample queue information, queue comparison information, and sample methylation sequencing data, wherein the sample queue information includes the sample and the queue to which the sample belongs, and the queue comparison information includes one or more queues to be compared, where a queue includes two queues;
[0095] The data processing and analysis module 302 is configured to perform data processing and data analysis according to the generated task script, wherein the data processing includes sequencing data quality control, sequencing data alignment, and alignment result statistics, and the data analysis includes sample methylation analysis, inter-group methylation analysis, and differential methylation analysis. The sample methylation analysis is used to analyze the methylation level of each sample, the inter-group methylation analysis is used to compare the methylation levels between samples of different cohorts, and the differential methylation analysis is used to identify differentially methylated regions between different cohorts and perform multi-dimensional analysis on them.
[0096] A visualization module 303 is configured to generate graphs and / or tables based on the results of data processing and data analysis;
[0097] The report generation module 304 is configured to output an interactive report based on the generated graphs and / or tables.
[0098] The methylation data processing and analysis platform of the embodiment of the present disclosure provides a comprehensive platform for integrated methylation data processing, analysis and visualization by highly integrating modules such as task generation, data processing and analysis, visualization, and report generation. It only requires inputting multiple sample information to output a complete interactive report on the interpretation of the methylation sequencing data analysis results in one stop. It has a high degree of automation, is convenient for non-professionals to operate, and can produce in-depth analysis reports on sample methylation sequencing data in a short time, and supports multi-group difference analysis based on the provided group information. In addition, the present disclosure provides sample methylation analysis, inter-group methylation analysis, differential methylation analysis and other content, which is more comprehensive in function. The interactive report output by the embodiment of the present disclosure has a strong sample quantity accommodating capacity, includes graphics and table interactions, is easy for readers to view and understand, and supports insertion and removal of content expansion and deletion.
[0099] In the disclosed embodiments, sample cohort information is used to determine the cohort to which a sample belongs, and cohort comparison information is used to provide any one or more cohorts for which a difference comparison is required. Each cohort includes two cohorts, and each cohort includes multiple samples. For each sample, there is sample methylation sequencing data, which is used for downstream data processing and data analysis. When the interactive report is subsequently output, the sample number, cohort number, and comparison scheme number (i.e., the two cohort numbers being compared) will be used as identifiers.
[0100] In some exemplary embodiments, the task generation module 301 is further configured to: receive background information, wherein the background information includes at least one of the following: file dependency information, software dependency information, custom script information and parameter information, the file dependency information includes information about dependent files required for data processing and data analysis, the software dependency information includes information about software required for data processing and data analysis, the custom script information includes intermediate data connection scripts and data statistics scripts required for data processing and data analysis, and the parameter information includes parameters required for data processing and data analysis.
[0101] In the embodiment of the present disclosure, as shown in Figure 4, the background information includes file dependency information, software dependency information, custom script information and parameter information. The file dependency information includes but is not limited to the reference genome fasta file, the genome chromosome length file, the genome annotation file gff3, the genome gene interval bed file, the methylation CpG island file, the gene transcription start site file, and the result output path; the software dependency information mainly includes but is not limited to sequencing data quality control software, methylation sequencing data comparison software, differential methylation region identification software, Python and R programming software; the custom script information includes but is not limited to intermediate data connection scripts and data statistics scripts; the parameter information includes various parameters in the customizable data analysis process, including but not limited to the number of process running cores, the difference identification threshold, etc.
[0102] As shown in Figure 5, the data analysis module includes the following processing flows: sequencing data quality control flow, sequencing data comparison flow, comparison result statistical flow, sample methylation analysis flow, inter-group methylation analysis flow, and differential methylation analysis flow. Among them, the sequencing data quality control, sequencing data comparison, and comparison result statistical logical relationships are progressive in sequence, and the sample methylation analysis, inter-group methylation analysis, and differential methylation analysis can be performed in series or in parallel. The disclosed embodiment connects the sequencing data quality control flow, sequencing data comparison flow, and comparison result statistical flow of a single sample in series, performs parallel processing on the above-mentioned flow of multiple samples, and then performs tasks such as sample methylation analysis, inter-group methylation analysis, and differential methylation analysis in parallel, which can speed up the processing and analysis of methylation data.
[0103] Specifically, the raw sample methylation sequencing data is input into the sequencing data quality control process to obtain cleaned sample methylation sequencing data. The cleaned sample methylation sequencing data is then input into the sequencing data alignment process to obtain alignment results. The alignment results are statistically analyzed to determine the methylation level of the sample. Sequencing data quality control is performed on both the raw sample methylation sequencing data and the filtered sample methylation sequencing data. Quality control includes quality testing of the sample methylation sequencing data, including but not limited to sequence base quality, sequence duplication level, and sequence GC content. Both the raw sample methylation sequencing data and the filtered sample methylation sequencing data are subject to these tests. Filtering of the raw sample methylation sequencing data includes quality trimming, 3'-end trimming, adapter sequence removal, polymorphic nucleotide (poly-N) removal, and short sequence filtering. The alignment result statistical process includes alignment data statistics, methylation statistics, data deduplication, methylation data deduplication statistics, M-deviation statistics, depth and uniformity statistics, etc. The sample methylation analysis process includes statistics on methylation levels of genomic chromosomes, methylation levels of functional genomic regions, and methylation levels of gene bodies. The intergroup methylation analysis process includes intergroup comparison of methylation levels in functional genomic regions, correlation analysis of methylation levels in functional genomic regions, and principal component cluster analysis of methylation in functional genomic regions. The differential methylation analysis process includes, but is not limited to, identification of differentially methylated regions (DMRs), significant distribution of hypermethylated and hypomethylated regions, DMR length and level distribution, clustering of methylation levels in DMR regions, DMR annotation, distribution of DMR coverage regions, and enrichment analysis of DMR-associated genes.
[0104] As shown in Figure 6 , visualization module 303 generates graphs and tables for the data generated during data analysis. These graphs include, but are not limited to, boxplots, line graphs, curve graphs, pie charts, donut charts, bar charts, circos charts, violin plots, cluster heat maps, and scatter plots. Tables are output sequentially numbered by sample or cohort name based on the provided sample cohort information and cohort comparison information, facilitating comparison and understanding of samples and cohorts.
[0105] As shown in Figure 7, the report generation module 304 includes a directory module, a text content module, a graphic interaction module, and a table interaction module. The directory module allows users to quickly jump to the corresponding content by clicking on multi-level titles; the text content module can be dynamically updated based on analysis results; the graphic interaction module allows for interactive operations such as dragging, zooming, and selecting; and the table interaction module allows for basic functions such as searching and sorting directly on the interactive report, and supports one-click downloading for tables with large rows and columns.
[0106] Taking the methylation sequencing data, analysis, and visualization of clinical tumor patient prognosis as an example, the methylation data processing and analysis platform of the disclosed embodiment includes: a task generation module, a data processing and analysis module, a visualization module, and a report generation module, wherein: the task generation module receives multiple sample information and generates a task script based on the received sample information; the data processing and analysis module performs data processing and analysis based on the generated task script; the visualization module generates graphs and / or tables based on the results of the data processing and analysis; and the report generation module outputs an interactive report based on the generated graphs and / or tables.
[0107] In the disclosed embodiment, the plurality of sample information includes sample cohort information, cohort comparison information, and sample methylation sequencing data.
[0108] Sample cohort information: According to the changes in the size of the target lesions after treatment, patients can be divided into four types: CR (complete remission), PR (partial remission), SD (stable), and PD (progression). In this embodiment, the number of patients of each type is 2, and each patient is numbered as "Type-01" and "Type-02". For example, patient No. 1 of the CR type is "CR-1". If there are more patients, then the same applies. For example, if there are more than 100 patients or 1000 patients, the numbering starts at 001 or 0001. The type of patient is the cohort information of the patient. For example, the cohort numbers of "CR-01" and "CR-02" patients are both "CR". Finally, the sample methylation sequencing data path of each sample is attached.
[0109] Cohort comparison information: Users can manually set the cohort comparison schemes they wish to observe. For example, two scheme numbers can be set: "CR_vs_PR" and "SD_vs_PD." The scheme numbers are set as "cohort number 1_vs_cohort number 2." More schemes can also be set based on specific circumstances. Comparison schemes are typically specified based on clinical phenotypes, traits, and so on. However, when biological information is ambiguous, scheme comparisons can also be set to determine whether differences exist between any two cohorts. This can be used to identify differences between cohorts with unclear clinical characteristics, enabling in-depth biomarker discovery.
[0110] In the disclosed embodiment, when processing and analyzing methylation data from different species, task generation module 301 receives information about multiple samples and also receives corresponding background information. The background information includes at least one of the following: file dependency information, software dependency information, custom script information, and parameter information.
[0111] File dependency information: For human clinical samples, use hg19 or hg38 versions of the fasta file as the reference genome. Reference genome files can be downloaded from websites such as UCSC, NCBI, and GeneCode. Use samtools to index the fasta file, and the reference genome chromosome length file can be generated based on the constructed fasta.fai file. The human genome annotation file gff3, human genome gene interval bed file, human methylation CpG island file, and human gene transcription start site file can also be downloaded from one or more of the above websites.
[0112] Software dependency information can be dynamically adjusted according to different application scenarios. In this embodiment, fastqc is used for quality detection, trim_galore is used for data cleaning, bismark is used for alignment software, picard, samtools, seqkit, and bedtools are used for alignment and multiple indicator statistics, and metilene software is used for difference identification.
[0113] Custom script information is used for the format connection of input and output of software analysis results, as well as content statistics of various indicators, graph drawing, table generation, etc. Custom scripts can be dynamically added according to specific needs.
[0114] The parameter information is used to dynamically adjust relevant parameters during the analysis process without requiring modification from the original command side. In this example, the number of CPU cores is defined as 5, the differential DMR identification threshold is set to 0.05, the minimum number of CpG sites constituting a DMR region is set to 4, and the minimum average methylation difference is set to 10.
[0115] The disclosed embodiments are applicable to data analysis of most species for which reference genomes exist. For different species, it is only necessary to modify the multiple sample information and corresponding background information received by the task generation module 301. Once the above content has been modified, one-click data processing, analysis, and visualization can be performed. When analyzing the same species, since the file dependency information in the background information used is the same, the file dependency information for a specific species only needs to be introduced once.
[0116] In the disclosed embodiments, data processing includes sequencing data quality control, sequencing data alignment, and alignment result statistics; data analysis includes sample methylation analysis, intergroup methylation analysis, and differential methylation analysis. Sequencing data quality control, sequencing data alignment, and alignment result statistics are logically linked in a sequential and progressive manner, and sample methylation analysis, intergroup methylation analysis, and differential methylation analysis can be performed in tandem or in parallel.
[0117] Sequencing data quality control includes quality testing and data cleaning of sample methylation sequencing data. Data quality testing includes sequence base quality, sequence duplication level, sequence GC content, total number of reads, total base number, Q20 and Q30 content, etc. Data cleaning includes trimming low-quality bases, removing low-quality regions at the 3' end of sequencing reads, removing adapter sequences from the data, and filtering out short sequencing reads. After data cleaning, the cleaned data is re-tested with the same quality testing content as above to ensure the reliability and stability of the sequences used in subsequent data analysis.
[0118] Sequencing data alignment utilizes cleaned sample methylation sequencing data to align to the human reference genome, and sequence duplication is removed by detecting fragment start and length information. Sequence duplication is caused by the characteristics of PCR amplification or sequencing technology. Removing sequence duplication improves the accuracy of methylation level identification. Sequencing data alignment and deduplication in this example were both performed by Bismark.
[0119] The comparison result statistics include comparison data statistics, methylation statistics, data deduplication statistics, methylation statistics (after deduplication), M-deviation statistics, depth and uniformity statistics, etc. (Among them, comparison data statistics, methylation statistics, data deduplication statistics, and methylation statistics after deduplication are in series, and M-deviation statistics, depth and uniformity statistics can be in series or parallel.). Alignment statistics include: the total number of sequence pairs analyzed, the number of paired-end alignments with unique optimal alignments, the number of paired sequences that were not aligned under any conditions, and the number of sequences for which genomic sequence information could not be extracted. Methylation statistics include: the total number of cytosines (C) analyzed, the number of methylated Cs in CpGs, the number of methylated Cs in CHGs, the number of methylated Cs in CHHs, the number of undetermined methylated Cs, the number of unmethylated Cs in CpGs, the number of unmethylated Cs in CHGs, the number of unmethylated Cs in CHHs, the number of undetermined unmethylated Cs, the percentage of methylated Cs in CpGs, the percentage of methylated Cs in CHGs, the percentage of methylated Cs in CHHs, and the percentage of undetermined methylated Cs. Data deduplication statistics include: the number of aligned sequences considered during analysis, the number of duplicate alignments removed, the number of sequences with duplicate positions (coordinates) in the alignment results, the number of unique aligned sequences in the alignment results, the proportion of sequences remaining after deduplication, and the proportion of sequences deleted. The methylation statistics (after deduplication) are the same as those in the methylation statistics section. The M-bias statistic displays the deviation in methylation levels by sequence position between R1 and R2 in the sequencing results. Depth and uniformity statistics include: 0.2x uniformity, the total number of sequencing bases falling within the target region (the probe design target), the percentage of bases falling within the target region, the average number of reads per base during sequencing, the number of total reads (sequencing fragments), the number of bases with sequencing data coverage less than 80% of the expected coverage, and the percentage of bases in the target region with at least 1x coverage.
[0120] Sample methylation analysis includes statistics on chromosomal methylation, functional genomic regions, and gene body methylation. Chromosome methylation statistics are presented using circos plots, which divide the genome into 500kb windows and calculate the GC content, gene content, and CpG methylation level within each window. The circles in the circos plot, from the outermost to the innermost circles, represent chromosome number and length, GC content histogram, gene content histogram, and methylation heatmap. The methylation heatmap colors range from light to dark, indicating low to high methylation levels. Functional region methylation statistics are presented using violin plots, with CpG methylation statistics calculated for different genomic elements: CpGi, enhancer, exon, intron, promoter, 3'UTR, and 5'UTR. Gene body methylation levels are presented using a scatter plot. CpG methylation levels are calculated for all genes within the 2K bp upstream of the transcription start site, exons, and introns. Each region of each gene is divided into 20 bins by length, and the CpG methylation level is calculated for each bin of each region of each gene.
[0121] Intergroup methylation analysis included comparison of methylation levels in functional genomic regions between groups, correlation analysis of methylation levels in functional genomic regions, and principal component cluster analysis of methylation in functional genomic regions. The comparison of methylation levels in functional genomic regions between groups was presented using a scatter plot. CpG methylation levels were statistically analyzed for all enhancers, 2K bp upstream of the gene transcription start site (up2k), promoters, 5'UTRs, exons, introns, 3'UTRs, and CpGi regions in different groups. Each region of each gene was divided into 20 bins based on length, and the CpG methylation levels were calculated for each bin in each region of each gene. Correlation analysis of methylation levels in functional genomic regions was presented using a cluster heat map. Correlation heat maps were plotted for functional elements in different samples, with darker blocks indicating higher correlation. Principal component cluster analysis of methylation in functional genomic regions is represented by a two-dimensional PCA scatter plot. The first principal component (PC1) is the direction with the largest variance, and the second principal component (PC2) is perpendicular to PC1 and has the second largest variance. Points closer to PC1 and PC2 represent samples with higher correlation.
[0122] Differential methylation analysis included identification of differentially methylated regions (DMRs), significance distribution of hypermethylated and hypomethylated regions, DMR length and level distribution, clustering of DMR methylation levels, DMR annotation, distribution of DMR coverage, and DMR-associated gene enrichment analysis. The identification of differentially methylated regions (DMRs) included statistics for the number of DMRs under each comparison scheme, average DMR length, average number of CpG sites covered by DMRs, number of hypermethylated DMRs in the latter comparison scheme relative to the former, number of hypomethylated DMRs in the latter comparison scheme relative to the former, as well as the chromosome number of the DMR under each comparison scheme, DMR start site, DMR end site, P value, differential methylation level, number of CpG sites included, average methylation rate in the latter comparison scheme, and average methylation rate in the former comparison scheme. The genome was divided into 500 kb windows. Within each window, the significance of Hyper DMRs, genomic GC content, gene content, and Hypo DMRs was calculated and plotted as circos diagrams. The circles in the circos diagram represent, from the outermost to the innermost circles, chromosome number and length, Hyper DMR significance (the more outward, the stronger the significance), genomic GC content heatmap, gene content heatmap, and Hypo DMR significance (the more inward, the stronger the significance). Heatmap colors range from light to dark, indicating the correlation value from low to high. DMR length distribution was presented using a bar chart, with the horizontal axis representing DMR length and the vertical axis representing the number of DMRs of a specific length. Methylation levels of DMR regions across different groups were visualized using violin plots, categorized as hypermethylated or hypomethylated DMRs. Methylation levels of DMR regions were clustered using heatmaps based on the average methylation levels between groups, demonstrating differential methylation between groups. DMR annotation software was used to obtain information such as the location of the gene closest to the DMR, its positive and negative strands, gene length, gene ID, distance from the DMR to the transcription start site (TSS), and gene name. A circular plot was used to calculate the percentage of DMRs covering different genomic elements. DMR-associated gene enrichment analysis included GO and KEGG enrichment analysis, with separate GO and KEGG analyses performed for promoter regions with clear biological significance for methylation. Results included enriched pathway identifiers, pathway descriptions, the proportion of genes with pathway function within the entire background gene set, P-values, and the number of DMR-associated genes enriched in that pathway.
[0123] The visualization module generates graphs and / or tables based on the results of data processing and data analysis.
[0124] Visualization includes graphics generation process and table generation process. Since the report generation module cannot directly display the graphics and tables generated by the data processing and analysis module, this module connects the graphics and tables generated by the data processing and analysis module with the report generation module to ensure that the report generation module can correctly read a variety of complex pictures and tables. In the embodiment of the present disclosure, the picture and table generation process includes organizing the picture and table paths under each section to the corresponding specific samples and comparison schemes, and marking the picture and table results, realizing carousel viewing of the picture results of multiple samples or comparison schemes, and realizing scrolling and page turning viewing of the table results of multiple samples or comparison schemes.
[0125] The report generation module includes: directory module, text content module, graphic interaction module, and table interaction module. Directory module: In order to make the report more interactive, the identification directory index is categorized according to the results generated by data analysis. Click the multi-level title to quickly jump to the corresponding part of the content for viewing. Text content module: In order to save the time of manually changing the text description of the content of each part of the result each time, the text content is passed as a parameter to dynamically modify the text content of each report. Graphic interaction module: In order to avoid the inability to effectively understand the results due to excessive content displayed in some graphics, the zoom function, frame selection and magnification function, selection and undo function, etc. are implemented. The table interaction module can interactively adjust the number of rows displayed per page according to the number of rows in the table, and set the search window to quickly locate the results. When there are many rows and columns of results, you can click the button to download the table in csv or xlsx format for external viewing.
[0126] The report generation module of the embodiment of the present disclosure sets a directory module, a text content module, a graphic interaction module, and a table interaction module. Compared with static PDF or Word reports, it can use an interactive visual interface to perform customized exploration and analysis of methylation data by setting specific conditions.
[0127] In summary, the disclosed embodiments provide a comprehensive platform for integrated methylation data processing, analysis, and visualization, comprising the following modules: a task generation module, a data analysis module, a visualization module, and a reporting module. The task generation module generates task scripts based on the provided sample and background information. The data analysis module performs processes based on the generated task scripts, including but not limited to sequencing data quality control, sequencing data alignment, alignment result statistics, sample methylation analysis, intergroup methylation analysis, and differential methylation analysis. The visualization module generates graphs and tables based on the results generated by the data analysis module. The reporting module outputs the graphs, tables, and text content generated by the visualization module.
[0128] Since the methylation data processing and analysis platform of the disclosed embodiment builds the underlying framework for each process in a modular manner, it supports the insertion and deletion of modules, thereby conveniently supporting customized expansion. This embodiment still takes the aforementioned clinical tumor sample classification machine learning model construction content as an example, and adds a machine learning modeling module to the methylation data processing and analysis platform to expand the content of inter-group comparative analysis. It should be noted that since the sample size determines the lower limit of the performance of the machine learning model, when using the machine learning modeling module, the number of samples in the two cohorts to be compared should be as large as possible, and the number of samples in each cohort is recommended to be no less than 20 cases.
[0129] In an embodiment of the present disclosure, the method for enabling the machine learning modeling module may be: when inputting queue comparison information, if the provided comparison scheme number includes a preset machine learning keyword (for example, adding "_ML" after the character string of the "XX_vs_XX" numbering method, that is, "XX_vs_XX_ML"), then while processing and analyzing the aforementioned data, the machine learning modeling module is called to perform model construction and display. For example, if the input queue comparison information is: "CR_vs_PR_ML" and "SD_vs_PD_ML", the machine learning modeling module will respectively construct classification models between the two groups of queues, CR and PR, and SD and PD.
[0130] In the embodiment of the present disclosure, the machine learning modeling module constructs a classification model based on the generated task script, and splits multiple sample information into training sets and test sets; the constructed classification model is trained using the training set, and multiple classification algorithms are traversed during the training process and the accuracy of different classification algorithms is tested on the test set, and the classification model corresponding to the classification algorithm with the highest accuracy on the test set is selected and saved.
[0131] In the disclosed embodiment, the machine learning modeling module uses a hierarchical splitting method to split the sample data of two different queues into training sets and test sets, ensuring that the ratio of the label columns of the data sets before and after the split is consistent, and automatically fits the model by traversing most of the classification algorithms using the training set, and selects the model with the highest accuracy on the test set after modeling for saving. The classification algorithm can include logistic regression, multi-layer perceptron (MLP Classifier), K nearest neighbor (K Neighbors Classifier), support vector machine (SVC), Gaussian process (Gaussian Process Classifier), decision tree (Decision Tree Classifier), Gaussian naive Bayes (Gaussian NB), random forest (Random Forest Classifier), discriminant analysis (Discriminant Analysis), multi-layer perceptron (MLP Classifier), etc.
[0132] The report generation module can display the ROC curve and PR curve of the constructed machine learning model, as well as the specificity, sensitivity, accuracy, and AUC result tables of the model on the training set and test set.
[0133] An embodiment of the present disclosure also provides a methylation data processing and analysis platform, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the methylation data processing and analysis method as described in any embodiment of the present disclosure based on the instructions stored in the memory.
[0134] As shown in FIG8 , in one example, a methylation data processing and analysis platform may include: a processor 810, a memory 820, a bus system 830, and a transceiver 840. The processor 810, the memory 820, and the transceiver 840 are connected via the bus system 830. The memory 820 is configured to store instructions, and the processor 810 is configured to execute the instructions stored in the memory 820 to control the transceiver 840 to transmit and receive signals. Specifically, under the control of the processor 810, the transceiver 840 may receive multiple sample information. The processor 810 generates a task script based on the received sample information and performs data processing and data analysis based on the generated task script. The data processing includes sequentially performing sequencing data quality control, sequencing data alignment, and alignment result statistics. The data analysis includes at least one of the following: sample methylation analysis, inter-group methylation analysis, and differential methylation analysis. Graphs and / or tables are generated based on the results of the data processing and analysis. An interactive report is output based on the generated graphs and / or tables.
[0135] It should be understood that the processor 810 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0136] The memory 820 may include a read-only memory and a random access memory, and provides instructions and data to the processor 810. A portion of the memory 820 may also include a non-volatile random access memory. For example, the memory 820 may also store information about the device type.
[0137] In addition to the data bus, the bus system 830 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as the bus system 830 in FIG.
[0138] During implementation, the processing performed by the processing device can be completed by hardware integrated logic circuits in the processor 810 or by instructions in the form of software. That is, the method steps of the embodiment of the present disclosure can be embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 820, and the processor 810 reads the information in the memory 820 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0139] The present disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methylation data processing and analysis method described in any of the embodiments of the present disclosure. The methylation data processing and analysis method driven by the execution of executable instructions is substantially identical to the methylation data processing and analysis method described in the aforementioned embodiments of the present disclosure and is not further described here.
[0140] In some possible embodiments, various aspects of the methylation data processing and analysis method provided by the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the methylation data processing and analysis method according to various exemplary embodiments of the present disclosure described above in this specification. For example, the computer device may execute the methylation data processing and analysis method described in the embodiments of the present disclosure.
[0141] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0142] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0143] It should be noted that the above-described embodiments or implementations are merely illustrative and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions may be made to the forms and details of the implementations without departing from the scope of the present disclosure.
Claims
1. A method for processing and analyzing methylation data, comprising: Receive multiple sample information and generate a task script based on the received sample information, wherein the multiple sample information includes: sample queue information, queue comparison information, and sample methylation sequencing data, wherein the sample queue information includes samples and queues to which the samples belong, and the queue comparison information includes one or more queues to be compared, wherein one queue includes two queues; According to the generated task script, data processing and data analysis are performed. The data processing includes sequencing data quality control, sequencing data alignment, and alignment result statistics. The data analysis includes sample methylation analysis, inter-group methylation analysis, and differential methylation analysis. The sample methylation analysis is used to analyze the methylation level of each sample. The inter-group methylation analysis is used to compare the methylation levels between samples of each different cohort. The differential methylation analysis is used to identify differentially methylated regions between each different cohort and perform multi-dimensional analysis on them. generating graphs and / or tables based on the results of the data processing and data analysis; Output interactive reports based on generated graphs and / or tables.
2. The method according to claim 1, wherein The one or more groups of queues to be compared include at least one of the following: complete response cohort versus partial response cohort; disease progression cohort vs. stable disease cohort; The disease remission cohort and the disease non-remission cohort, the disease remission cohort includes the complete remission cohort and the partial remission cohort, and the disease non-remission cohort includes the stable disease cohort and the disease progression cohort; Patient cohort and normal cohort.
3. The method according to claim 1, further comprising: Receive background information, the background information including at least one of the following: file dependency information, software dependency information, script information and parameter information, the file dependency information including information of dependent files required for data processing and data analysis, the software dependency information including information of software required for data processing and data analysis, the script information including intermediate data connection scripts and data statistics scripts, and the parameter information including parameters required for data processing and data analysis.
4. The method according to claim 1, wherein The sequencing data quality control includes: performing quality inspection and data cleaning on the sample methylation sequencing data, and performing quality inspection on the cleaned data again; The quality detection includes: sequence base quality detection, sequence repeat level detection, sequence GC content detection, and the data cleaning includes quality trimming, 3' end trimming, adapter sequence removal, polymorphic nucleotide removal, and short sequence filtering.
5. The method according to claim 1, wherein The comparison result statistics include comparison data statistics, methylation statistics, data deduplication statistics, methylation statistics after deduplication, M-deviation statistics, depth and uniformity statistics; The content of the comparison data statistics includes: the total number of sequence pairs analyzed, the number of paired-end alignments with unique optimal alignments, the number of paired sequences that were not aligned under any conditions, and the number of sequences for which genomic sequence information could not be extracted; The methylation statistics include: the total number of cytosine Cs analyzed, the number of methylated Cs in CpG, the number of methylated Cs in CHG, the number of methylated Cs in CHH, the number of methylated Cs that cannot be determined, the number of unmethylated Cs in CpG, the number of unmethylated Cs in CHG, the number of unmethylated Cs in CHH, the number of unmethylated Cs that cannot be determined, the percentage of methylated Cs in CpG, the percentage of methylated Cs in CHG, percentage of methylated C in CHH, percentage of methylated C that could not be determined; The data deduplication statistics include: the number of aligned sequences considered during the analysis, the number of duplicate aligned sequences removed, the number of sequences with duplicate positions in the alignment results, the number of unique aligned sequences in the alignment results, the proportion of sequences remaining after deduplication, and the proportion of sequences deleted; The content of the methylation statistics after deduplication is the same as the content of the methylation statistics; The M-deviation statistic is used to display the methylation level deviation of R1 and R2 by sequence position in the sequencing results; The depth and uniformity statistics include: 0.2x uniformity, the total number of sequencing data bases that fall on the target region, the percentage of bases that fall on the target region, the average number of times each base is read in sequencing, the number of all reads, the number of times the sequencing data coverage is less than 80% of the expected coverage, and the percentage of bases in the target region with at least 1x coverage.
6. The method according to claim 1, wherein The sample methylation analysis, inter-group methylation analysis and differential methylation analysis were performed in parallel.
7. The method according to claim 1, wherein The sample methylation analysis includes at least one of the following: genomic chromosome methylation level statistics, genomic functional region methylation level statistics, and gene body methylation level statistics. The genomic chromosome methylation level statistics use preset windows to divide the genome and count the genomic GC content, gene content, and methylation level under each window; the genomic functional region methylation level statistics perform CpG methylation level statistics for different genomic elements according to CpGi, enhancers, exons, introns, promoters, 3'UTRs, and 5'UTRs; the gene body methylation level statistics perform CpG methylation level statistics in preset regions upstream of the transcription start site of all genes, exons, and introns, and divide each region of each gene into equal parts according to length, and calculate the CpG methylation level of each equally divided region.
8. The method according to claim 1, wherein The inter-group methylation analysis includes at least one of the following: inter-group methylation level comparison of genomic functional regions, genomic functional region methylation level correlation analysis, and genomic functional region methylation principal component cluster analysis; the inter-group genomic functional region methylation level comparison performs CpG methylation level statistics on all enhancers, preset regions upstream of gene transcription start sites, promoters, 5'UTRs, exons, introns, 3'UTRs and CpGi regions of different groups, and divides each region of each gene into equal parts according to length, and calculates the CpG methylation level of each equally divided region; the genomic functional region methylation level correlation analysis draws a correlation heat map of functional elements of different samples; the genomic functional region methylation principal component cluster analysis performs principal component analysis on functional elements of different samples.
9. The method according to claim 1, wherein The differential methylation analysis includes at least one of the following: DMR identification of differentially methylated regions, significant distribution of identification of hypermethylated and hypomethylated regions, DMR length and level distribution, DMR regional methylation level clustering, DMR annotation, DMR coverage region distribution, and DMR-related gene enrichment analysis.
10. The method according to claim 1, wherein The interactive report output based on the generated graphs and / or tables includes: Establishing a browsing directory index, wherein the browsing directory index corresponds to the process of data processing and data analysis; generating text content according to the results of the data processing and data analysis; The text content and corresponding graphics and / or tables are output to the corresponding browsing directory index.
11. The method according to claim 10, further comprising: In response to the user's interactive operation, at least one of the following operations is performed on the graph and / or table: dragging, scaling, searching, sorting, selecting, unselecting, and downloading.
12. The method according to claim 1, further comprising: According to the generated task script, a classification model is constructed, and the plurality of sample information is split into a training set and a test set; The constructed classification model is trained using the training set. During the training process, multiple classification algorithms are traversed and the accuracy of different classification algorithms is tested on the test set. The classification model corresponding to the classification algorithm with the highest accuracy on the test set is selected for saving and display.
13. A methylation data processing and analysis platform, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the methylation data processing and analysis method according to any one of claims 1 to 12 based on the instructions stored in the memory. 14 . A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the methylation data processing and analysis method according to claim 1 is implemented. 15 . A computer program product comprising instructions, wherein when the computer program product is executed by a computer, the instructions execute the methylation data processing and analysis method according to claim 1 .
16. A methylation data processing and analysis platform, comprising a task generation module, a data processing and analysis module, a visualization module, and a report generation module, wherein: The task generation module is configured to receive a plurality of sample information and generate a task script based on the received sample information, wherein the plurality of sample information includes: sample queue information, queue comparison information, and sample methylation sequencing data, wherein the sample queue information includes samples and queues to which the samples belong, and the queue comparison information includes one or more queues to be compared, wherein a queue includes two queues; The data processing and analysis module is configured to perform data processing and data analysis according to the generated task script. The data processing includes sequencing data quality control, sequencing data alignment, and alignment result statistics performed in sequence. The data analysis includes sample methylation analysis, inter-group methylation analysis, and differential methylation analysis. The sample methylation analysis is used to analyze the methylation level of each sample. The inter-group methylation analysis is used to compare the methylation levels between samples of each different cohort. The differential methylation analysis is used to identify differentially methylated regions between each different cohort and perform multi-dimensional analysis on them. The visualization module is configured to generate graphics and / or tables based on the results of data processing and data analysis by the data analysis module; The report generation module is configured to output an interactive report based on the graphs and / or tables generated by the visualization module.
Citation Information
Patent Citations
DNA (Deoxyribo-Nucleic Acid) methylation sequencing data calculation and interpretation method
CN107273663A
Biological cloud platform-based methylation data analysis application system
CN107563152A
Biological analysis process of MeRIP-seq high-throughput sequencing data
CN111261229A
Transcriptome and DNA methylation data correlation analysis method and system
CN112201302A
Thyroid cancer diagnosis by DNA methylation analysis
US20180051343A1