Methylation-based tumor data processing system
Through the combination of high-throughput sequencing and microfluidic chip technology, the problem of noise masking in low-concentration DNA samples is solved, and accurate identification and quantitative analysis of tumor-related methylation sites is achieved, reducing the false negative rate and improving the sensitivity and accuracy of detection.
Patent Information
- Application Number
- CN202510781204.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When traditional tumor detection methods treat low-concentration DNA samples, the methylation signal is easily masked by background noise, resulting in a high false negative rate and making it difficult to accurately identify tumor-related methylation changes.
A tumor data processing system based on methylation is adopted, combined with high-throughput sequencing technology and microfluidic chips, efficient separation and enrichment of DNA samples is achieved through enzyme cutting, PCR amplification and bisulfite treatment, and accurate identification and quantitative analysis of methylation sites are performed by combining negative binomial distribution model and convolutional neural network model.
It improves the sensitivity and accuracy of tumor detection, reduces the false negative rate, and realizes in-depth mining and accurate analysis of methylation information in low-content clinical tumor DNA samples, which can more accurately identify tumor-related methylation sites.
Smart Images

Figure CN120299531A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor data processing, and particularly to a tumor data processing system based on methylation. Background Art
[0002] As an important epigenetic modification, DNA methylation plays an increasingly important role in the occurrence, development, diagnosis, and treatment of tumors. DNA methylation refers to the process in which a methyl group is covalently bound to the 5'-carbon atom of cytosine under the action of DNA methyltransferases (DNMTs) to form 5-methylcytosine (5mC). This modification usually occurs in CpG islands (DNA sequences rich in cytosine-guanine dinucleotides) and can regulate gene expression, affect chromatin structure, and participate in various biological processes such as cell differentiation and genomic imprinting. In tumor cells, the DNA methylation pattern changes significantly, showing global genomic hypomethylation and local CpG island hypermethylation. Global genomic hypomethylation leads to genomic instability, activation of oncogenes, etc.; while hypermethylation of CpG islands in the promoter regions of specific genes leads to silencing of tumor suppressor genes and promotes the occurrence of tumors. However, when the content of clinical tumor DNA samples is low, traditional tumor detection data processing methods are prone to submerge the detection signal in background noise, resulting in a high false negative rate, that is, the failure to detect real tumor-related methylation changes. Summary of the Invention
[0003] Based on this, the present invention provides a tumor data processing system based on methylation to solve at least one of the above technical problems.
[0004] To achieve the above object, a tumor data processing system based on methylation includes the following modules: A methylation sample sequencing module, including a high-throughput sequencer, a microfluidic chip, a restriction enzyme kit, and a PCR instrument. The restriction enzyme kit is used to perform restriction enzyme digestion on a clinical tumor DNA sample to be tested and separate it using the microfluidic chip to obtain preliminary DNA sample separation fragments; the high-throughput sequencer is used to perform high-throughput sequencing on the preliminary DNA sample separation fragments to obtain methylation tumor sequencing data; A sequencing difference site identification module, which is used to filter low-quality sequencing data from the methylation tumor sequencing data to generate high-quality sequencing data; quantify the methylation level based on the high-quality sequencing data to generate methylation site frequency data; analyze tumor difference sites based on the methylation site frequency data to generate tumor difference site data; The tumor gene sequence analysis module includes a genomic database, which is used to extract the genomic sequence of the corresponding region according to the tumor differential site data to generate tumor differential region sequence data; identify the key gene sequences related to tumor occurrence based on the tumor differential region sequence data, and perform sequence feature analysis to generate tumor methylation feature data. The tumor subtype classification module is used to construct a tumor risk prediction model; transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction, and perform tumor subtype label identification to generate tumor type label data.
[0005] The present invention uses high-throughput sequencing technology to deeply sequence the sample, which can capture more methylation information, and even weak methylation signals can be effectively detected. The application of the microfluidic chip further improves the efficiency and accuracy of sample processing. By precisely controlling the micro-droplets, the efficient separation of DNA samples is realized, reducing sample loss and cross-contamination. When traditional detection methods process low-concentration DNA samples, the methylation signals are easily masked by background noise, making it difficult to accurately identify tumor-related methylation changes. However, this system can efficiently separate and enrich DNA fragments through the microfluidic chip, combined with high-throughput sequencing technology, which can effectively improve the detection sensitivity. Even when the DNA content in clinical samples is low, weak methylation signals can be captured, thereby reducing the false negative rate and more accurately identifying tumor-related methylation sites. Through the sequencing differential site recognition module, strict quality control and filtering are performed on the sequencing data to exclude the interference caused by low-quality data and ensure the accuracy of subsequent analysis. On this basis, the system quantifies the methylation level of high-quality sequencing data and combines tumor differential site analysis to accurately identify tumor-related differentially methylated sites. This avoids the error caused by judging only based on the methylation status of a single site and improves the reliability of tumor data. By constructing a tumor risk prediction model, tumor subtype classification and risk prediction are realized. Inputting the extracted tumor methylation feature data into a pre-trained tumor risk prediction model can achieve accurate prediction of tumor types and perform tumor subtype label identification, which can more accurately reflect the molecular heterogeneity of tumors and realize in-depth mining and accurate analysis of methylation information in low-content clinical tumor DNA samples. Therefore, a methylation-based tumor data processing system of the present invention precisely separates DNA fragments through microfluidic chip technology and combines high-throughput sequencing to achieve accurate identification and quantitative analysis of methylation differential sites, reduces the false negative rate, and effectively overcomes the problem that signals are easily submerged by noise in traditional methods.
[0006] Preferably, the methylation sample sequencing module includes the following functions: Perform enzymatic digestion on the clinical tumor DNA sample to be tested to obtain a mixture of digested fragments to be tested. Introduce the mixture of enzyme-digested fragments to be tested into a microfluidic chip with a microchannel of a set length, and separate them based on a preset electrophoretic mobility to obtain preliminary DNA sample separation fragments; Perform PCR amplification on the preliminary DNA sample separation fragments to generate a methylated enrichment DNA solution; Based on the methylated enrichment DNA solution, perform high-throughput sequencing to obtain methylated tumor sequencing data.
[0007] Through enzymatic digestion, the present invention can break the clinical tumor DNA sample to be tested into fragments of specific sizes, which lays a foundation for subsequent fragment separation and enrichment. The setting of the length of the microchannel and the preset of the electrophoretic mobility ensure that only DNA fragments with specific sizes and charge characteristics can pass through, thus realizing the preliminary separation of the DNA sample. Performing PCR amplification on the preliminarily separated DNA fragments can significantly increase the number of target DNA fragments, which is crucial for subsequent high-throughput sequencing. Since methylated sites usually only account for a small part of the genome, through PCR amplification, DNA fragments containing methylation information can be effectively enriched, improving the signal-to-noise ratio of sequencing, avoiding the waste of sequencing in a large number of non-target regions, reducing the sequencing cost, and improving the data utilization rate. Unbiased detection of methylated sites across the entire genome realizes the efficient conversion from tumor DNA samples to methylated sequencing data.
[0008] Preferably, the PCR amplification of the preliminary DNA sample separation fragments includes: Mix the preliminary DNA sample separation fragments with a bisulfite solution to obtain a bisulfite reaction mixture; Perform a conversion reaction on the bisulfite reaction mixture and perform desulfonation treatment to obtain a desulfonated DNA solution; Perform magnetic bead adsorption purification on the desulfonated DNA solution to obtain a purified DNA solution; Perform DNA target capture on the purified DNA solution with a preset methylated biotin-labeled probe to obtain a sample-captured DNA solution; Perform PCR amplification on the sample-captured DNA solution to generate a methylated enrichment DNA solution.
[0009] The present invention uses bisulfite treatment to initially separate DNA fragments, which can convert unmethylated cytosine into uracil, while methylated cytosine remains unchanged, thereby achieving specific recognition of methylation sites. Magnetic microspheres are used to adsorb and purify desulfonated DNA, which can efficiently remove residual impurities and unbound bisulfite in the reaction system, thereby improving the efficiency and specificity of subsequent target capture. DNA target capture is performed using a preset methylation biotin-labeled probe, which can specifically enrich the target methylation region, thereby increasing the sequencing depth and sensitivity and reducing the sequencing cost. Through target capture, a large number of unmethylated regions in the genome can be effectively excluded, and sequencing resources can be concentrated on methylation regions related to tumors, thereby more effectively detecting low-abundance methylation sites and improving the ability to analyze the tumor methylation map. After target capture, PCR amplification is performed to further increase the copy number of target methylated DNA, thereby meeting the requirements of high-throughput sequencing. PCR amplification can exponentially amplify the target region, significantly improving the detection sensitivity. Even in the case of a low starting DNA sample amount, sufficient DNA can be obtained for sequencing analysis, thereby ensuring the reliability and accuracy of methylation detection.
[0010] Preferably, the sequencing differential site recognition module includes the following functions: Perform base recognition on methylated tumor sequencing data to obtain sequencing base sequence data; Remove adapter sequences according to the sequencing base sequence data and perform sequencing quality assessment to obtain sequencing quality assessment data; Filter out low-quality sequencing from the methylated tumor sequencing data through the sequencing quality assessment data to generate high-quality sequencing data; Quantify the methylation level according to the high-quality sequencing data to generate methylation site frequency data; Analyze tumor differential sites according to the methylation site frequency data to generate tumor differential site data.
[0011] The present invention converts unmethylated cytosine (C) into uracil (U), while methylated cytosine remains unchanged. This differential treatment lays the foundation for subsequent differentiation between methylated and non-methylated sites. Through subsequent PCR amplification, U will be replaced by thymine (T), so that in the sequencing results, the originally unmethylated C will be shown as T, while methylated C will remain as C. This conversion reaction transforms the originally difficult-to-distinguish methylation status into easily detectable sequence differences. Magnetic microspheres can specifically bind to DNA and separate DNA from other impurities by applying an external magnetic field. Using pre-set methylated biotin-labeled probes for DNA target capture is the core technology for achieving methylation enrichment. These probes can specifically bind to the methylated regions of interest and capture the target DNA fragments through the biotin-streptavidin system. This target capture method can greatly enrich the DNA in the target region, reduce the interference from non-target regions, and improve the targeting and efficiency of sequencing.
[0012] Preferably, the quantification of methylation level according to high-quality sequencing data includes: Performing a transformed sequence-specific genomic alignment on the high-quality sequencing data to obtain methylation sequencing alignment data; Removing repetitive sequences based on the methylation sequencing alignment data to obtain unique sequencing alignment sequence data; Extracting cytosine context sequences based on the unique sequencing alignment sequence data to obtain cytosine site context sequence data; Performing a bisulfite conversion judgment on both strands of the cytosine site context sequence data to generate cytosine site conversion status data; Counting the number of methylated / non-methylated sequencing fragments based on the cytosine site conversion status data for the unique sequencing alignment sequence data to obtain the number of methylated sequencing fragments and the number of non-methylated sequencing fragments respectively; Calculating the methylation frequency based on the number of methylated sequencing fragments and the number of non-methylated sequencing fragments to generate methylation site frequency data.
[0013] By aligning high-quality sequencing data with a reference genome and removing repetitive sequences, the present invention can accurately locate the positions of methylation sites in the genome and exclude the interference of repetitive sequences caused by PCR amplification or sequencing errors, thereby ensuring the accuracy of methylation level quantification. Since untreated cytosine (C) is converted to uracil (U, which ultimately appears as thymine (T) in the sequencing data) after bisulfite treatment, a specialized alignment algorithm is required to accurately align the sequencing data to the reference genome. This specific alignment can distinguish true methylation variations from sequence changes caused by bisulfite conversion. The bisulfite conversion-based judgment can effectively identify and exclude false positive or false negative results caused by incomplete or excessive conversion, thereby improving the accuracy of methylation status determination. By comparing the sequencing information of the forward and reverse strands, it is possible to more reliably determine whether cytosine has been converted, and thus more accurately evaluate the methylation level.
[0014] Preferably, the calculation of methylation frequency based on the number of methylated sequencing fragments and the number of unmethylated sequencing fragments includes: Counting the number of sequencing fragment coverages at the site based on the number of methylated sequencing fragments and the number of unmethylated sequencing fragments to generate the total sequencing depth data of cytosine sites; Calculating the methylation ratio at the site based on the number of methylated sequencing fragments and the number of unmethylated sequencing fragments to generate the initial methylation frequency; Establishing a mapping relationship between the sequencing depth and the methylation frequency, and constructing a frequency background noise model using a preset negative binomial distribution model; Setting a low-depth threshold parameter based on the frequency background noise model; filtering the initial methylation frequency at low sites through the total sequencing depth data of cytosine sites based on the low-depth threshold parameter to obtain filtered methylation frequency data; Calculating the frequency confidence interval of the filtered methylation frequency data using the frequency background noise model to generate methylation confidence interval data; Performing confidence interval correction on the filtered methylation frequency data through the methylation confidence interval data to generate methylation site frequency data.
[0015] By statistically counting the coverage times of sequencing fragments at loci, the present invention can obtain the total sequencing depth data of cytosine loci, establish a mapping relationship between the sequencing depth and the methylation frequency, and use the negative binomial distribution model to construct a frequency background noise model, which can effectively identify and remove the background noise caused by low sequencing depth or random errors, thereby improving the accuracy of the methylation frequency. The negative binomial distribution model can better fit the statistical characteristics of the sequencing data, so as to more accurately estimate the background noise level. Setting a low-depth threshold parameter and performing low-locus frequency filtering can remove the unreliable methylation frequency data caused by insufficient sequencing depth, and further improve the reliability of the results. Using the frequency background noise model to calculate the confidence interval of the filtered methylation frequency data can evaluate the statistical significance of the methylation frequency and provide a more reliable methylation level estimate. The confidence interval can reflect the precision of the methylation frequency estimate. The narrower the confidence interval, the more accurate the estimate. Through confidence interval correction, the methylation frequency can be finely adjusted to obtain more accurate and reliable methylation site frequency data.
[0016] Preferably, the analysis of tumor differential loci according to the methylation site frequency data includes: Dividing the high-quality sequencing data into functional regions by using the preset genomic annotation information to obtain the genomic functional region list data; Dividing the genomic functional region list data into tumor / normal samples by using the methylation site frequency data, and performing locus frequency annotation to generate regional methylation frequency data; Calculating the average methylation levels of tumor samples / normal samples according to the regional methylation frequency data to obtain the tumor region methylation mean and the normal region methylation mean respectively; Calculating the methylation difference of the tumor region methylation mean by using the normal region methylation mean, and performing a significance test of the difference to obtain the regional methylation difference P value; Screening the tumor differential methylation regions from the genomic functional region list data by using the regional methylation difference P value, and performing tumor differential locus annotation to generate tumor differential locus data.
[0017] The present invention divides genomic functional regions into tumor / normal samples based on methylation site frequency data and performs site frequency annotation, enabling comparison of methylation level differences between tumor samples and normal samples in different functional regions, thereby identifying methylation changes related to tumors. By calculating the average methylation levels of tumor samples and normal samples in each functional region respectively, the methylation differences between the two groups of samples can be compared more intuitively, providing a data basis for subsequent statistical analysis. The average methylation level can reflect the overall methylation status of a specific functional region, thus more clearly showing the differences between tumor and normal samples. By calculating the methylation difference between the average methylation of the tumor region and that of the normal region, and performing a significance test for the difference, the statistical significance of the methylation difference between tumor samples and normal samples can be evaluated, and methylation regions with significant differences can be identified. By screening for tumor differentially methylated regions in genomic functional regions based on the P-value of regional methylation differences and performing tumor differential site annotation, the differentially methylated sites related to tumors can be accurately located.
[0018] Preferably, the tumor gene sequence analysis module includes the following functions: Perform genomic coordinate mapping based on tumor differential site data, extract genomic sequences of corresponding regions, and generate tumor differential region sequence data; Construct a differential methylation site expression network based on the tumor differential region sequence data to generate a gene co-expression network; Perform network topology analysis on the gene co-expression network to obtain node centrality indicators; Based on the node centrality indicators, identify key gene sequences for tumorigenesis in the gene co-expression network to obtain tumor recognition gene sequences; Perform sequence feature analysis based on the tumor recognition gene sequences to generate a tumor gene sequence feature vector; Extract the methylation level based on the tumor recognition gene sequences, and perform feature fusion on the tumor gene sequence feature vector to generate tumor methylation feature data.
[0019] By mapping the tumor differential site data to the genomic coordinates and extracting the genomic sequences of the corresponding regions, the tumor differential region sequence data can be obtained. Constructing a differential methylation site expression network can reveal the complex interaction relationships between genes. Conducting network topology analysis on the gene co-expression network and calculating the node centrality index can identify the key genes in the core position of the network. Identifying the key gene sequences in tumorigenesis based on the node centrality index can screen out the genes closely related to tumors from the complex gene network. Conducting sequence feature analysis on the tumor recognition gene sequences and generating tumor gene sequence feature vectors can convert the gene sequence information into feature vectors that can be used in machine learning models. The sequence feature vectors can capture the specific patterns and information of the gene sequences, thereby improving the prediction performance of the model.
[0020] Preferably, the sequence feature analysis according to the tumor recognition gene sequences includes: Calculating the GC content of the tumor recognition gene sequences to obtain the GC content value; Identifying CpG islands in the tumor recognition gene sequences and calculating the CpG island density to generate the CpG island density value; Concatenating the GC content value and the CpG island density value to generate an initial sequence feature vector; Performing regularization processing on the initial sequence feature vector to generate the sequence feature vector; Performing K-Mer decomposition on the tumor differential region sequence data and integrating features according to the sequence feature vector to generate the tumor gene sequence feature vector.
[0021] The present invention calculates the GC content of tumor recognition gene sequences, which is the first step in extracting the basic features of the sequences. The GC content is an important feature of genomic sequences and is closely related to gene stability, expression regulation, etc. The difference in GC content reflects the changes in genes during tumorigenesis. By identifying CpG islands and calculating their density in tumor recognition gene sequences, information on CpG islands closely related to gene regulation can be obtained. CpG islands are regions in the genome rich in CG dinucleotides, usually located in gene promoter regions and are closely related to gene expression regulation. Alterations in the methylation status of CpG islands are often associated with the occurrence and development of tumors. Therefore, CpG island density can be an important feature of tumor gene sequences. CpG islands are regions in the genome rich in CG dinucleotides, usually associated with gene promoter regions and are important targets for methylation regulation. Changes in CpG island density reflect alterations in the gene methylation status. Concatenating the GC content value and the CpG island density value to generate an initial sequence feature vector is the first step in integrating sequence features of different dimensions. This integration combines the overall features (GC content) and local features (CpG island density) of the genome, providing more comprehensive sequence information. Performing K-Mer decomposition on the sequence data of tumor differential regions and integrating features based on the sequence feature vector, K-Mer decomposition divides the sequence into short fragments of length K, which can capture local patterns in the sequence. Integrating K-Mer features with features such as GC content and CpG island density can more comprehensively describe the features of gene sequences and improve the information content of the feature vector.
[0022] Preferably, the tumor subtype classification module includes the following functions: Obtain historical tumor gene sequence feature data; Perform stratified sampling on the historical tumor gene sequence feature data to obtain a training set and a test set respectively; Perform transfer learning on a preset convolutional neural network model using the training set to generate a trained tumor risk prediction model; Perform model cross-validation on the trained tumor risk prediction model using the test set and optimize the parameters to obtain a tumor risk prediction model; Transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction and perform classification probability conversion to generate a tumor classification probability matrix; Identify the maximum probability value based on the tumor classification probability matrix and perform tumor subtype label identification to generate tumor type label data.
[0023] The present invention utilizes existing data resources, which can avoid training a model from scratch, thereby saving time and computing resources. High-quality historical data can improve the training effect and generalization ability of the model. By performing transfer learning on a preset convolutional neural network model using a training set, the knowledge of the pre-trained model can be utilized to accelerate the model training process and improve the model's performance. Transfer learning can transfer the knowledge learned by the pre-trained model on other tasks to the current task, thereby reducing the training time and data requirements. By performing cross-validation and parameter optimization on the trained model using a test set, the generalization ability and robustness of the model can be further improved. Cross-validation can evaluate the performance of the model under different data partitions, and parameter optimization can find the optimal model parameters, thereby enabling the model to achieve the best prediction effect. Inputting tumor methylation feature data into the trained tumor risk prediction model can predict the tumor type and generate a tumor classification probability matrix, and the probability matrix can reflect the prediction confidence of the model for different tumor types. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic diagram of the module process of the tumor data processing system based on methylation of the present invention; Figure 2 is Figure 1 a detailed implementation process schematic diagram of S3 in Figure 3 is Figure 1 a detailed implementation process schematic diagram of S4 in The implementation, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0026] In addition, the drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0027] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.
[0028] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a methylation-based tumor data processing system, including the following modules: The methylation sample sequencing module includes a high-throughput sequencer, a microfluidic chip, a restriction enzyme digestion kit, and a PCR instrument. The restriction enzyme digestion kit is used to perform restriction enzyme digestion on the clinical tumor DNA sample to be tested, and the microfluidic chip is used for separation to obtain preliminary DNA sample separation fragments; the high-throughput sequencer is used to perform high-throughput sequencing on the preliminary DNA sample separation fragments to obtain methylation tumor sequencing data; The sequencing differential site identification module is used to filter out low-quality sequencing from the methylation tumor sequencing data to generate high-quality sequencing data; quantify the methylation level according to the high-quality sequencing data to generate methylation site frequency data; analyze the tumor differential sites according to the methylation site frequency data to generate tumor differential site data; The tumor gene sequence analysis module includes a genomic database, which is used to extract the genomic sequence of the corresponding region according to the tumor differential site data to generate tumor differential region sequence data; identify the key gene sequence of tumorigenesis according to the tumor differential region sequence data and perform sequence feature analysis to generate tumor methylation feature data; The tumor subtype classification module is used to construct a tumor risk prediction model; transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction and perform tumor subtype label identification to generate tumor type label data.
[0029] In the embodiment of the present invention, the methylation-based tumor data processing system includes the following modules: S1: The methylation sample sequencing module includes a high-throughput sequencer, a microfluidic chip, a restriction enzyme digestion kit, and a PCR instrument. The restriction enzyme digestion kit is used to perform restriction enzyme digestion on the clinical tumor DNA sample to be tested, and the microfluidic chip is used for separation to obtain preliminary DNA sample separation fragments; the high-throughput sequencer is used to perform high-throughput sequencing on the preliminary DNA sample separation fragments to obtain methylation tumor sequencing data; In the embodiment of the present invention, 1 μg of clinical tumor DNA sample is obtained. Then, the extracted DNA is digested using a digestion kit. The DNA sample is mixed with 10-fold concentrated CutSmart buffer and enzyme, and incubated at 37°C for 2 hours, and then incubated at 65°C for 20 minutes to inactivate the enzyme. The digested DNA fragment mixture is purified using magnetic beads. The purified DNA fragment mixture is loaded onto a microfluidic chip, which contains a microchannel with a preset length of 10 cm and is filled with a prefabricated colloid. A voltage of 1000 V is applied for 30 minutes for electrophoresis separation, and DNA fragments of 100 - 500 bp are collected. Next, the collected DNA fragments are amplified by PCR using primers with specific adapters and a PCR premix. The PCR reaction program is as follows: pre-denaturation at 95°C for 3 minutes; denaturation at 98°C for 20 seconds, annealing at 60°C for 15 seconds, extension at 72°C for 30 seconds, for 30 cycles; final extension at 72°C for 5 minutes; incubation at 4°C. The purified DNA is loaded onto a high-throughput sequencer for paired-end 150 bp sequencing to generate methylation tumor sequencing data (FASTQ file).
[0030] S2: A sequencing differential site recognition module, which is used to filter out low-quality sequencing data from the methylation tumor sequencing data to generate high-quality sequencing data; quantify the methylation level based on the high-quality sequencing data to generate methylation site frequency data; analyze tumor differential sites based on the methylation site frequency data to generate tumor differential site data; In the embodiment of the present invention, the Cutadapt tool is used to remove the sequencing adapter sequences, and the FastQC tool is used to evaluate the sequencing quality to generate a quality report. According to the FastQC report, the Trimmomatic tool is used to filter out low-quality sequencing reads, such as removing bases with an average quality value lower than 20 and reads with a length less than 50 bp. High-quality sequencing data is obtained after filtering. The Bismark tool is used to align the high-quality sequencing data to a reference genome (e.g., hg38), and calculate the methylation level of each CpG site to generate methylation site frequency data. The methylKit tool kit is used to perform differential methylation analysis on the methylation site frequency data of tumor samples and normal samples, such as using a logistic regression model, and use the false discovery rate (FDR) to control multiple test corrections. Differentially methylated sites with an FDR less than 0.05 are screened to generate tumor differential site data, including the position information of the differentially methylated sites, the difference in methylation level, p-value, FDR value, etc.
[0031] S3: Tumor gene sequence analysis module, including a genomic database, which is used to extract genomic sequences of corresponding regions according to tumor differential site data to generate tumor differential region sequence data; identify key gene sequences related to tumorigenesis based on the tumor differential region sequence data, and perform sequence feature analysis to generate tumor methylation feature data. In the embodiment of the present invention, according to the genomic coordinate information in the tumor differential site data, the Bedtools tool is used to extract the genomic sequences of corresponding regions from the human reference genomic database (for example, GENCODE) to generate tumor differential region sequence data. Based on the tumor differential region sequence data and gene expression data (for example, RNA-seq data), the WGCNA toolkit is used to construct a gene co-expression network. Topological analysis is performed on the gene co-expression network, such as calculating the degree centrality, betweenness centrality, etc. of nodes. Genes with top-ranked centrality metrics are selected as key genes related to tumorigenesis, and their sequences are extracted to generate tumor recognition gene sequences. Sequence feature analysis is performed on the tumor recognition gene sequences, such as calculating the GC content, CpG island density, etc., and these features are integrated with the methylation levels of the key genes to generate tumor methylation feature data.
[0032] S4: Tumor subtype classification module, which is used to construct a tumor risk prediction model; transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction, and perform tumor subtype label identification to generate tumor type label data.
[0033] In the embodiment of the present invention, historical tumor gene sequence feature data, including gene sequence feature vectors of different tumor subtypes and corresponding tumor subtype labels, is obtained from the TCGA database or other public databases. The historical data is divided into a training set and a test set using the stratified sampling method, for example, in a ratio of 8:2. A pre-trained convolutional neural network model (for example, ResNet) is used to perform transfer learning in combination with the training set data to generate a trained tumor risk prediction model. The trained model is cross-validated and parameter-optimized using the test set data, for example, using the Adam optimizer and the cross-entropy loss function, and parameters such as the learning rate and batch size are adjusted. Finally, a tumor risk prediction model is obtained. The new tumor methylation feature data is input into the trained tumor risk prediction model to obtain the prediction probabilities of each tumor subtype, generating a tumor classification probability matrix. For each sample, the subtype with the largest probability value is selected as the predicted subtype to generate tumor type label data.
[0034] Preferably, the methylation sample sequencing module includes the following functions: Perform enzymatic digestion on the clinical tumor DNA sample to be tested to obtain a mixture of digested fragments to be tested; Introduce the mixture of enzyme-digested fragments to be tested into a microfluidic chip with a microchannel of a set length, and separate them based on a preset electrophoretic mobility to obtain preliminary DNA sample separation fragments; Perform PCR amplification on the preliminary DNA sample separation fragments to generate a methylation-enriched DNA solution; Perform high-throughput sequencing based on the methylation-enriched DNA solution to obtain methylation tumor sequencing data.
[0035] In an embodiment of the present invention, 1 μg of a clinical tumor DNA sample is taken and added to a 20-μL reaction system containing 5 units of HhaI restriction endonuclease (which recognizes and cleaves non-methylated CCGG sites). The reaction system also includes a 1x concentration of the HhaI enzyme-compatible buffer (which provides the optimal pH value, ionic strength, and magnesium ion concentration for enzyme activity) and nuclease-free water supplemented to 20 μL. The mixture is incubated in a 37°C constant temperature water bath for 16 hours to ensure complete digestion. The reaction system is mixed evenly with 1.8 times the volume of magnetic beads and left to stand at room temperature for 5 minutes to allow the DNA fragments to bind fully to the magnetic beads. The centrifuge tube is placed on a magnetic stand, and after the solution becomes clear, the supernatant is carefully aspirated. The magnetic beads are washed twice with 70% ethanol to remove residual salt ions and other impurities. Finally, 30 μL of elution buffer (such as 10 mM Tris-HCl, pH 8.0) is added and incubated at room temperature for 2 minutes to elute the DNA fragments from the magnetic beads. The eluate containing the digested DNA fragments is collected, which is the mixture of digested fragments to be tested. It is precisely injected into the injection hole of the microfluidic chip using a micropipette. The microfluidic chip has a pre-designed microchannel network with a microchannel length of 10 cm, a width of 50 μm, and a depth of 20 μm, and is filled with a polyacrylamide gel electrophoresis medium (concentration of 6%). After injection, a constant voltage of 100 volts is applied across the two ends of the microchannel. Since DNA fragments of different lengths have different migration rates in the gel medium under the action of an electric field, shorter fragments migrate faster and longer fragments migrate slower, thus achieving the separation of DNA fragments of different lengths. The migration of DNA fragments in the microchannel is monitored in real time through an ultraviolet transilluminator. According to a pre-established DNA molecular weight standard curve (established by electrophoresis of DNA standards with known molecular weights), DNA fragments within a specific migration distance range are intercepted. For example, fragments with a migration distance between 4 cm and 6 cm are collected, and this range corresponds to DNA fragments of a specific size (such as 200 bp - 500 bp). They are added to a 25-μL PCR reaction system. The reaction system contains: a 1x concentration of PCR buffer (providing the ionic strength and pH value required for the PCR reaction), a 2.5 mmol / L deoxynucleoside triphosphate mixture (dNTPs, as raw materials for DNA synthesis), a 0.5 μmol / L universal PCR primer pair (which can complementarily bind to the sequences at both ends of the DNA fragment to be amplified, and the primer length is 20 bases), 1.25 units of Taq DNA polymerase (catalyzing the extension of DNA strands), and nuclease-free water supplemented to 25 μL. The PCR reaction tube is placed in a PCR instrument for PCR amplification.The PCR program is set as follows: pre-denaturation at 95 °C for 3 minutes; then 30 cycles are carried out, each cycle including denaturation at 95 °C for 30 seconds (to unwind the DNA double strand), annealing at 58 °C for 30 seconds (for the primer to bind to the DNA template), and extension at 72 °C for 45 seconds (for the DNA polymerase to synthesize a new DNA strand along the template); finally, extension at 72 °C for 5 minutes to ensure that all DNA strands are fully extended. The PCR products are detected by agarose gel electrophoresis to confirm the size and purity of the amplified products. End repair of the DNA is performed by incubating with an end repair enzyme mixture at 37 °C for 30 minutes to phosphorylate the 5' end of the DNA fragment and fill in the 3' end. Then, an A-tailing reaction is carried out by incubating at 72 °C for 30 minutes to add an adenine base (A) to the 3' end of the DNA fragment. Next, a ligation reaction is performed to ligate the adapters with specific sequence tags (index) to both ends of the DNA fragment for subsequent differentiation of different samples in the same sequencing reaction. The ligation reaction is carried out at room temperature for 2 hours. The ligation products are purified using a magnetic bead purification kit. The purified DNA library is accurately quantified by quantitative PCR. Then, the quantified library is denatured at a concentration of 10 picomoles per liter and diluted to the final concentration required for sequencing. The diluted library is loaded into the flow cell of a sequencer (such as the Illumina NovaSeq 6000). In the sequencer, the single-stranded DNA fragments in the DNA library bind to the oligonucleotide probes on the surface of the flow cell and are amplified by bridge PCR to form DNA clusters. Then, using the Sequencing-by-Synthesis (SBS) technology, four fluorescently labeled deoxynucleotides (dNTPs) are added in sequence. Each time a dNTP is added, the fluorescent group is excited by a laser, and the fluorescent signal is captured by a high-resolution CCD camera. According to the different fluorescent signals corresponding to different bases, the base sequence of the DNA fragment is determined. After sequencing is completed, the raw sequencing data (FASTQ format) is generated.
[0036] Preferably, the PCR amplification of the separated fragments of the preliminary DNA sample includes: Mix the separated fragments of the preliminary DNA sample with a bisulfite solution to obtain a bisulfite reaction mixture; Perform a conversion reaction on the bisulfite reaction mixture and carry out a desulfonation treatment to obtain a desulfonated DNA solution; Perform magnetic microsphere adsorption purification on the desulfonated DNA solution to obtain a purified DNA solution; Perform DNA target capture on the purified DNA solution using a preset methylated biotin-labeled probe to obtain a sample-captured DNA solution; Perform PCR amplification on the sample-captured DNA solution to generate a methylation-enriched DNA solution.
[0037] In an embodiment of the present invention, the preparation method of the bisulfite conversion reagent is as follows: weigh 4.75 grams of sodium bisulfite (Sigma-Aldrich, item number S9000), add it to 8 milliliters of nuclease-free water, and after fully dissolving, add 1.25 milliliters of 2M sodium hydroxide solution (prepared by solid sodium hydroxide and nuclease-free water) to adjust the pH value to 5.0. Then add 0.5 milliliters of newly prepared 10mM hydroquinone solution (Sigma-Aldrich, item number H9003, dissolved in nuclease-free water), and finally make up to 10 milliliters with nuclease-free water, filter and sterilize with a 0.22μm filter membrane, and store at 4°C away from light. The bisulfite reaction mixture is placed in a PCR instrument and the following conversion reaction is performed: denaturation at 95°C for 5 minutes; incubation at 60°C for 25 minutes; denaturation at 95°C for 5 minutes; incubation at 60°C for 25 minutes; denaturation at 95°C for 5 minutes; incubation at 60°C for 1 hour and 45 minutes; storage at 4°C. This program is cycled twice. After the conversion reaction is completed, desulfonation treatment is performed. Use the DNA binding buffer (Buffer BL) in the EpiTect Bisulfite Kit (QIAGEN, Cat. No. 59104) to add 600 μl of Buffer BL to every 20 μl of bisulfite reaction solution, and mix the mixture after the conversion reaction with Buffer BL. Subsequently, transfer the mixture to the EpiTect adsorption column provided by the kit and centrifuge at 12000g for 1 minute at room temperature. Discard the flow-through, add 500 μl of wash buffer (Buffer BW) to the adsorption column, and centrifuge at 12000g for 1 minute. Repeat the washing step once. Add 500 μl of desulfonation buffer (Buffer BD) to the adsorption column and let it stand at room temperature for 15 minutes. The function of the desulfonation buffer is to remove the sulfonic acid groups formed during the bisulfite conversion process, so that uracil is completely converted into uracil. Place the adsorption column in a clean 1.5 ml centrifuge tube, add 20 μl of elution buffer (Buffer EB), let stand at room temperature for 1 minute, centrifuge at 12000g for 1 minute, and collect the eluate, which is the desulfonated DNA solution. Incubate the mixture at room temperature for 5 minutes to allow the DNA to fully bind to the magnetic beads. After incubation, place the centrifuge tube on the magnetic stand for 2 minutes. After the magnetic beads are completely adsorbed to one side of the tube wall and the solution is clear, carefully aspirate and discard the supernatant, taking care not to absorb the magnetic beads. Keep the centrifuge tube on the magnetic stand, add 200 μl of 80% ethanol solution (freshly prepared using anhydrous ethanol and nuclease-free water), gently rotate the centrifuge tube several times to wash the magnetic beads, let stand for 30 seconds, and carefully aspirate and discard the supernatant. Repeat the ethanol washing step once. Remove the centrifuge tube from the magnetic stand, briefly centrifuge at low speed for a few seconds to allow the residual ethanol to gather at the bottom of the tube, then put it back on the magnetic stand and aspirate and discard the residual ethanol with a micropipette. Dry the magnetic beads at room temperature for 5-10 minutes until the ethanol is completely evaporated.Add 20 μL of elution buffer (10 mM Tris-HCl, pH 8.0) to the magnetic beads, vortex thoroughly to resuspend the magnetic beads, incubate at room temperature for 2 minutes. After the magnetic beads are completely adsorbed, carefully aspirate the supernatant into a new 1.5 mL centrifuge tube, which is the purified DNA solution. Place the mixture in a PCR instrument and perform the following hybridization reaction: denature at 95 °C for 10 minutes; then slowly cool down to 58 °C (cooling at 0.1 °C per second) and incubate at 58 °C for 4 hours. During the incubation, the biotin-labeled probe binds complementarily to the target DNA fragment. After hybridization is completed, prepare streptavidin magnetic beads (such as Dynabeads MyOne Streptavidin C1, Thermo Fisher Scientific). Take 50 μL of the magnetic bead suspension, place it on a magnetic stand, and discard the supernatant. Wash the magnetic beads twice with 200 μL of magnetic bead wash buffer (formula refer to the IDT manual). Add the hybridized mixture to the washed streptavidin magnetic beads, incubate at room temperature for 30 minutes, and gently oscillate to allow the biotin-labeled probe-DNA complex to bind to the streptavidin magnetic beads. After incubation, place the centrifuge tube on a magnetic stand and discard the supernatant. Wash the magnetic beads three times with 200 μL of wash buffer I (formula refer to the IDT manual), then wash the magnetic beads three times with 200 μL of wash buffer II (formula refer to the IDT manual), and finally wash the magnetic beads once with 200 μL of wash buffer III (formula refer to the IDT manual). For each wash, let it stand on the magnetic stand for 2 minutes and carefully aspirate and discard the supernatant. After washing is completed, it is the sample-captured DNA solution. The primers are the universal adapter sequences of the Illumina sequencing platform. Place the PCR reaction tube in a PCR instrument and run the following PCR program: pre-denature at 95 °C for 3 minutes; then perform 10 - 15 cycles (denature at 98 °C for 20 seconds, anneal at 60 °C for 30 seconds, extend at 72 °C for 45 seconds); finally extend at 72 °C for 5 minutes. The number of cycles is adjusted according to the capture efficiency and the initial DNA amount. After PCR is completed, purify it using AMPure XP magnetic beads. Dissolve the purified PCR product in 20 μL of elution buffer, which is the methylation-enriched DNA solution.
[0038] Preferably, the sequencing differential site recognition module includes the following functions: Perform base recognition on the methylated tumor sequencing data to obtain sequencing base sequence data; Remove the adapter sequences according to the sequencing base sequence data and perform sequencing quality assessment to obtain sequencing quality assessment data; Filter the low-quality sequencing of the methylated tumor sequencing data through the sequencing quality assessment data to generate high-quality sequencing data; Quantify the methylation level according to the high-quality sequencing data to generate methylation site frequency data; Perform tumor differential site analysis based on methylation site frequency data to generate tumor differential site data.
[0039] In an embodiment of the present invention, the image processing software provided by the Illumina NovaSeq 6000 sequencing platform is used to convert the fluorescent signal image generated during the sequencing process into a base sequence. Specifically, the software identifies the fluorescent signal of each cluster according to the fluorescent image of each cycle and converts it into the corresponding base (A, T, C, G). Each sequencing cycle reads one base, and through multiple cycles of reading, the complete base sequence of each DNA fragment is finally obtained. The Cutadapt tool (version, for example, v3.0) is used to remove the adapter sequence in the sequencing base sequence data. The adapter sequence is specified as the adapter sequence used in PCR amplification, and the parameters are set to allow up to 3 base mismatches. After the adapter is removed, the FastQC tool (version, for example, v0.11.9) is used to evaluate the quality of the sequence after the adapter is removed. FastQC generates a report containing information such as the length distribution of the sequencing reads, the base quality distribution, the GC content distribution, the residual adapter, and the content of the repeated sequence. According to the quality statistics in the FastQC report, such as the average quality value of each base position (Phred quality value), the average quality value of the read segment, the residual rate of the adapter, etc., the quality of the sequencing data is evaluated to obtain the sequencing quality evaluation data. Run the Trimmomatic command line tool, specify the input FASTQ file, output file, and quality filtering parameters. Set the sliding window filter (SLIDINGWINDOW), for example, the window size is 4 bases, and the quality threshold is 20, which means that if the average quality value in a 4-base window is lower than 20, the bases in the window and after it are cut off. Set the head trimming (LEADING), for example, the trimming value is 3, which means that the 3 bases at the head of the sequencing fragment are cut off, because the quality of the head base is usually low. Set the tail trimming (TRAILING), for example, the trimming value is 3, which means that the 3 bases at the tail of the sequencing fragment are cut off. Set the minimum length filter (MINLEN), for example, the minimum length is 50 bases, which means that if the length of the filtered sequencing fragment is less than 50 bases, the fragment is discarded. Since bisulfite treatment converts unmethylated cytosine to uracil, it is necessary to convert all C in the reference genome to T and convert all C in the sequencing data to T for alignment. After the alignment is completed, Bismark can identify the methylation status of each CpG site and calculate the methylation level of each CpG site. The methylation level is defined as the ratio of the number of methylated reads to the total number of reads, expressed as a percentage. The methylation level information of each CpG site is summarized to generate methylation site frequency data. The methylation site frequency data of the tumor samples are compared with the methylation site frequency data of the normal control samples to identify differentially methylated sites. The differential methylation analysis uses a logistic regression model and uses the false discovery rate (FDR) to control multiple testing corrections. The FDR threshold is set to 0.05, that is, differentially methylated sites with an FDR less than 0.05 are screened out.Summarize the position information of differentially methylated sites, methylation level differences, p-values, FDR values, etc. to generate tumor differential site data.
[0040] Preferably, the quantification of methylation level according to high-quality sequencing data includes: Perform transformed sequence-specific genomic alignment on high-quality sequencing data to obtain methylation sequencing alignment data; Remove repetitive sequences based on the methylation sequencing alignment data to obtain unique sequencing alignment sequence data; Extract cytosine context sequences based on the unique sequencing alignment sequence data to obtain cytosine site context sequence data; Perform bisulfite conversion judgment on both strands of the cytosine site context sequence data to generate cytosine site conversion status data; Based on the cytosine site conversion status data, count the number of methylated and non-methylated sequencing fragments in the unique sequencing alignment sequence data to obtain the number of methylated sequencing fragments and the number of non-methylated sequencing fragments respectively; Calculate the methylation frequency according to the number of methylated sequencing fragments and the number of non-methylated sequencing fragments to generate methylation site frequency data.
[0041] In the embodiments of the present invention, all cytosines C are converted to thymines T). At the same time, Bismark also constructs a reverse complementary reference genome for "G-to-A" conversion (converting all guanines G to adenines A) to handle the case where the sequencing fragments come from the negative strand (because the DNA strand after bisulfite treatment has directionality). Then, each high-quality sequencing fragment is aligned with the original reference genome, the "C-to-T" conversion reference genome, and the "G-to-A" conversion reference genome respectively. Specific operation: Run the Bismark command-line tool, specifying the reference genome directory, FASTQ file path, output directory, alignment tool (Bowtie2 or HISAT2) and its parameters (such as the number of allowed mismatches, multiple alignment handling methods, etc.). Bismark will generate four sets of alignment results: the alignment result of the original reference genome, the positive strand alignment result of the C-to-T conversion genome, the negative strand alignment result of the C-to-T conversion genome, and the positive strand alignment result of the G-to-A conversion genome. According to the quality of the alignment results (such as alignment score, uniqueness of alignment, etc.), the best alignment result is selected. The alignment results are stored in SAM / BAM format files, which are the methylation sequencing alignment data. Use the markdup command of SAMtools or the MarkDuplicates command of Picard, input the BAM file, and output the BAM file after marking duplicate sequences. These tools determine whether they are PCR duplicates based on the alignment position and direction of the sequencing fragments. If multiple sequencing fragments have the same alignment start position and direction, they are considered PCR duplicates, and only one of them is retained (usually the one with the highest sequencing quality). The sequencing fragments with marked duplicates will be marked in the FLAG field of the BAM file (such as adding the DUP mark). Then, use the view command of SAMtools to filter out the sequencing fragments with the DUP mark, and only retain the unmarked sequencing fragments. Traverse each sequencing fragment in the BAM file, and determine the coordinates of each cytosine site on the reference genome according to the alignment position and the CIGAR string (recording alignment operations, such as insertions, deletions, matches, etc.). For each cytosine site, extract a certain number of bases before and after it (such as 2 bases before and after each, forming a 5-base context sequence). At the same time, record whether this cytosine site is in a CpG, CHG, or CHH sequence context. For the context sequence data of each cytosine site, judge its conversion status according to the bisulfite conversion rule. If the cytosine is at a CpG site and this site is converted to thymine (T), then judge that this site is unmethylated; if the cytosine is at a CpG site and remains cytosine (C), then judge that this site is methylated. If the cytosine is not at a CpG site, no methylation status judgment is made. Based on the cytosine site conversion status data and the unique sequencing alignment sequence data, count the number of methylated and unmethylated sequencing fragments for each CpG site.Traverse the unique sequencing alignment sequence data. For each sequencing read covering a CpG site, determine whether the read is a methylated read or an unmethylated read according to the cytosine site conversion status data. Count the number of methylated reads and the number of unmethylated reads for each CpG site to obtain the number of methylated sequencing fragments and the number of unmethylated sequencing fragments respectively. Calculate the methylation frequency according to the number of methylated sequencing fragments and the number of unmethylated sequencing fragments for each CpG site. The methylation frequency calculation formula is: Methylation frequency = Number of methylated sequencing fragments / (Number of methylated sequencing fragments + Number of unmethylated sequencing fragments). Store the methylation frequency information of each CpG site in a text file to generate methylation site frequency data.
[0042] Preferably, the calculation of the methylation frequency according to the number of methylated sequencing fragments and the number of unmethylated sequencing fragments includes: Statistically count the number of site sequencing fragment coverages according to the number of methylated sequencing fragments and the number of unmethylated sequencing fragments to generate cytosine site total sequencing depth data; Calculate the site methylation ratio according to the number of methylated sequencing fragments and the number of unmethylated sequencing fragments to generate an initial methylation frequency; Establish a mapping relationship between the sequencing depth and the methylation frequency, and construct a frequency background noise model using a preset negative binomial distribution model; Set a low-depth threshold parameter based on the frequency background noise model; filter the initial methylation frequency for low-site frequencies through the cytosine site total sequencing depth data based on the low-depth threshold parameter to obtain filtered methylation frequency data; Calculate the frequency confidence interval for the filtered methylation frequency data using the frequency background noise model to generate methylation confidence interval data; Perform confidence interval correction on the filtered methylation frequency data through the methylation confidence interval data to generate methylation site frequency data.
[0043] In the embodiments of the present invention, all CpG sites are traversed. The number of methylated sequencing fragments and the number of unmethylated sequencing fragments at each site are added together to obtain the total number of sequencing fragment coverage times at that site. The coverage times information of each CpG site is stored in a data structure, such as a dictionary or a list, to generate the total sequencing depth data of cytosine sites. For example, if the number of methylated sequencing fragments at a certain CpG site is 30 and the number of unmethylated sequencing fragments is 20, then the total sequencing depth at this site is 50. For each CpG site, the number of methylated sequencing fragments is divided by the total number of sequencing fragment coverage times to obtain the initial methylation frequency at that site. The initial methylation frequency information of each CpG site is stored in a data structure to generate the initial methylation frequency data. For example, if the number of methylated sequencing fragments at a certain CpG site is 30 and the total sequencing depth is 50, then the initial methylation frequency at this site is 0.6. A set of CpG site data with known methylation status is collected. For example, methylation data in a public database or sequencing data of artificially synthesized methylated DNA fragments is used. Based on this data, a scatter plot of sequencing depth and methylation frequency is drawn. By observing the scatter plot, it can be found that in the low sequencing depth region, the methylation frequency fluctuates greatly and there is obvious background noise. A negative binomial distribution model is used to fit the relationship between sequencing depth and methylation frequency to construct a frequency background noise model. The negative binomial distribution model can describe the influence of sequencing depth on methylation frequency and is used to estimate the background noise level. According to the frequency background noise model, a low-depth threshold parameter is set. For example, the threshold is set to 10, that is, CpG sites with a sequencing depth lower than 10 are considered low-depth sites. According to the total sequencing depth data of cytosine sites, the sites with a sequencing depth lower than the threshold in the initial methylation frequency data are filtered out to obtain the filtered methylation frequency data. Using the frequency background noise model (negative binomial distribution model), the confidence interval of the methylation frequency is calculated for each CpG site in the filtered methylation frequency data. For example, a 95% confidence interval is calculated. The calculation method of the confidence interval can be based on the probability density function of the negative binomial distribution. For example, the Wald method or the Wilson score interval method is used. The methylation frequency confidence interval information of each CpG site is stored in a data structure to generate the methylation confidence interval data. According to the methylation confidence interval data, confidence interval correction is performed on the filtered methylation frequency data. For example, if the methylation frequency of a certain CpG site is 0.6 and its 95% confidence interval is [0.4, 0.8], then the methylation frequency value of this site is retained. If the confidence interval of a certain CpG site contains 0.5, it is considered that the methylation status of this site is not clear, and its methylation frequency is set to 0.5, or it is removed from the data. The corrected methylation frequency information is stored in a data structure to generate the final methylation site frequency data.
[0044] Preferably, the tumor differential site analysis based on the methylation site frequency data includes: Dividing the high-quality sequencing data into functional regions by using preset genomic annotation information to obtain genomic functional region list data; Dividing the genomic functional region list data into tumor / normal samples through the methylation site frequency data, and performing site frequency annotation to generate regional methylation frequency data; Calculating the average methylation levels of tumor samples / normal samples according to the regional methylation frequency data to obtain the mean methylation value of the tumor region and the mean methylation value of the normal region respectively; Calculating the methylation difference value of the tumor region methylation mean through the normal region methylation mean, and performing a significance test of the difference to obtain the regional methylation difference P value; Screening the tumor differential methylation regions from the genomic functional region list data through the regional methylation difference P value, and performing tumor differential site annotation to generate tumor differential site data.
[0045] In the embodiments of the present invention, a genomic annotation file (e.g., the annotation file of the GENCODE database) is used to divide the reference genome into different functional regions, such as promoter regions, gene body regions, enhancer regions, intron regions, intergenic regions, etc. The genomic annotation file usually contains information such as the start position, end position, exon region, and intron region of genes. According to this information, the functional region to which each CpG site belongs can be determined. The genomic functional region information is stored in a data structure, such as a list or a dictionary, to generate genomic functional region list data, which contains the name, start position, and end position of each region. According to the sample information, the samples in the methylation site frequency data are divided into tumor samples and normal samples. Then, according to the genomic functional region list data, the methylation frequency information of each CpG site is added to the corresponding functional region. Specifically, for each CpG site, the functional region to which it belongs is found, and the methylation frequency of this site is added to the tumor sample or normal sample data of this functional region. Finally, regional methylation frequency data is generated, which contains the methylation frequency information of all CpG sites in tumor samples and normal samples for each functional region. For each genomic functional region, the average methylation level of all CpG sites in tumor samples and normal samples is calculated respectively. Specifically, the methylation frequencies of all CpG sites in this region in tumor samples or normal samples are added up, and then divided by the total number of CpG sites in this region to obtain the average methylation level of this region in tumor samples or normal samples. The tumor region methylation mean and the normal region methylation mean are obtained respectively. For each genomic functional region, the tumor region methylation mean is subtracted from the normal region methylation mean to obtain the methylation difference of this region. Then, a statistical method (e.g., t-test or Wilcoxon rank sum test) is used to perform a significance test on the methylation difference, and the p-value is calculated. The p-value indicates whether there is a significant difference in the methylation level between tumor samples and normal samples in this region. The methylation difference and p-value of each region are stored in a data structure to obtain regional methylation difference P-value. According to a preset significance level threshold (e.g., 0.05), the genomic functional regions with p-values less than the threshold in the regional methylation difference P-value are screened out. These regions are considered to be tumor differentially methylated regions. Then, all CpG sites in these regions are marked as tumor differential sites. The position information of the tumor differential sites, the functional regions to which they belong, the average methylation levels of tumor samples and normal samples, the methylation difference, and the p-value, etc. are stored in a data structure to generate tumor differential site data.
[0046] Preferably, the tumor gene sequence analysis module includes the following functions: Perform genomic coordinate mapping according to the tumor differential site data, and extract the genomic sequence of the corresponding region to generate tumor differential region sequence data; Construct a differential methylation site expression network based on the tumor differential region sequence data to generate a gene co-expression network; Perform network topology analysis on the gene co-expression network to obtain node centrality indicators; Identify the key gene sequences for tumorigenesis in the gene co-expression network based on the node centrality indicators to obtain tumor recognition gene sequences; Perform sequence feature analysis based on the tumor recognition gene sequences to generate tumor gene sequence feature vectors; Extract the methylation level based on the tumor recognition gene sequences and perform feature fusion on the tumor gene sequence feature vectors to generate tumor methylation feature data.
[0047] As an example of the present invention, refer to Figure 2 shown in Figure 1 the detailed implementation process schematic diagram of S3 in S31: Perform genome coordinate mapping based on the tumor differential site data and extract the genomic sequences of the corresponding regions to generate tumor differential region sequence data; In the embodiments of the present invention, the tumor differential site data contains the genomic coordinate information of the differential methylation sites. Using these coordinate information, combined with the human reference genome sequence (for example, GRCh38 version), the genomic sequences of the regions where each differential methylation site is located are extracted. The length of the extracted sequences can be set according to specific requirements. For example, extract the sequences with 1000 bases upstream and downstream centered on the differential methylation sites. The samtools faidx command or the getfasta command in the Bedtools toolkit can be used to extract the corresponding sequences from the reference genome according to the genomic coordinates.
[0048] S32: Construct a differential methylation site expression network based on the tumor differential region sequence data to generate a gene co-expression network; In the embodiments of the present invention, based on the tumor differential region sequence data, a differential methylation site expression network is constructed. First, according to the genomic annotation information, determine the genes to which each differential methylation site belongs. Then, calculate the correlation of the expression levels between genes. For example, use the Pearson correlation coefficient or the Spearman correlation coefficient. RNA-seq data can be used to calculate the gene expression levels. Connect the gene pairs with a correlation coefficient higher than a preset threshold (for example, 0.8) to construct a gene co-expression network. The nodes in the network represent genes, and the edges represent the co-expression relationships between genes. Tools such as igraph and NetworkX can be used to construct and visualize the gene co-expression network.
[0049] S33: Perform network topology analysis on the gene co-expression network to obtain node centrality metrics; In the embodiments of the present invention, topology analysis is performed on the constructed gene co-expression network, and the centrality metrics of each node in the network are calculated. For example, degree centrality, betweenness centrality, closeness centrality, eigenvector centrality, etc. Degree centrality represents the number of edges connected to a node; betweenness centrality represents the number of shortest paths passing through the node; closeness centrality represents the average distance from the node to other nodes; eigenvector centrality represents the influence of the node in the network. These centrality metrics can be calculated using toolkits such as igraph and NetworkX.
[0050] S34: Identify key gene sequences for tumorigenesis in the gene co-expression network based on the node centrality metrics to obtain tumor recognition gene sequences; In the embodiments of the present invention, for each centrality metric, a threshold is set (for example, genes ranked in the top 10% of the centrality metric are selected). Genes that are above the threshold in multiple centrality metrics are selected as key genes. For example, genes ranked in the top 10% of degree centrality, betweenness centrality, and eigenvector centrality can be selected. These genes play important roles in the network, participate in regulating the expression of multiple genes, or are at key positions in signal pathways, and are closely related to the occurrence and development of tumors. The sequence information of these key genes (such as the coding sequence CDS, promoter sequence, UTR sequence, etc. of the gene) is obtained from genomic annotation information or gene sequence databases (such as NCBI RefSeq or Ensembl).
[0051] S35: Perform sequence feature analysis based on the tumor recognition gene sequences to generate tumor gene sequence feature vectors; In the embodiments of the present invention, various features of tumor recognition gene sequences are analyzed, including but not limited to: k-mer frequency (the frequency of subsequences of length k in the sequence, for example, when k = 3, the frequencies of AAA, AAC, AAG, etc.), GC content (the proportion of guanine G and cytosine C in the sequence), CpG island density (the number and length of CpG islands in the sequence), the number and type of repetitive sequences (such as short tandem repeats STR, interspersed repeats SINE / LINE, etc.), the length and position of conserved regions (conserved sequences identified by aligning with genomic sequences of other species), prediction of transcription factor binding sites (TFBS) (prediction using known TFBS patterns or motifs), etc. These sequence features are represented as numerical vectors. For example, the k-mer frequency can be represented by a vector of length 4^k, and each element represents the frequency of a k-mer; the GC content can be represented by a numerical value; the CpG island density can be represented by the number of CpG islands per kb of the sequence. The numerical values of all features are combined into a vector as the feature vector of the gene sequence.
[0052] S36: Extract the methylation level according to the tumor recognition gene sequence, and perform feature fusion on the tumor gene sequence feature vector to generate tumor methylation feature data.
[0053] In the embodiments of the present invention, according to the genomic coordinates of the tumor recognition gene sequence, the methylation level information of these gene regions is extracted from the previous methylation site frequency data or regional methylation frequency data. For example, the average methylation level of gene promoter regions, gene body regions, exon regions, intron regions, etc. can be calculated. The methylation level information is spliced or combined with the tumor gene sequence feature vector. For example, the methylation level can be added as an additional feature to the sequence feature vector to form a longer vector. It is also possible to perform dimensionality reduction or feature selection on the methylation level vector and the sequence feature vector respectively, and then combine the features after dimensionality reduction. An output vector file containing the fused features is generated, and this file is the tumor methylation feature data. Each row in this file represents a gene, and each column represents a sequence feature or methylation feature.
[0054] Preferably, the sequence feature analysis according to the tumor recognition gene sequence includes: Calculate the GC content of the tumor recognition gene sequence to obtain the GC content value; Identify CpG islands in the tumor recognition gene sequence and calculate the CpG island density to generate the CpG island density value; Splice the GC content value and the CpG island density value to generate an initial sequence feature vector; Perform regularization processing on the initial sequence feature vector to generate the sequence feature vector; Perform K-Mer decomposition on the sequence data of tumor differential regions, and perform feature integration according to the sequence feature vectors to generate tumor gene sequence feature vectors.
[0055] In the embodiments of the present invention, for each tumor recognition gene sequence, each base in the sequence is traversed to count the number of guanine (G) and cytosine (C). Then, the sum of the numbers of G and C is divided by the total length of the sequence to obtain the GC content value of the sequence, expressed as a percentage. For example, in a sequence with a length of 1000 bp, the numbers of G and C are 400 and 300 respectively, then the GC content value of the sequence is (400 + 300) / 1000 = 0.7, that is, 70%. Use a CpG island recognition tool, such as CpG IslandSearcher, to perform CpG island recognition on each tumor recognition gene sequence. CpG island recognition tools usually perform recognition based on parameters such as the ratio of the observed frequency to the expected frequency of CpG sites, GC content, and length. After identifying the CpG islands, calculate the total length of the CpG islands in each gene sequence, and then divide it by the total length of the gene sequence to obtain the CpG island density value. For example, in a sequence with a length of 1000 bp, the total length of the CpG islands is 200 bp, then the CpG island density value of the sequence is 200 / 1000 = 0.2. Concatenate the GC content value and the CpG island density value of each tumor recognition gene sequence to construct an initial sequence feature vector. For example, if the GC content value of a sequence is 0.7 and the CpG island density value is 0.2, then the initial sequence feature vector of this sequence is [0.7, 0.2]. Perform regularization processing on the initial sequence feature vector. For example, use the min-max normalization or Z-score normalization method to scale the feature values to a specific range, such as [0, 1] or [-1, 1]. This can avoid the influence of the difference in the order of magnitude between different feature values on subsequent analysis. For example, using the min-max normalization method, the feature value x is converted to x'=(x - min) / (max - min), where min and max are the minimum and maximum values of this feature in all sequences respectively. Perform K-Mer decomposition on the sequence data of tumor differential regions. K-Mer refers to a DNA subsequence with a length of K. For example, set the K value to 3, then each sequence is decomposed into all subsequences with a length of 3. Count the frequency of each K-Mer in each sequence, and add these frequency values as new features to the sequence feature vector. For example, the sequence "ATGC" will be decomposed into two 3-Mers, "ATG" and "TGC". Count the frequency of each 3-Mer in all sequences, and add these frequency values to the regularized sequence feature vector to generate the final tumor gene sequence feature vector.
[0056] Preferably, the tumor subtype classification module includes the following functions: Obtain historical tumor gene sequence feature data; Perform stratified sampling on the historical tumor gene sequence feature data to obtain a training set and a test set respectively; Perform transfer learning on a preset convolutional neural network model through the training set to generate a trained tumor risk prediction model; Perform model cross-validation on the trained tumor risk prediction model through the test set and optimize the parameters to obtain a tumor risk prediction model; Transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction and perform classification probability conversion to generate a tumor classification probability matrix; Identify the maximum probability value according to the tumor classification probability matrix and perform tumor subtype label identification to generate tumor type label data.
[0057] As an example of the present invention, refer to Figure 3 shown, for Figure 1 a detailed implementation process schematic diagram of S4 in S41: Obtain historical tumor gene sequence feature data; In the embodiment of the present invention, determine the tumor type and data type (such as gene expression data, methylation data, copy number variation data, etc.) for which data needs to be obtained. Download the Level 3 data (preprocessed and standardized data) of the corresponding tumor type from the TCGA database (portal.gdc.cancer.gov). For example, download the gene expression RNAseq data (HTSeq - FPKM-UQ format) and methylation data (Illumina HumanMethylation450BeadChip data) of breast cancer (BRCA). Download similar data of other tumor types from the ICGC database (dcc.icgc.org). Search and download relevant gene expression or methylation data sets from the GEO database (www.ncbi.nlm.nih.gov / geo). If collaborating with a hospital, genomic data and clinical information (including tumor subtypes) of tumor samples can be obtained from the hospital's electronic medical record system or biobank. Ensure that all data has undergone quality control and preprocessing and has a consistent format and identifier. Integrate all data into a data set that contains the feature information (such as gene expression values, methylation levels) of genes / regions and the subtype labels of samples.
[0058] S42: Perform stratified sampling on the historical tumor gene sequence feature data to obtain a training set and a test set respectively; In the embodiments of the present invention, according to the subtype labels of the samples, the number of samples of each subtype in the dataset is counted. Calculate the number of samples that should be allocated to the training set and the test set for each subtype. For example, it is set that the training set accounts for 70% and the test set accounts for 30%. If a certain subtype has 100 samples, then 70 samples should be allocated to the training set and 30 samples should be allocated to the test set. For each subtype, randomly select a specified number of samples (without replacement sampling) from the samples of this subtype to add to the training set, and the remaining samples are added to the test set. The train_test_split function in a programming language (such as the scikit-learn library in Python) can be used, and the stratify parameter is set to the subtype label to achieve stratified sampling. Output two files, which respectively contain the sample IDs, gene / region feature data, and subtype labels of the training set and the test set.
[0059] S43: Perform transfer learning on a preset convolutional neural network model through the training set to generate a training tumor risk prediction model; In the embodiments of the present invention, a pre-trained CNN model is selected. For example, the ResNet50 or VGG16 model pre-trained on the ImageNet dataset can be selected. These models have learned the general features of image recognition and can be transferred to genomic data analysis. A deep learning framework (such as TensorFlow or PyTorch) can be used to load the pre-trained model. Modify the last layer (fully connected layer) of the pre-trained model so that the number of output classes is the same as the number of tumor subtypes. For example, if the tumor has 4 subtypes, the last layer should output a 4-dimensional vector, and each dimension represents the probability of a subtype. Freeze most of the layers of the pre-trained model (such as convolutional layers), and only train the last few layers (such as fully connected layers) or fine-tune some convolutional layers. In this way, the feature extraction ability of the pre-trained model can be utilized while avoiding overfitting. Use the gene / region feature data of the training set as the input and the subtype labels as the output to train the modified CNN model. Use a suitable optimization algorithm (such as Adam) and loss function (such as cross-entropy loss function). Set the training parameters (such as learning rate, batch size, number of training epochs, etc.). During the training process, monitor the performance of the model on the validation set (a part of the data is divided from the training set as the validation set), and select the model with the best performance as the training tumor risk prediction model.
[0060] S44: Perform model cross-validation on the training tumor risk prediction model through the test set and perform parameter optimization to obtain the tumor risk prediction model; In the embodiments of the present invention, the gene / region feature data of the test set is used as input and input into the trained tumor risk prediction model. The model will output the prediction probabilities of each sample belonging to different subtypes. The subtype with the highest prediction probability is used as the predicted subtype of the sample. Calculate indicators such as accuracy, precision, recall rate, F1 value, and area under the ROC curve (AUC) of the model on the test set to evaluate the performance of the model. The cross-validation method can be used to divide the training set into multiple subsets (such as 5-fold or 10-fold). Each time, one subset is used as the validation set, and the remaining subsets are used as the training set. The model is trained and validated repeatedly, and the average performance is taken as the final performance of the model. Methods such as grid search, random search, or Bayesian optimization are used to adjust the hyperparameters of the model (such as learning rate, regularization coefficient, number of layers of CNN, and size of convolutional kernels, etc.). For each set of hyperparameters, the model is trained and validated repeatedly, and the combination of hyperparameters with the best performance is selected. The model is retrained on the entire training set using the best combination of hyperparameters to obtain the final tumor risk prediction model.
[0061] S45: Transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction, and perform classification probability conversion to generate a tumor classification probability matrix; In the embodiments of the present invention, the tumor methylation feature data of the sample to be predicted (for example, a vector containing gene / region methylation levels and sequence features) is converted into a format consistent with the training data. The feature data is input into the tumor risk prediction model. The model will output a vector, and each element of the vector represents the probability of the sample belonging to each tumor subtype. For example, if the model predicts that the probabilities of a sample belonging to subtype 1, subtype 2, subtype 3, and subtype 4 are 0.1, 0.6, 0.2, and 0.1 respectively, the output vector is [0.1, 0.6, 0.2, 0.1]. For multiple samples to be predicted, the above process is repeated to obtain a matrix. Each row of the matrix represents a sample, each column represents a subtype, and the elements in the matrix are the probabilities of the sample belonging to the subtype.
[0062] S46: Identify the maximum probability value according to the tumor classification probability matrix, and perform tumor subtype label identification to generate tumor type label data.
[0063] In the embodiments of the present invention, for each row (representing a sample) in the tumor classification probability matrix, find the column (subtype) corresponding to the maximum probability value in that row. Use the index (or subtype name) of this column as the predicted subtype label for this sample. For example, if the probability vector of a sample is [0.1, 0.6, 0.2, 0.1], then the maximum probability value is 0.6, and the corresponding column is the second column. Therefore, the predicted subtype label for this sample is subtype 2. Repeat the above process for all samples to obtain a vector or list containing the predicted subtype labels for all samples. This vector or list is the tumor type label data.
[0064] This application lies in that, by combining the efficient separation ability of the microfluidic chip, it is possible to perform deep sequencing and precise separation on low-concentration clinical tumor DNA samples. The application of the microfluidic chip not only improves the efficiency and accuracy of sample processing, but also reduces the risk of sample loss and cross-contamination. This combination of technologies enables even weak methylation signals to be effectively detected, significantly reducing the false negative rate in the detection of low-concentration samples. A frequency background noise model is constructed using the negative binomial distribution model to correct the confidence interval of the methylation frequency, thereby effectively removing background noise and improving the accuracy and reliability of the methylation frequency data. This precise differential site analysis avoids the errors caused by the judgment of the methylation state of a single site in traditional methods. By combining network topology analysis and sequence feature analysis, tumor methylation feature data is generated, which not only captures the methylation level of the gene sequence, but also integrates the gene expression regulation information. By constructing a tumor risk prediction model, precise prediction of tumor types and subtype label identification are achieved. The various modules of the system are seamlessly connected, ensuring the efficiency and continuity of sample processing and data analysis. Through an automated data processing flow, the errors caused by manual operations are reduced, and the efficiency and accuracy of data processing are improved, with significant advantages such as high sensitivity, high specificity, low false negative rate, and high efficiency.
[0065] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to encompass all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.
[0066] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A methylation-based tumor data processing system, characterized in that It includes the following modules: The methylation sample sequencing module includes a high-throughput sequencer, a microfluidic chip, a restriction enzyme digestion kit, and a PCR instrument. The restriction enzyme digestion kit is used to perform restriction enzyme digestion on the clinical tumor DNA sample to be tested, and the microfluidic chip is used for separation to obtain preliminary DNA sample separation fragments; the high-throughput sequencer is used to perform high-throughput sequencing on the preliminary DNA sample separation fragments to obtain methylation tumor sequencing data; The sequencing differential site recognition module is used to filter out low-quality sequencing from the methylation tumor sequencing data to generate high-quality sequencing data; Quantify the methylation level based on the high-quality sequencing data to generate methylation site frequency data; analyze the tumor differential sites based on the methylation site frequency data to generate tumor differential site data; The tumor gene sequence analysis module includes a genomic database, which is used to extract the genomic sequence of the corresponding region based on the tumor differential site data to generate tumor differential region sequence data; identify the key gene sequences for tumorigenesis based on the tumor differential region sequence data, and perform sequence feature analysis to generate tumor methylation feature data; The tumor subtype classification module is used to construct a tumor risk prediction model; transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction, and perform tumor subtype label identification to generate tumor type label data.
2. The methylation-based tumor data processing system according to claim 1, wherein The methylation sample sequencing module includes the following functions: Perform restriction enzyme digestion on the clinical tumor DNA sample to be tested to obtain a mixture of restriction enzyme digestion fragments to be tested; Import the mixture of restriction enzyme digestion fragments to be tested into a microfluidic chip with a microchannel of a set length, and separate based on the preset electrophoretic mobility to obtain preliminary DNA sample separation fragments; Perform PCR amplification on the preliminary DNA sample separation fragments to generate a methylation-enriched DNA solution; Perform high-throughput sequencing based on the methylation-enriched DNA solution to obtain methylation tumor sequencing data.
3. The methylation-based tumor data processing system according to claim 2, wherein The PCR amplification of the preliminary DNA sample separation fragments includes: Mix the preliminary DNA sample separation fragments with a bisulfite solution to obtain a bisulfite reaction mixture; Perform a conversion reaction on the bisulfite reaction mixture and perform desulfonation treatment to obtain a desulfonated DNA solution; Perform magnetic bead adsorption purification on the desulfonated DNA solution to obtain a purified DNA solution; Perform DNA target capture on the purified DNA solution through a preset methylation biotin-labeled probe to obtain a sample-captured DNA solution; Perform PCR amplification on the sample-captured DNA solution to generate a methylation-enriched DNA solution.
4. The methylation-based tumor data processing system according to claim 1, wherein The sequencing differential site recognition module includes the following functions: Perform base recognition on the methylation tumor sequencing data to obtain sequencing base sequence data; Remove the adapter sequence based on the sequencing base sequence data and perform sequencing quality assessment to obtain sequencing quality assessment data; Filter out low-quality sequencing from the methylation tumor sequencing data through the sequencing quality assessment data to generate high-quality sequencing data; Quantify the methylation level based on the high-quality sequencing data to generate methylation site frequency data; Analyze the tumor differential sites based on the methylation site frequency data to generate tumor differential site data.
5. The methylation-based tumor data processing system according to claim 4, wherein, The quantification of methylation levels based on high-quality sequencing data includes: Performing a transformed sequence-specific genomic alignment on the high-quality sequencing data to obtain methylation sequencing alignment data; Removing repetitive sequences based on the methylation sequencing alignment data to obtain unique sequencing alignment sequence data; Extracting cytosine context sequences based on the unique sequencing alignment sequence data to obtain cytosine site context sequence data; Performing a bisulfite conversion judgment on the cytosine site context sequence data to generate cytosine site conversion status data; Counting the number of methylated and non-methylated sequencing fragments based on the cytosine site conversion status data for the unique sequencing alignment sequence data, respectively obtaining the number of methylated sequencing fragments and the number of non-methylated sequencing fragments; Calculating the methylation frequency based on the number of methylated sequencing fragments and the number of non-methylated sequencing fragments to generate methylation site frequency data.
6. The methylation-based tumor data processing system according to claim 5, wherein The calculation of methylation frequency based on the number of methylated sequencing fragments and the number of non-methylated sequencing fragments includes: Counting the number of sequencing fragment coverage times at the site based on the number of methylated sequencing fragments and the number of non-methylated sequencing fragments to generate cytosine site total sequencing depth data; Calculating the site methylation ratio based on the number of methylated sequencing fragments and the number of non-methylated sequencing fragments to generate an initial methylation frequency; Establishing a mapping relationship between the sequencing depth and the methylation frequency, and constructing a frequency background noise model using a preset negative binomial distribution model; Setting a low-depth threshold parameter based on the frequency background noise model; filtering the initial methylation frequency based on the cytosine site total sequencing depth data through the low-depth threshold parameter to obtain filtered methylation frequency data; Calculating the frequency confidence interval for the filtered methylation frequency data using the frequency background noise model to generate methylation confidence interval data; Performing confidence interval correction on the filtered methylation frequency data through the methylation confidence interval data to generate methylation site frequency data.
7. The methylation-based tumor data processing system according to claim 4, wherein The analysis of tumor differential sites based on the methylation site frequency data includes: Dividing the functional regions of the high-quality sequencing data using preset genomic annotation information to obtain genomic functional region list data; Dividing the genomic functional region list data into tumor / normal samples based on the methylation site frequency data and performing site frequency annotation to generate regional methylation frequency data; Calculating the average methylation levels of tumor samples / normal samples based on the regional methylation frequency data, respectively obtaining the tumor region methylation mean and the normal region methylation mean; Calculating the methylation difference between the tumor region methylation mean and the normal region methylation mean through the normal region methylation mean, and performing a significance test for the difference to obtain the regional methylation difference P value; Screening for tumor differentially methylated regions in the genomic functional region list data through the regional methylation difference P value and performing tumor differential site annotation to generate tumor differential site data.
8. The methylation-based tumor data processing system according to claim 1, wherein The tumor gene sequence analysis module includes the following functions: Performing genomic coordinate mapping based on the tumor differential site data and extracting the genomic sequence of the corresponding region to generate tumor differential region sequence data; Construct a differential methylation site expression network based on the sequence data of tumor differential regions to generate a gene co-expression network; Perform network topology analysis on the gene co-expression network to obtain node centrality indicators; Based on the node centrality indicators, identify the key gene sequences for tumorigenesis in the gene co-expression network to obtain tumor recognition gene sequences; Conduct sequence feature analysis based on the tumor recognition gene sequences to generate tumor gene sequence feature vectors; Extract the methylation level according to the tumor recognition gene sequences, and perform feature fusion on the tumor gene sequence feature vectors to generate tumor methylation feature data.
9. The methylation-based tumor data processing system according to claim 8, wherein The sequence feature analysis according to the tumor recognition gene sequences includes: Calculate the GC content of the tumor recognition gene sequences to obtain the GC content value; Identify CpG islands in the tumor recognition gene sequences and calculate the CpG island density to generate the CpG island density value; Concatenate the GC content value and the CpG island density value to generate an initial sequence feature vector; Perform regularization processing on the initial sequence feature vector to generate a sequence feature vector; Perform K-Mer decomposition on the tumor differential region sequence data and perform feature integration according to the sequence feature vectors to generate tumor gene sequence feature vectors.
10. The methylation-based tumor data processing system according to claim 1, wherein The tumor subtype classification module includes the following functions: Obtain historical tumor gene sequence feature data; Perform stratified sampling on the historical tumor gene sequence feature data to obtain a training set and a test set respectively; Perform transfer learning on a preset convolutional neural network model through the training set to generate a trained tumor risk prediction model; Perform model cross-validation on the trained tumor risk prediction model through the test set and perform parameter optimization to obtain a tumor risk prediction model; Transmit the tumor methylation feature data to the tumor risk prediction model for tumor type prediction and perform classification probability conversion to generate a tumor classification probability matrix; Identify the maximum probability value according to the tumor classification probability matrix and perform tumor subtype label identification to generate tumor type label data.
Citation Information
Cited By
Tumor typing method and system based on proteome detection
CN120564832A
A method and system for tumor typing based on proteomic detection
CN120564832B