A multi-modal pan-cancer early prediction method based on cfDNA methylation, a prediction device and an electronic device
By integrating fragmented analysis of methylation sequencing data and methylation entropy detection, a multimodal pan-cancer early prediction model was constructed, which solved the accuracy problem of joint detection of multiple cancers in existing technologies and achieved efficient early screening and personalized diagnosis and treatment of multiple cancers.
Patent Information
- Application Number
- CN202411803571.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-12-09
AI Technical Summary
Existing technologies struggle to achieve highly accurate combined detection of multiple cancers in cfDNA methylation testing, especially in early-stage cancer detection where false negatives and false positives are common, and the testing costs and difficulties are high.
By integrating fragmentation analysis and methylation entropy detection of methylation sequencing data, a multimodal pan-cancer early prediction model was constructed. Multiple analytical methods were used to perform multidimensional analysis on cfDNA samples, extract differentially methylated regions and chromosomal fragmentation features, and train multiple prediction models for joint detection.
It improves the accuracy and versatility of combined detection of multiple cancers, reduces detection costs, adapts to early screening of different types of cancer, and provides more efficient personalized treatment plans.
Smart Images

Figure BDA0005180232330000031 
Figure BDA0005180232330000131 
Figure FDA0005585241360000011
Abstract
Description
Technical Field
[0001] The present application belongs to the field of molecular biomedical technology, and specifically relates to a multimodal pan-cancer early prediction method, prediction device, and electronic equipment based on cfDNA methylation. Background Art
[0002] Cancer is one of the leading causes of death worldwide and a major public health issue that seriously threatens human health. Numerous studies have shown that early detection and timely diagnosis can effectively improve the survival rate of patients with various cancers. In the field of cancer detection, multi-cancer screening is often more efficient than single-cancer screening. This is because people who meet the criteria for cancer screening are often at potential risk for multiple cancers, and it is difficult for them to simply determine the likely site of disease based on their condition before the examination, especially for cancers occurring in the peritoneal cavity or blood. cfDNA is derived from DNA released after damaged and ruptured cells and circulates in the blood. Tumor genomes often carry characteristic gene mutation sites. Detecting mutation sites, particularly methylation sites, in cfDNA can enable cancer detection.
[0003] In order to continuously improve the performance of cfDNA methylation in cancer detection, researchers have made great efforts in technological innovation and detection model construction. Although some achievements have been made in the early detection of cancer using methylation panels, its accuracy in the joint detection of multiple cancers needs to be improved. Based on cfDNA sequencing data, a multimodal detection model is constructed by integrating multiple features such as methylation, fragmentation, and methylation entropy to more comprehensively focus on sample data from the test population, which can effectively help improve the versatility and accuracy of multi-cancer detection. Therefore, this application designs a multimodal pan-cancer detection method based on methylation panel sequencing data through multi-dimensional analysis. Summary of the Invention
[0004] In order to solve the problems existing in the prior art, the purpose of this application is to provide a multimodal pan-cancer early prediction method, prediction device and electronic equipment. Based on the cfDNA methylation difference analysis, fragmentation analysis of methylation sequencing data and detection analysis of methylation entropy are introduced, and a multimodal pan-cancer early prediction model is obtained by integrating multiple analysis methods. The model is used to analyze the sequencing samples of the test population to achieve high-accuracy joint detection of seven cancers, including lung cancer, intestinal cancer, gastric cancer, liver cancer, esophageal cancer, thyroid cancer, and ovarian cancer, providing a more accurate and feasible solution for pan-cancer early screening.
[0005] Specifically, this application involves the following aspects:
[0006] 1. A multimodal pan-cancer early prediction method, comprising: collecting multiple cfDNA samples from i types of cancer populations and healthy populations, respectively, and extracting methylation data from each of the multiple cfDNA samples; merging CpG sites based on the methylation data of the multiple cfDNA samples to obtain multiple methylation intervals, extracting differential methylation intervals between each of the i types of cancer populations and the healthy population from the multiple methylation intervals, selecting differential methylation intervals from the obtained multiple differential methylation intervals that appear in at least m types of cancer populations as candidate marker intervals, and screening the obtained multiple candidate marker intervals; training a first prediction model using the screened multiple candidate marker intervals, training a second prediction model using chromosome fragmentation features extracted based on a reference genome, and training a third prediction model using chromosome methylation entropy features; and training a multimodal pan-cancer early prediction model using the cancer population prediction values of the first, second, and third prediction models for pan-cancer early detection.
[0007] 2. The multimodal pan-cancer early prediction method according to item 1, wherein extracting methylation data of multiple cfDNA samples separately includes: extracting multiple CpG sites of the multiple cfDNA samples and the methylation value of each CpG site in the multiple CpG sites; merging CpG sites based on the methylation data of multiple cfDNA samples includes: calculating the difference in methylation value of each CpG site between each cancer population and a healthy population in i types of cancer populations; and merging corresponding CpG sites in response to the difference being not zero.
[0008] 3. The multimodal pan-cancer early prediction method according to item 1, wherein screening the obtained multiple candidate marker intervals includes: calculating an importance value for each candidate marker interval in the multiple candidate marker intervals, eliminating the corresponding candidate marker interval in response to the importance value being less than a first threshold, and retaining the corresponding candidate marker interval in response to the importance value being greater than the first threshold.
[0009] 4. The multimodal pan-cancer early prediction method according to item 1, wherein the second prediction model is trained based on the fragmentation features of chromosomes extracted from the reference genome, including: tiling the reference genome autosomes and evenly dividing them into multiple first intervals; recording the number of short-length fragments, medium-length fragments, and long-length fragments in each of the multiple first intervals; sequentially merging the multiple first intervals to obtain multiple second intervals, calculating the coverage of each of the multiple second intervals, and performing dimensionality reduction on the obtained multiple coverages; and constructing a second prediction model using the multiple coverages after dimensionality reduction.
[0010] 5. The multimodal pan-cancer early prediction method according to item 4, wherein the coverage is the short-length fragments, medium-length fragments, and long-length fragments contained in each of the multiple second intervals; and the second prediction model is constructed using elastic network regression using the multiple coverages after dimensionality reduction.
[0011] 6. According to the multimodal pan-cancer early prediction method of item 1, extracting the methylation entropy features of chromosomes to train the third prediction model includes: extracting multiple insertion fragments from the methylation data of multiple cfDNA samples to obtain methylation pattern entropy values of the multiple insertion fragments; calculating the methylation pattern entropy value of each chromosome separately; and constructing the third prediction model using logistic regression based on the methylation pattern entropy values of all chromosomes.
[0012] 7. The multimodal pan-cancer early prediction method according to item 6, wherein extracting multiple inserts from the methylation data of multiple cfDNA samples comprises: extracting inserts in which both reads read from left to right and reads read from right to left of the cfDNA samples can be compared with a reference genome in a CpG pattern; and obtaining methylation pattern entropy values for the multiple inserts using the following formula:
[0013]
[0014] Where BiEn(s) is the methylation pattern entropy of the insert, n represents the number of all CpG sites in the insert, k is a set of values from 0 to n-2, and p is the probability value of k at a given value.
[0015] 8. The multimodal pan-cancer early prediction method according to item 6, wherein calculating the methylation pattern entropy value of each chromosome separately includes: calculating the mean of the methylation pattern entropy values of all inserted fragments on each chromosome separately as the methylation pattern entropy value of each chromosome.
[0016] 9. A multimodal pan-cancer early prediction device comprises: a data acquisition unit, which collects multiple cfDNA samples from i types of cancer populations and healthy populations, and extracts methylation data from the multiple cfDNA samples; a marker screening unit, which merges CpG sites based on the methylation data of the multiple cfDNA samples to obtain multiple methylation intervals, extracts differential methylation intervals between each of the i types of cancer populations and the healthy population from the multiple methylation intervals, uses differential methylation intervals that appear in no less than m types of cancer populations among the obtained multiple differential methylation intervals as candidate marker intervals, and screens the obtained multiple candidate marker intervals; a model construction unit, which trains a first prediction model using the screened multiple candidate marker intervals, trains a second prediction model using chromosome fragmentation features extracted based on a reference genome, and trains a third prediction model using chromosome methylation entropy features; a model integration unit, which trains a multimodal pan-cancer early prediction model using the cancer population prediction values of the first prediction model, the second prediction model, and the third prediction model for pan-cancer early detection.
[0017] 10. An electronic device comprising: a processor; and a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the multimodal pan-cancer early prediction method as described in any one of items 1-8.
[0018] The multimodal pan-cancer early prediction method of this application integrates information from different modalities, such as the methylation data and chromosome fragment data mentioned in this application. The model can capture data features in more dimensions. This not only helps the model better cope with noise and interference during prediction and enhances the model's generalization ability, but also reveals hidden associations between different modalities, thereby further improving the model's prediction accuracy based on a single modality, helping doctors assess the test subject's risk of cancer and provide them with personalized diagnosis and treatment plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG2 shows a flowchart of a multimodal pan-cancer early prediction method according to an embodiment of the present application.
[0020] Figure 2 The figure shows a flowchart of a multimodal pan-cancer early prediction method according to an embodiment of the present application, which uses three modality features to construct three prediction models.
[0021] Figure 3 FIG2 is a schematic diagram of an ROC curve of a multimodal pan-cancer early prediction method on a test sample set according to an embodiment of the present application.
[0022] Figure 4 FIG2 shows a block diagram of a multimodal pan-cancer early prediction device according to an embodiment of the present application.
[0023] Figure 5 The figure shows a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] Methylation data
[0025] Methylation data is a collection of data that measures and analyzes the methylation levels of DNA or RNA molecules using various techniques. Methylation data are particularly important in cancer research, as they can reveal the epigenetic mechanisms behind changes in gene expression and help understand the occurrence and progression of cancer.
[0026] Methylation value
[0027] A metric used to measure DNA methylation levels, commonly used in biostatistics and genomics research. The methylation value reflects the proportion of methylated cytosine at a specific CpG site, ranging from 0 to 1, where 0 indicates complete unmethylation and 1 indicates complete methylation. In practical applications, methylation values can be used to compare methylation differences between different samples and cell types, and are of great significance in disease diagnosis and gene expression regulation research.
[0028] The present application is further described below with reference to examples. It should be understood that the examples are only used to further illustrate and explain the present application and are not intended to limit the present application.
[0029] Unless otherwise defined, technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art. Although methods and materials similar or identical to those described herein may be used in experiments or practical applications, the materials and methods are described herein below. In the event of a conflict, the present specification, including definitions, will prevail. In addition, the materials, methods, and examples are provided for illustrative purposes only and are not intended to be limiting. The present application is further described below with reference to specific examples, which are not intended to limit the scope of this application.
[0030] Application Overview
[0031] As mentioned above, in the field of cancer detection, due to the heterogeneity of cfDNA abnormalities across different cancer types, subtypes, stages, and etiologies, technologies using cfDNA sequencing data to characterize pan-cancer genetic signatures are subject to a certain degree of false negatives and false positives. Due to inconsistent detection standards for tumor tissue and ctDNA, the consistency data provided by existing studies varies widely, making it difficult to provide useful information for extracting cancer signatures. This makes it difficult to advance experiments to refine effective biomarker combinations for specific cancer types, thus limiting the use of cfDNA tests, such as cfDNA methylation tests, in predicting individual cancer risk. Furthermore, in some cancers at early stages or at the minimal residual disease stage, ctDNA levels in plasma are extremely low, significantly increasing the cost and difficulty of detection.
[0032] Patent CN114045345A proposes a genomic cancerous change information detection system that uses enzymatic transfection to perform whole-genome methylation sequencing on cfDNA from plasma samples. This system analyzes genomic methylation density, fragment length distribution, fragment 5' end motifs, and / or chromosomal stability, enabling early detection and screening for multiple cancers. However, this monitoring system neglects to capture subtle changes in DNA or RNA molecules within the genome, including changes in fragment length and the loss or gain of specific sequences. These fragmented features often reflect changes in the entire genome or transcriptome, and their analysis does not rely on specific cancer markers or genes. This makes it difficult for this monitoring system to provide strong support for pan-cancer detection.
[0033] Patent CN116356021A provides GutSeer, a multi-cancer early screening and localization technology for five digestive system cancers with a high mortality rate, and proves that using a relatively small second-generation sequencing panel can utilize multiple dimensional features including methylation, copy number changes, and terminal motifs to achieve more accurate cancer detection.
[0034] Methylation entropy is the entropy value generated during the DNA methylation process, that is, the degree of disorder or uncertainty of the DNA methylation state. Since methylation entropy is an indicator to measure the complexity and stability of DNA methylation patterns, it can be used as a specific indicator to reflect the distribution and variation of methylation sites on cfDNA molecules. This makes the methylation entropy feature helpful to help the prediction model understand the differences in methylation patterns between target cancer patients, thereby learning the heterogeneity between these cancers. Therefore, relative to the detection method of methylation panel sequencing feature construction of CN116356021A, the present application proposes to construct a multimodal detection model based on the cfDNA methylation rate feature and the genomic fragmentation feature in combination with the methylation entropy feature, so that the prediction method, device and electronic device of the present application are adapted to detect cancer types not limited to the digestive system, but cover more cancers in multiple systems, and can achieve a wider range of high-accuracy joint detection compared to the above-mentioned method or system.
[0035] Specifically, the present application provides a multimodal pan-cancer early prediction method, prediction device and electronic device, which collect cfDNA samples from i types of cancer populations and healthy populations, and extract methylation data of the cfDNA samples respectively; based on the methylation data of the cfDNA samples, the CpG sites in the cfDNA samples are merged to obtain multiple methylation intervals, and multiple differential methylation intervals between each cancer population in the i types of cancer populations and the healthy population are screened from the multiple methylation intervals, and the differential methylation intervals that appear in no less than m types of cancer populations in the multiple differential methylation intervals are extracted as multiple candidate marker intervals, and the multiple candidate marker intervals are screened; a first prediction model is trained using the multiple candidate marker intervals, and a second prediction model is trained by extracting fragmentation features on the reference genome, and a third prediction model is trained by extracting methylation entropy features of chromosomes; a multimodal pan-cancer early prediction model is trained by the cancer population prediction values of the first prediction model, the second prediction model and the third prediction model for pan-cancer early detection.
[0036] By screening multiple differential methylation intervals between i cancer populations and healthy populations, and further extracting differential methylation intervals that appear simultaneously in cfDNA samples of cancer populations belonging to multiple types of cancer as multiple candidate marker intervals in these differential methylation intervals, these candidate marker intervals can better reflect the overall genetic characterization characteristics of cancer in i cancer populations, providing a basic predictive model for the joint detection of i cancers; by extracting the fragmentation characteristics of the reference genome, we can capture subtle changes in genomic DNA fragments, which are closely related to the occurrence and development of multiple cancers, thereby improving the sensitivity and accuracy of detecting each cancer, and at the same time providing rich biological information to reveal the heterogeneity between i cancers, providing a predictive model for cancer differentiation between i cancers.
[0037] Furthermore, entropy reflects the degree of disorder in a system: the more disordered a system, the greater its entropy; the more ordered a system, the smaller its entropy. Methylation pattern entropy refers to the overall state of the methylation pattern at CpG sites on the insert, reflecting the distribution and variability of the methylation state at CpG sites and can be used to assess epigenetic heterogeneity within the cell population containing cfDNA. Therefore, extracting methylation pattern entropy features is used to comprehensively assess the complexity and stability of DNA methylation patterns and accurately identify methylation abnormalities associated with various cancers. Predictive models constructed using methylation entropy features can help predictive models based on methylation panel data features and genomic fragment features further improve the accuracy of overall detection and classification of these cancers, thereby providing a highly feasible pan-cancer early screening method.
[0038] After introducing the basic principles of the present application, various non-limiting embodiments of the present application will be described in detail with reference to the accompanying drawings.
[0039] Exemplary Methods
[0040] Figure 1 The figure illustrates a multimodal pan-cancer early prediction method according to an embodiment of the present application.
[0041] like Figure 1 As shown, the multimodal pan-cancer early prediction method according to an embodiment of the present application includes the following steps.
[0042] S110, collects cfDNA samples of i kinds of cancer populations and healthy populations respectively, and extracts methylation data of cfDNA samples respectively.In the present application, methylation data refers to the data obtained by routine analysis of fastq data obtained using methylation panel, preferably a processable bam file. According to the multimodal pan-cancer early prediction method of the present application embodiment, up to 7 kinds of cancer populations can be included, namely lung cancer, intestinal cancer, gastric cancer, liver cancer, esophageal cancer, thyroid cancer, and ovarian cancer. Based on these 7 kinds of cancer populations, all samples of the training set in the cancer sample obtained by sequencing the methylation data sample of its cfDNA sample by the methylation panel are candidate groups, and all healthy samples are control groups, and methylation intervals are screened. It is understandable that methylation data can be data obtained by further screening the methylation sequencing data obtained by the methylation panel, such as data quality preprocessing and evaluation (fastp software), genome alignment (Bismark software) or methylation sequencing data after removing the duplicate data brought by sample / experimental technology. The reference genome used for sequencing is the human genome. In the art, there are multiple versions of human genome gene sequencing versions. The commonly used version is hg19. Those skilled in the art can select the appropriate version.
[0043] In this way, the methylation data samples obtained for the seven types of cancer populations can be the position information of all CpG sites and the methylation values of these CpG sites obtained by methylation panel sequencing from all cancer population samples and healthy population samples; the CpG sites can be CpG sites on the cancer-high-correlation target intervals obtained based on existing public research, public testing or sequencing products, such as sites on the intervals used in Boercheng's seven-cancer NGS methylation detection kit, or CpG sites on the entire genome. These CpG sites can be merged according to certain rules to form methylation intervals containing multiple CpG sites, such as intervals with high methylation levels or intervals with low methylation levels, and these methylation intervals may therefore have differentiated methylation status manifestations in cancer patients and healthy people, thereby becoming favorable methylation markers.
[0044] S120, based on the methylation data of the cfDNA sample, CpG sites are merged to obtain multiple methylation intervals, and multiple differential methylation intervals between each cancer population and the healthy population are screened from the multiple methylation intervals. The differential methylation intervals that appear in no less than m cancer populations are extracted from the multiple differential methylation intervals as multiple candidate marker intervals, and the multiple candidate marker intervals are screened. As can be seen from the above, the position information and methylation values of all CpG sites of the two populations are obtained through step S110. The expression difference of each CpG site in the two populations can be obtained by calculating the difference in methylation values of the cancer population and the healthy population at each CpG site; wherein, the difference in methylation values can be the difference between the average methylation value of each CpG site in one cancer population and the average methylation value of the healthy population.
[0045] In this way, differentiated CpG sites are obtained, and differentiated methylation intervals can be constructed based on this. A feasible method for merging differentiated CpG sites to obtain methylation intervals is to merge the corresponding CpG sites in response to each CpG site whose methylation value difference is not zero, that is, to merge these adjacent differentiated CpG sites in sequence. It should be noted that the methylation interval obtained in this way may have the problem of being too long. Since the length of human cfDNA is generally around 167bp, it is necessary to further set the maximum length of the methylation interval so that it does not significantly exceed the length of cfDNA. For example, the maximum length of the methylation interval can be set to 200bp, thereby ensuring that each methylation interval is close to 167bp of cfDNA, so that the methylation level of the CpG site in the calculated interval is closer to the actual level, ensuring data quality. In addition, since the fewer the number of CpG sites in a given interval, the higher the false positive probability of the methylation value of the CpG site obtained by the existing technology, it is also possible to additionally set the number of CpG sites in each methylation interval to at least 3, and to screen the CpG sites in the methylation interval by any feasible CpG site outlier processing or missing value processing method, so as to further improve the robustness of the obtained methylation interval.
[0046] In some embodiments, screening multiple methylation intervals to obtain multiple differentially methylated intervals between each cancer population and a healthy control group within i types of cancer includes the following screening method: specifying that the detection sensitivity of the methylation intervals for each cancer population within i types of cancer is no less than a given threshold value at a given detection specificity. While maintaining a certain level of specificity and sensitivity for a given test, for example, in a preferred embodiment, ensuring a sensitivity of no less than 70%, such that more than 70% of the methylation intervals from the cancer population exhibit methylation differences between the populations, and ensuring a specificity of no less than 80%, such that more than 80% of the methylation intervals from the healthy control group exhibit methylation differences between the populations, differentially methylated intervals with high sensitivity and specificity can be obtained, while also reducing the impact of outliers in the methylation intervals on the differences and ensuring the stability of the differences.
[0047] In addition, the following screening method can be further included: specifying that the detection AUC for the cancer population and the healthy population based on the methylation interval detection of each cancer type i is not less than a given threshold; and that the difference in the average methylation values of all methylation sites within the methylation interval between each cancer population and the healthy population is not less than a given threshold. For example, in a preferred embodiment, the AUC for detecting the cancer population using the methylation interval is guaranteed to be no less than 0.7, and the absolute value of the difference in the average methylation values of methylation sites within the methylation interval between the cancer population and the healthy population is guaranteed to be no less than 0.02. In this way, while ensuring that the screened differentially methylated intervals have good population classification performance, methylation direction interference can be avoided, i.e., the methylation levels of each CpG site in these differentially methylated intervals are different between the two populations, and the methylation direction of the difference is not specified, thereby ensuring that some differentially methylated intervals are not screened out due to the total difference in CpG sites being too small or zero.
[0048] In this way, multiple differential methylation intervals with high population classification performance are obtained from multiple methylation intervals through one or more of the aforementioned multiple screening methods. Since these differential methylation intervals are used for pan-cancer detection of i different types of cancer, in order to improve the detection speed and the generalization performance of the constructed prediction model, it is also necessary to screen and obtain candidate marker intervals with wider population adaptability relative to a single cancer population. Specifically, in some favorable embodiments, differential methylation intervals that appear in no less than m types of cancer populations from multiple differential methylation intervals can be extracted as multiple candidate marker intervals, and the value of m can be 2, 3, 4, 5, 6 or 7; for example, differential methylation intervals that appear in cfDNA samples of at least two types of cancer populations are selected as candidate marker intervals, so that the candidate marker intervals can show good classification performance for at least two cancer types.
[0049] Considering that candidate marker intervals based on a larger targeted interval or a merging of CpG sites across the entire genome may still have a high data dimension even after being screened by one or more of the aforementioned screening methods, placing an excessive burden on the prediction model to learn classification features and optimize parameters, the multimodal pan-cancer early prediction method according to an embodiment of the present application further includes screening the candidate marker intervals, i.e., interval dimensionality reduction. According to some feasible implementation schemes, the screening method may include: calculating an importance value for each candidate marker interval in a plurality of candidate marker intervals, deleting the corresponding candidate marker interval from the plurality of candidate marker intervals in response to the importance value being no greater than a first threshold, and retaining the corresponding candidate marker interval from the plurality of candidate marker intervals in response to the importance value being greater than the first threshold.
[0050] Here, the importance value is an indicator that can reflect the degree of contribution of each candidate marker interval to the prediction result of the prediction model. The larger the importance value, the greater the influence of the candidate marker interval on the prediction result of the model. The importance value can be obtained in a variety of ways, for example, by inputting multiple candidate marker intervals into a linear regression model, a Lasso model, a random forest, a gradient boosting tree, etc. to screen intervals with high importance values. According to a feasible implementation scheme, multiple candidate marker intervals are input into a random forest model and the importance value of each candidate marker interval is calculated respectively. The multiple candidate marker intervals are sorted according to the importance value and the first l candidate marker intervals whose importance value is greater than the first threshold are selected, and all candidate marker intervals starting from the l+1th are deleted; the value of l is preferably 45, so as to ensure that the classification performance and construction cost of the simple two-class prediction model can be optimized.
[0051] In step S130, a first prediction model is trained using multiple candidate marker intervals, a second prediction model is trained based on chromosome fragmentation features extracted from the reference genome, and a third prediction model is trained based on chromosome methylation entropy features. Step S130 can be further divided into three prediction steps: the first prediction model, the second prediction model, and the third prediction model.
[0052] In step S1301, a first prediction model is trained using multiple candidate marker intervals, and a prediction value for the cancer population is obtained based on the first prediction model. Here, since multiple candidate marker intervals have been obtained after multiple screening and dimensionality reduction, a concise and efficient binary classification model can be selected, and the binary classification model is trained using the candidate marker intervals to obtain the first prediction model. Therefore, in one advantageous embodiment, a logistic regression algorithm is selected to construct a binary classification model for cancer and healthy populations. Logistic regression has relatively low computational complexity and is suitable for processing large-scale sequencing data. In addition, it can be used in conjunction with an optimization algorithm to quickly iterate classification model parameters, making it suitable for real-time analysis and processing of cancer detection.
[0053] In this way, the first prediction model is obtained through the logistic regression algorithm. By providing the model with a fully connected layer or any possible classifier, it can output the probability value of each individual in the cancer population for each of the i types of cancer, which will serve as part of the cancer population prediction value for the final training of the multimodal pan-cancer early prediction model.
[0054] Step S1302, based on the fragmentation characteristics of the chromosomes extracted from the reference genome, a second prediction model is trained, and a prediction value for the cancer population is obtained based on the second prediction model. Specifically, the reference genome autosomes selected for methylation panel sequencing are tiled to divide them into adjacent, non-overlapping first intervals, for example, multiple intervals of 100kb in length. In particular, the length of ctDNA fragments in plasma samples will be shorter than the length of normal cfDNA fragments. Specifically, the distribution of ctDNA in cancer populations and cfDNA in healthy populations is different in the three types of fragments of 90-150bp, 180-220bp and 250-320bp.
[0055] Therefore, in each first interval, the short length fragment is set as the short fragment, and its length can be set between 100-150bp, which is included in the first type of differential fragments; the medium length fragment is the middle fragment, and its length is set between 150-260bp, including the second type of differential fragments; and the long length fragment is the long fragment, and its length is set between 260-320bp, which is included in the third type of differential fragments; in this way, more differential fragments are obtained through the length type, so that the fragmentation feature has higher specificity. In addition, the above three fragments can also be integrated to obtain the overall length fragment, that is, the nfrags fragment, whose interval length is between 100-320bp. Then, the number of short, middle, long, and nfrags in each first interval of each interval is counted to obtain the fragmentation feature.
[0056] Then, the first intervals are sequentially merged to obtain multiple second intervals of a certain length, for example, multiple 1MB second intervals, resulting in 2608 non-overlapping second intervals. A coverage is defined for each second interval: j represents the jth second interval among the multiple second intervals, and the value of j can be 1, 2, ..., 2608. Then, the various types of fragments (short, middle, long, and nfrags) and their corresponding numbers in the jth second interval can be set as fragment features and corresponding feature values. For example, a total of 10,432 fragment features and corresponding numbers of fragments are obtained from the 2608 second intervals. In this way, the elastic net regression algorithm can be used to construct a fragmented binary classification model for cancer and healthy subjects as the second prediction model.
[0057] As can be understood, after obtaining the segment features, the PCA method is used to reduce the dimensionality of the massive segment features, constructing a segment feature matrix with principal components that can explain 95% of the variance. The second prediction model is then constructed using the elastic net regression algorithm. Elastic net regression is chosen because the number of segment features in the second interval may be much higher than the number in the first interval, and elastic net regression can effectively select segment features to avoid overfitting of the second prediction model.
[0058] In this way, the second prediction model can better handle the high collinearity problem between chromosome segments than the first prediction model. This is because chromosome segments often have complex interactions and associations, which the first prediction model constructed using candidate marker intervals does not take into account. In addition, elastic net regression is more suitable for building sparse models to reduce model complexity and computational complexity. Therefore, the second prediction model can show higher specificity in classification and has a faster fitting speed. By constructing the second prediction model and integrating it with the first prediction model, the specificity of the first prediction model can be effectively improved.
[0059] Step S1303 extracts chromosome methylation entropy features to train a third prediction model, and then obtains a cancer population prediction value based on the third prediction model. Specifically, multiple inserts from the methylation data of the cfDNA sample are extracted based on the CpG pattern to obtain methylation pattern entropy values for each insert. The methylation pattern entropy value for each chromosome is calculated as the methylation entropy feature. First, inserts are extracted from the methylation panel sequencing data of the cfDNA sample, and unqualified inserts are filtered according to specific criteria.
[0060] It should be noted that extracting insert fragments refers to extracting fragments in which both the reads read from left to right and the reads read from right to left during the sequencing process can be compared with the reference genome (e.g., the same fragment on the genome). Examples include fragments of about 300 bp in length that have been interrupted, or fragments of about 170 bp in length for free DNA. In some advantageous embodiments, filtering unqualified insert fragments may include: removing insert fragments containing less than 3 or more than 32 CpG sites and / or removing insert fragments with missing CpG sites in the middle, thereby eliminating erroneous fragments or fragments with too low methylation levels.
[0061] In particular, the CpG pattern refers to the pattern of obtaining the methylation status of the entire CpG site on the insert. Using the CpG pattern, we can statistically obtain the chromosome position, the search position of the first and last CpG sites, the CpG pattern diagram (the methylation status of the entire CpG site on the insert, that is, the pattern diagram of the methylation value size), and the frequency of CpG pattern occurrence. Then, the methylation pattern entropy value of each insert is calculated separately on a chromosome basis. The calculation method is shown in the following formula:
[0062]
[0063] Wherein, BiEn(s) represents the methylation pattern entropy of the inserted fragment, n represents the number of all CpG sites in the inserted fragment, k is a set of values from 0 to n-2, and p is the probability value of k under a given value. In this way, the calculated methylation pattern entropy can judge the disorder or randomness of the methylation pattern of the inserted fragment. The inserted fragments screened by the multimodal pan-cancer early prediction method according to the embodiment of the present application are the collective effect of the entropy of cfDNA fragments from different sources, while the methylation pattern entropy in the prior art is calculated based on the site. This is not like the inserted fragments extracted in the present application, which is a methylation pattern entropy of cfDNA shed from a certain tumor cell. Instead, it is an entropy of several continuous CpG sites of DNA with multiple sources as a whole. Therefore, the methylation pattern entropy of the present application will not cover up the CpG site information of the target tumor cell cfDNA.
[0064] In this way, the methylation pattern entropy value of each chromosome or specified interval can be further calculated using the following formula:
[0065] RE=1 / N*(e1+e2+e3+…+eN)
[0066] (Formula 2)
[0067] Here, RE represents the methylation entropy value of a region, such as a chromosome, N represents the total number of inserts within the region, and e represents the methylation entropy value of each insert within the region. Thus, the mean of the methylation pattern entropy values of all inserts for each chromosome can be used as the methylation pattern entropy value for that chromosome. Finally, using the methylation entropy value as the methylation entropy feature, a logistic regression algorithm was used to construct a classification model for cancer and healthy subjects as the third prediction model.
[0068] In another embodiment, multiple target intervals can be selected and the average of the methylation pattern entropy values of all inserted fragments in each target interval can be used as the methylation pattern entropy value of each target interval. The target interval can also be a high-correlation target interval for cancer obtained based on existing public research, public detection or sequencing products, such as the high-throughput methylation sequencing target interval used in Boercheng's seven-cancer NGS methylation detection kit. This can reduce the calculation cost of the methylation pattern entropy value of the inserted fragment, prevent data redundancy and speed up the fitting speed of the third prediction model; it can be understood that researchers in this field can select appropriate target intervals according to actual research or production conditions and use the multimodal pan-cancer early prediction method of the implementation scheme of the fundamental application to calculate the methylation pattern entropy value of the inserted fragment.
[0069] Step S140: A multimodal pan-cancer early prediction model is trained using the cancer population prediction values of the first prediction model, the second prediction model, and the third prediction model for pan-cancer early detection. For the case of a small sample size and a large number of features in this application, a base model constructed from multiple samples, i.e., the cancer population prediction values of the first prediction model and the third prediction model, is selected to combine to form a cancer population prediction feature set. This cancer population prediction feature set is then passed as input to the multimodal pan-cancer early prediction model for training, and the final prediction result is output. This helps integrate base models of different modalities, making the multimodal pan-cancer early prediction model adaptable to different types of data and classification problems. Furthermore, since the sample predictions of the three base models can be fully considered during training, the risk of overfitting that may exist in a single model is reduced, and model complementarity is utilized to provide more stable classification results and improve robustness.
[0070] In one advantageous embodiment, a binary multimodal pan-cancer early prediction model can be constructed using a logistic regression algorithm, and predictions can be made for each type of cancer in each cancer population. It should be noted that after obtaining the predicted values for the cancer population through steps S110-S130, i.e., the predicted labels for the cancer population by at least one prediction model, any feasible algorithm can be used to train a three-class or even more prediction model as a multimodal pan-cancer early prediction model, or the predicted values for the cancer population can be used to fine-tune a large language model to obtain multi-classification results. This application does not impose any restrictions on this.
[0071] The present application provides a multimodal pan-cancer early prediction method based on cfDNA methylation, which includes a first prediction model constructed using the methylation level of the methylation interval detected by a methylation panel, a fragmented second prediction model constructed using genomic fragment features, and a third prediction model constructed using chromosome methylation entropy. By integrating the three modal models, the present application obtains a multimodal pan-cancer early prediction model, which can show better pan-cancer classification accuracy than any single modality prediction model in the detection of multiple cancer types, while also maintaining high sensitivity and specificity, and is suitable for early non-invasive screening of multiple tumors.
[0072] Exemplary devices
[0073] Figure 4 FIG2 is a block diagram of a multimodal pan-cancer early prediction device according to an embodiment of the present application. Figure 4 As shown, the multimodal pan-cancer early prediction device 200 according to an embodiment of the present application includes:
[0074] A data collection unit 210 collects multiple cfDNA samples from i types of cancer populations and healthy populations, and extracts methylation data of the multiple cfDNA samples respectively; a marker screening unit 220 merges CpG sites based on the methylation data of the multiple cfDNA samples to obtain multiple methylation intervals, extracts differential methylation intervals between each cancer population and healthy population in the i types of cancer populations from the multiple methylation intervals, uses the differential methylation intervals that appear in no less than m types of cancer populations among the obtained multiple differential methylation intervals as candidate marker intervals, and screens the obtained multiple candidate marker intervals; a model construction unit 230 trains a first prediction model using the screened multiple candidate marker intervals, trains a second prediction model based on the fragmentation features of chromosomes extracted from the reference genome, and trains a third prediction model by extracting the methylation entropy features of chromosomes; and a model integration unit 240 trains a multimodal pan-cancer early prediction model using the cancer population prediction values of the first prediction model, the second prediction model, and the third prediction model for pan-cancer early detection.
[0075] Here, those skilled in the art will appreciate that the specific functions and operations of the various units and modules in the multimodal pan-cancer early prediction device 200 have been described above with reference to FIG. Figures 1 to 2 The multimodal pan-cancer early prediction method has been introduced in detail in the description of the present invention, and therefore, its repeated description will be omitted.
[0076] As described above, the multimodal pan-cancer early prediction apparatus 200 according to the embodiments of the present application can be implemented in various terminal devices, such as a server for training any prediction model or a multimodal pan-cancer early prediction model. In one example, the multimodal pan-cancer early prediction apparatus 200 according to the embodiments of the present application can be integrated into the terminal device as a software module and / or a hardware module. For example, the multimodal pan-cancer early prediction apparatus 200 can be a software module in the terminal device's operating system, or an application developed specifically for the terminal device. Of course, the multimodal pan-cancer early prediction apparatus 200 can also be one of the terminal device's many hardware modules.
[0077] Alternatively, in another example, the multimodal pan-cancer early prediction apparatus 200 and the terminal device may be separate devices, and the multimodal pan-cancer early prediction apparatus 200 may be connected to the terminal device via a wired and / or wireless network and transmit interactive information in a predetermined data format.
[0078] Exemplary electronic devices
[0079] Below, reference Figure 5 To describe the electronic device according to the embodiment of the present application.
[0080] Figure 5 The figure shows a block diagram of an electronic device according to an embodiment of the present application.
[0081] like Figure 5 As shown, the electronic device 10 includes one or more processors 11 and a memory 12 .
[0082] The processor 13 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0083] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute the program instructions to implement the multimodal pan-cancer early prediction method of each embodiment of the present application described above and / or other desired functions. Various content such as cfDNA samples, candidate marker intervals, fragmentation features, methylation entropy features, etc. may also be stored in the computer-readable storage medium.
[0084] In one example, the electronic device 10 may further include an input device 13 and an output device 14 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0085] The input device 13 may include, for example, a keyboard, a mouse, and the like.
[0086] The output device 14 can output various information to the outside, including the trained multimodal pan-cancer early prediction model, etc. The output device 14 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.
[0087] Of course, to simplify, Figure 5 Only some of the components related to the present application in the electronic device 10 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device 10 may further include any other appropriate components according to specific application scenarios.
[0088] Example
[0089] This application provides general and / or specific descriptions of the materials and experimental methods used in the experiments. Raw materials or instruments used without manufacturer indication are all commercially available conventional raw materials or instruments.
[0090] Example 1: Construction of the first prediction model
[0091] For each of the seven types of cancer samples provided by Boercheng, including lung cancer, intestinal cancer, gastric cancer, liver cancer, esophageal cancer, thyroid cancer, and ovarian cancer, cfDNA samples were collected and divided into training set samples and test set samples. The candidate marker intervals were screened for the training set samples according to the following method:
[0092] Using the human genome hg19 as the reference genome, methylation sequencing was performed on cfDNA samples using the intervals used by the Boercheng Seven Cancer NGS Methylation Detection Kit to determine the methylation intervals. For highly methylated intervals, the sensitivity thresholds (Sensitivity ≥ 0.6), the difference in methylation values (DELTA) of CpG sites within the methylation interval, and the significance p-value thresholds for CpG sites were set to ≥ 0.6, ≥ 0.02, and p-value < 0.01, respectively.
[0093] For hypomethylated intervals, the sensitivity, delta value, and p-value thresholds were set to ≥0.55, ≤-0.02, and <0.01, respectively. The top 200 differentially methylated intervals were selected based on the absolute delta value. If there were fewer than 200 intervals, all were selected. Differentially methylated intervals that appeared in at least two cancer types were then selected as candidate marker intervals. The resulting candidate marker intervals contained 152 hypermethylated intervals and 313 hypomethylated intervals. Finally, the importance of the candidate marker intervals was calculated using a random forest algorithm, and 45 candidate marker intervals with importance values greater than 0.0055 were selected as features for post-selection modeling.
[0094] A first prediction model for binary classification was constructed using a logistic regression algorithm. Classification predictions were performed on the test set samples, and the model's predictive performance was evaluated. The results showed an AUC of 0.97, a sensitivity of 85.7%, and a specificity of 96.6% in the test set.
[0095] Example 2 Construction of the Second Prediction Model
[0096] The reference genome autosomes were tiled into adjacent, non-overlapping first intervals of 100 kb in length. Short fragments were designated short (100-150 bp), medium fragments were designated middle (150-260 bp), long fragments were designated long (260-320 bp), and nfrags were designated 100-320 bp. The number of each type of fragment (short, middle, long, and nfrags) was counted. The 100 kb first intervals were then merged into 1 MB second intervals, resulting in 2608 non-overlapping second intervals. j represents the jth second interval, and its value on the reference genome autosomes is 1, 2, ..., 2608. The number of each type of fragment (short, middle, long, nfrags) in the jth second interval is the coverage of that second interval. All coverages were used as fragment features, resulting in a total of 10,432 features.
[0097] Then, PCA was used to reduce the dimensionality of the fragment features. A fragment feature matrix was constructed using principal components that could explain 95% of the variance. Elastic Net Regression was then used to construct a binary classification discriminant model for the fragments, the second prediction model. Classification predictions were performed on the test set samples, and the model's predictive performance was evaluated. The evaluation results showed an AUC of 0.95, a sensitivity of 79.1%, and a specificity of 98.3% in the test set.
[0098] Example 3: Construction of the third prediction model
[0099] The methylation pattern entropy of each insert was calculated for each chromosome. Specifically, all inserts were extracted and unqualified inserts were filtered out. The methylation pattern entropy of each insert was used as the methylation entropy feature. A logistic regression algorithm was used to construct a binary classification discriminant model based on methylation entropy, the third prediction model. Classification predictions were performed on the test set samples, and the model's predictive performance was evaluated. The evaluation results showed an AUC of 0.85, a sensitivity of 65.9%, and a specificity of 84.5% in the test set.
[0100] Example 4 Construction of a multimodal pan-cancer early prediction model
[0101] Using cfDNA samples as input, the first prediction model, the second prediction model, and the third prediction model were used to obtain the positive predictive values of the seven cancer populations in Example 1. All positive predictive values were used as features, and a logistic regression algorithm was used to construct a multimodal two-class discrimination model, namely a multimodal pan-cancer early prediction model. The test set samples were classified and predicted, and the model prediction performance was evaluated. The evaluation results showed that the multimodal model had good cancer and non-cancer discrimination capabilities, and the ROC curve of the test set was as follows: Figure 3 As shown, the AUC in the test set was 0.98, the sensitivity was 89.0%, and the specificity was 98.3%.
[0102] Combined with Example 1-Example 4 and Figure 3 As can be seen, after integrating the three basic models to form a multimodal pan-cancer early prediction model, its prediction sensitivity for cancer samples has improved compared to the basic models, while its prediction specificity for healthy samples has remained at the highest level compared to the basic models. Therefore, the multimodal pan-cancer early prediction model of this application can be trained using small sample data and accurately distinguish plasma samples from seven cancer types and healthy people.
[0103] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.
[0104] In this application, words such as "including," "comprising," "having," and the like are open-ended words that mean "including but not limited to," and are used interchangeably therewith. The words "or" and "and" used herein mean the words "and / or," and are used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as, but not limited to," and is used interchangeably therewith.
[0105] It should also be noted that in the methods, systems, and devices of the present application, each step or module can be decomposed and / or recombined, and such decompositions and / or recombinations should be regarded as equivalent solutions of the present application.
[0106] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A multimodal pan-cancer early prediction method, comprising: Collecting multiple cfDNA samples from i types of cancer populations and healthy populations, respectively, and extracting methylation data of the multiple cfDNA samples; Merging CpG sites based on the methylation data of the multiple cfDNA samples to obtain multiple methylation intervals, extracting differential methylation intervals between each cancer population and a healthy population in i types of cancer populations from the multiple methylation intervals, selecting differential methylation intervals that appear in at least two of the required limited types of cancer populations among the multiple obtained differential methylation intervals as candidate marker intervals, and screening the multiple obtained candidate marker intervals; The first prediction model is trained using the screened multiple candidate marker intervals, the second prediction model is trained based on the chromosome fragmentation features extracted from the reference genome, and the third prediction model is trained based on the chromosome methylation entropy features extracted; Training a multimodal pan-cancer early prediction model using the cancer population prediction values of the first prediction model, the second prediction model, and the third prediction model for pan-cancer early detection; The extracting of the methylation entropy features of the chromosome to train the third prediction model comprises: Extracting multiple insert fragments from the methylation data of the multiple cfDNA samples to obtain methylation pattern entropy values of the multiple insert fragments; Calculate the methylation pattern entropy value of each chromosome separately; and The third prediction model is constructed by using logistic regression based on the methylation pattern entropy values of all chromosomes; Extracting multiple inserts from the methylation data of the multiple cfDNA samples includes: The reads extracted from the cfDNA sample in the CpG pattern, read from left to right and read from right to left, can both be compared with the insert fragments on the reference genome; The formula for obtaining the methylation pattern entropy values of the multiple inserted fragments is as follows: Where BiEn(s) is the methylation pattern entropy of the insert, n represents the number of all CpG sites in the insert, k is a set of values from 0 to n-2, and p is the probability value of k at a given value.
2. The multimodal pan-cancer early prediction method according to claim 1, wherein: The extracting methylation data of the plurality of cfDNA samples respectively includes: extracting a plurality of CpG sites from the plurality of cfDNA samples and a methylation value of each CpG site in the plurality of CpG sites; The merging of CpG sites based on the methylation data of the multiple cfDNA samples includes: Calculate the difference in methylation value of each CpG site between each cancer population and the healthy population in each type of cancer population; In response to the difference being non-zero, corresponding CpG sites are merged.
3. The multimodal pan-cancer early prediction method according to claim 1, wherein: The screening of the obtained multiple candidate marker intervals includes: An importance value of each candidate marker interval in the plurality of candidate marker intervals is calculated, and the corresponding candidate marker interval is eliminated in response to the importance value being not greater than a first threshold, and the corresponding candidate marker interval is retained in response to the importance value being greater than the first threshold.
4. The multimodal pan-cancer early prediction method according to claim 1, wherein: The training of the second prediction model based on the fragmentation features of the chromosome extracted from the reference genome includes: The reference genome autosomes are tiled and evenly divided into multiple first intervals; Recording the number of short-length segments, medium-length segments, and long-length segments of each first interval in the plurality of first intervals; sequentially merging the plurality of first intervals to obtain a plurality of second intervals, calculating the coverage of each of the plurality of second intervals, and performing dimensionality reduction on the obtained plurality of coverages; and The second prediction model is constructed using the multiple coverages after dimensionality reduction.
5. The multimodal pan-cancer early prediction method according to claim 4, wherein: The coverage is the number of short-length fragments, medium-length fragments, and long-length fragments contained in each second interval of the plurality of second intervals; The second prediction model is constructed using elastic network regression based on the multiple coverages after dimensionality reduction.
6. The multimodal pan-cancer early prediction method according to claim 1, wherein: Calculating the methylation pattern entropy value of each chromosome separately includes: The mean of the methylation pattern entropy values of all inserted fragments on each chromosome was calculated as the methylation pattern entropy value of each chromosome.
7. A multimodal pan-cancer early prediction device, comprising: a data collection unit, collecting a plurality of cfDNA samples from i types of cancer populations and healthy populations, and extracting methylation data of the plurality of cfDNA samples; a marker screening unit, merging CpG sites based on the methylation data of the plurality of cfDNA samples to obtain a plurality of methylation intervals, extracting differential methylation intervals between each of i cancer populations and a healthy population from the plurality of methylation intervals, selecting differential methylation intervals that appear in at least m cancer populations from the plurality of obtained differential methylation intervals as candidate marker intervals, and screening the plurality of obtained candidate marker intervals; A model construction unit is configured to train a first prediction model using the screened multiple candidate marker intervals, train a second prediction model based on chromosome fragmentation features extracted from a reference genome, and train a third prediction model based on chromosome methylation entropy features; a model integration unit, which trains a multimodal pan-cancer early prediction model for pan-cancer early detection using the cancer population prediction values of the first prediction model, the second prediction model, and the third prediction model; The extracting of the methylation entropy features of the chromosome to train the third prediction model comprises: Extracting multiple insert fragments from the methylation data of the multiple cfDNA samples to obtain methylation pattern entropy values of the multiple insert fragments; Calculate the methylation pattern entropy value of each chromosome separately; and The third prediction model is constructed by using logistic regression based on the methylation pattern entropy values of all chromosomes; Extracting multiple inserts from the methylation data of the multiple cfDNA samples includes: The reads extracted from the cfDNA sample in the CpG pattern, read from left to right and read from right to left, can both be compared with the insert fragments on the reference genome; The formula for obtaining the methylation pattern entropy values of the multiple inserted fragments is as follows: Where BiEn(s) is the methylation pattern entropy of the insert, n represents the number of all CpG sites in the insert, k is a set of values from 0 to n-2, and p is the probability value of k at a given value.
8. An electronic device comprising: processor; as well as A memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor is caused to perform the multimodal pan-cancer early prediction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
CfDNA targeted methylation sequencing multi-dimensional feature-based common digestive system cancer early detection technology
CN116356021A
Screening method and device for methylation diagnosis markers for early diagnosis of colon cancer
CN117524302A
Method and system for screening pan cancer markers based on cfDNA methylation as well as combination and application of markers
CN118064583A