Multi-cancer detection and tracing system based on high-throughput methylation sequencing of DNA

CN118298916BActive Publication Date: 2026-09-22ZHEJIANG SHENGTING MEDICAL LAB CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410386951.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-01
Publication Date
2026-09-22
Estimated Expiration
2044-04-01

AI Technical Summary

Technical Problem

[0003]目前,我国在临床中能够进行筛查的癌种主要包括肺癌、肠癌、乳腺癌、宫颈癌、胃癌、食管癌和肝癌等,而对于大多数其他癌种,目前仍缺乏可行的早筛方法

Benefits of technology

[0034]本发明提供了一种基于DNA高通量甲基化测序的多癌种检测及溯源系统,属于分子生物医学技术领域,包括:第一样本测序模块,用于对样本DNA进行测序,甲基化特征集合获取模块用于利用比较测序数据的甲基化水平,得到甲基化特征集合;末端基序特征获取模块,用于统计末端基序碱基片段在全部排列组合中所占比,得到末端基序特征;断点基序特征获取模块,用于统计断点基序碱基片段在全部排列组合中所占比,得到断点基序特征;模型构建模块,用于将甲基化特征集合、末端基序特征和断点基序特征作为输入,将是否患癌、患癌概率作为输出,对学习分类模块进行训练,得到检测及溯源模型。本发明具有无创检测,测序成本降低,检测特异性和敏感性高等特点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118298916B_ABST
    Figure CN118298916B_ABST
Patent Text Reader

Abstract

The application provides a multi-cancer detection and tracing system based on DNA high-throughput methylation sequencing, and belongs to the technical field of molecular biology and medicine, and comprises: a first sample sequencing module for sequencing sample DNA, a methylation feature set acquisition module for obtaining a methylation feature set by comparing the methylation level of sequencing data, an end motif feature acquisition module for counting the proportion of end motif base fragments in all permutations and combinations to obtain end motif features, a breakpoint motif feature acquisition module for counting the proportion of breakpoint motif base fragments in all permutations and combinations to obtain breakpoint motif features, and a model construction module for taking the methylation feature set, the end motif features and the breakpoint motif features as inputs, taking whether to have cancer and the probability of having cancer as outputs, training a learning classification module, and obtaining a detection and tracing model. The application has the characteristics of non-invasive detection, reduced sequencing cost, high detection specificity and sensitivity, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular biomedical technology, and in particular to a multi-cancer detection and tracing system based on high-throughput DNA methylation sequencing. Background Technology

[0002] Early diagnosis and treatment of malignant tumors are urgent issues that my country needs to address. Malignant tumors are a serious threat to human health. Currently, the average 5-year survival rate for malignant tumor patients in my country is about 35%, although this has improved in recent years. The treatment cost for early-stage tumors is significantly lower than that for late-stage tumors. Therefore, early screening and diagnosis of tumors can not only significantly improve patient survival rates but also greatly reduce the medical expenditure burden on the country and individuals. In particular, pan-cancer early screening technology is expected to significantly reduce the cost of tumor screening and greatly improve the early diagnosis rate of tumors.

[0003] Currently, the types of cancer that can be screened in clinical practice in my country mainly include lung cancer, colorectal cancer, breast cancer, cervical cancer, stomach cancer, esophageal cancer, and liver cancer. However, for most other types of cancer, there is still a lack of feasible early screening methods. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a multi-cancer detection and tracing system based on high-throughput DNA methylation sequencing. This multi-cancer detection and tracing system based on high-throughput DNA methylation sequencing features non-invasive detection, reduced sequencing costs, and high detection specificity and sensitivity.

[0005] To achieve the above objectives, the present invention provides the following solution:

[0006] A multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA, characterized in that it comprises:

[0007] The first sample sequencing module is used to collect cancer tissue and adjacent tissue samples of the target cancer and plasma samples of the target cancer patient and healthy person, extract the first genomic DNA of the cancer tissue and adjacent tissue samples of the target cancer and the first cell-free DNA of the plasma samples of the target cancer patient and healthy person, and perform methylation sequencing on the first genomic DNA and the first cell-free DNA respectively to obtain the first genomic DNA sequencing data and the first cell-free DNA sequencing data.

[0008] The first methylation feature set acquisition module is used to, after quality control of the first genomic DNA sequencing data and the first cell-free DNA sequencing data, respectively align the first genomic DNA sequencing data and the first cell-free DNA sequencing data to the human reference genome, obtain the first coordinates of the first genomic DNA sequencing data and the first cell-free DNA sequencing data on the human reference genome, divide the human reference genome into multiple CpG Island regions according to the first coordinates, and count the methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different preset haploid algorithms, compare the differences in the methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different haploid algorithms, obtain the cancer-specific regions of the first genomic DNA sequencing data with respect to each haploid algorithm, count the methylation level of the first cell-free DNA sequencing data in the cancer-specific regions with respect to the preset haploid algorithm, and obtain the methylation feature set of the first cell-free DNA.

[0009] The first terminal motif feature acquisition module is used to take all p base pairs of the 3' end of the first free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the first free DNA.

[0010] The first breakpoint motif feature acquisition module is used to take all q base pairs upstream and downstream of the 3' end of the first free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the first free DNA.

[0011] The detection model construction module is used to train the learning classification module by taking the methylation feature set of the first free DNA, the terminal motif feature of the first free DNA, and the breakpoint motif feature of the first free DNA as input vectors and whether or not cancer is present as output vectors, to obtain a trained multi-cancer detection model; the multi-cancer detection model is used to detect whether the target object has the target cancer.

[0012] Preferably, the target cancers include: lung cancer, stomach cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer.

[0013] Preferably, the preset haploid algorithm includes: MM, MHL, CHALM, PDR, and Entropy.

[0014] Preferably, the first sample sequencing module includes:

[0015] The first sample collection unit is used to collect 10ml whole blood samples from both the target cancer patient and a healthy person using a cell-free DNA blood collection tube.

[0016] The cell-free DNA extraction unit is used to extract the first cell-free DNA from the whole blood sample plasma using the GENFINE plasma DNA extraction kit;

[0017] The free DNA sequencing unit is used to construct a library and perform degenerate representative bisulfite sequencing on the first free DNA to obtain the sequencing data of the first free DNA.

[0018] Preferably, the first sample sequencing module further includes:

[0019] The second sample collection unit is used to collect cancer tissue and adjacent normal tissue samples of the target cancer.

[0020] The genomic DNA extraction unit is used to extract the first genomic DNA from cancer tissue and adjacent normal tissue samples of the target cancer using the TIANGEN genomic DNA extraction kit.

[0021] A genomic DNA sequencing unit is used to construct a library and perform degenerate representative bisulfite sequencing on the first genomic DNA to obtain sequencing data of the first genomic DNA.

[0022] Preferably, the first methylation feature set acquisition module includes:

[0023] The control group setting unit is used to set up samples of each type of cancer in the target cancer as positive control groups and samples of adjacent normal tissue corresponding to the target cancer as negative control groups.

[0024] The control group comparison unit is used to compare the methylation levels of the positive control group and the negative control group with respect to the haploid algorithm with p-values ​​after multiple test correction, to obtain the cancer-specific regions of all target cancer samples.

[0025] Preferably, the learning classification module is divided into a first sub-module and a second sub-module connected to the first sub-module. The first sub-module embeds an ensemble algorithm. The ensemble algorithm includes: logistic regression model algorithm, support vector machine algorithm, random forest algorithm, gradient boosting tree algorithm, Bayesian model algorithm, k-nearest neighbor algorithm, XGBoost algorithm, and CatBoost algorithm. The second sub-module has a built-in logistic regression model. The ensemble algorithm is used to train the free DNA methylation feature set, the free DNA terminal motif features, and the free DNA breakpoint motif features. The logistic regression model is used to integrate and output the training results of the ensemble algorithm.

[0026] Preferably, p is any integer between 4 and 10, and q is any integer between 2 and 5.

[0027] Preferably, a multi-cancer tracing system based on high-throughput methylation sequencing of cell-free DNA is characterized by comprising:

[0028] The second sample sequencing module is used to collect cancer tissue and adjacent tissue samples of the target cancer and plasma samples of the target cancer patient, extract the second genomic DNA of the cancer tissue and adjacent tissue samples of the target cancer and the second cell-free DNA of the plasma sample of the target cancer patient, and perform methylation sequencing on the second genomic DNA and the second cell-free DNA respectively to obtain the second genomic DNA sequencing data and the second cell-free DNA sequencing data.

[0029] The second methylation feature set acquisition module is used to, after quality control of the second genomic DNA sequencing data and the second cell-free DNA sequencing data, align the second genomic DNA sequencing data and the second cell-free DNA sequencing data to the human reference genome, obtain the second coordinates of the second genomic DNA sequencing data and the second cell-free DNA sequencing data on the human reference genome, divide the human reference genome into multiple CpG Island regions according to the second coordinates, and count the methylation level of the second genomic DNA sequencing data in the CpG Island regions with respect to different haplotype algorithms, compare the differences in the methylation level of the second genomic DNA sequencing data in the CpG Island regions with respect to the preset haplotype algorithm, obtain the tissue-specific regions of the second genomic DNA sequencing data with respect to each haplotype algorithm, count the methylation level of the second cell-free DNA sequencing data in the tissue-specific regions with respect to the preset haplotype algorithm, and obtain the methylation feature set of the second cell-free DNA.

[0030] The second terminal motif feature acquisition module is used to take all p base pairs of the 3' end of the second cell-free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the second cell-free DNA.

[0031] The second breakpoint motif feature acquisition module is used to take all q base pairs upstream and downstream of the 3' end of the second cell-free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the second cell-free DNA.

[0032] The source tracing model construction module is used to train the learning classification module by taking the methylation feature set of the second free DNA, the terminal motif feature of the second free DNA, and the breakpoint motif feature of the second free DNA as input vectors and the probability of cancer as output vectors, to obtain a trained multi-cancer source tracing model; the multi-cancer source tracing model is used to determine the type of cancer detected.

[0033] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0034] This invention provides a multi-cancer detection and source tracing system based on high-throughput DNA methylation sequencing, belonging to the field of molecular biomedical technology. It includes: a first sample sequencing module for sequencing sample DNA; a methylation feature set acquisition module for obtaining a methylation feature set by comparing the methylation levels of the sequencing data; a terminal motif feature acquisition module for statistically analyzing the proportion of terminal motif base fragments in all permutations and combinations to obtain terminal motif features; a breakpoint motif feature acquisition module for statistically analyzing the proportion of breakpoint motif base fragments in all permutations and combinations to obtain breakpoint motif features; and a model building module for training a learning classification module using the methylation feature set, terminal motif features, and breakpoint motif features as inputs, and the presence or absence of cancer and the probability of cancer as outputs, to obtain a detection and source tracing model. This invention features non-invasive detection, reduced sequencing costs, and high detection specificity and sensitivity. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 The detection model construction process provided in this embodiment of the invention;

[0037] Figure 2 The traceability model construction process provided in this embodiment of the invention;

[0038] Figure 3 This is a quality control chart provided in an embodiment of the present invention. Figure 3 (a) is a CGG / TGG ratio chart. Figure 3 (b) is the comparison rate. Figure 3 (c) is a graph showing the number of reads aligned to the reference genome. Figure 3 (d) is a map showing the number of CGI islands covering 10X and above. Figure 3 (e) is a map showing the number of CpG sites covering 10X or more;

[0039] Figure 4 The specific construction process of the learning model provided in the embodiments of the present invention;

[0040] Figure 5 AUC performance diagram provided for embodiments of the present invention;

[0041] Figure 6 This is a sensitivity diagram provided by an embodiment of the present invention under different stages;

[0042] Figure 7 This is a sensitivity diagram provided in an embodiment of the present invention. Figure 7 (a) is a graph showing the sensitivity of lung cancer at different stages. Figure 7 (b) is a graph showing the sensitivity of the model at different stages of colorectal cancer. Figure 7 (c) is a graph showing the sensitivity of the model at different stages of gastric cancer. Figure 7 (d) shows the sensitivity of the model at different stages of liver cancer. Figure 7 (e) shows the sensitivity of the model at different stages of thyroid cancer. Figure 7 (f) shows the sensitivity of the model at different stages of breast cancer;

[0043] Figure 8 This is a graph showing the performance of cancer scores on the model training set provided in this embodiment of the invention.

[0044] Figure 9 The confusion matrix of the model provided in this embodiment of the invention on the validation set;

[0045] Figure 10 A graph showing the sensitivity of the model to different cancers provided in the embodiments of the present invention;

[0046] Figure 11This is a graph showing the effect of the amount of free DNA injected on the model accuracy provided in an embodiment of the present invention;

[0047] Figure 12 This is a graph illustrating the effect of age on model accuracy provided in an embodiment of the present invention.

[0048] Figure 13 The graph shows the effect of gender on model accuracy, as provided in an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] The purpose of this invention is to provide a multi-cancer detection and tracing system based on high-throughput DNA methylation sequencing. This system features non-invasive detection, reduced sequencing costs, and high detection specificity and sensitivity.

[0051] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Figure 1 For the detection model building process, such as Figure 1 As shown, this invention provides a multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA, comprising:

[0053] The first sample sequencing module is used to collect cancer tissue and adjacent tissue samples of the target cancer and plasma samples of the target cancer patient and healthy person, extract the first genomic DNA of the cancer tissue and adjacent tissue samples of the target cancer and the first cell-free DNA of the plasma samples of the target cancer patient and healthy person, and perform methylation sequencing on the first genomic DNA and the first cell-free DNA respectively to obtain the first genomic DNA sequencing data and the first cell-free DNA sequencing data.

[0054] The first methylation feature set acquisition module is used to, after quality control of the first genomic DNA sequencing data and the first cell-free DNA sequencing data, respectively align the first genomic DNA sequencing data and the first cell-free DNA sequencing data to the human reference genome, obtain the first coordinates of the first genomic DNA sequencing data and the first cell-free DNA sequencing data on the human reference genome, divide the human reference genome into multiple CpG Island regions according to the first coordinates, and count the methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different preset haploid algorithms, compare the differences in the methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different haploid algorithms, obtain the cancer-specific regions of the first genomic DNA sequencing data with respect to each haploid algorithm, count the methylation level of the first cell-free DNA sequencing data in the cancer-specific regions with respect to the preset haploid algorithm, and obtain the methylation feature set of the first cell-free DNA.

[0055] The first terminal motif feature acquisition module is used to take all p base pairs of the 3' end of the first free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the first free DNA.

[0056] The first breakpoint motif feature acquisition module is used to take all q base pairs upstream and downstream of the 3' end of the first free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the first free DNA.

[0057] The detection model construction module is used to train the learning classification module by taking the methylation feature set of the first free DNA, the terminal motif feature of the first free DNA, and the breakpoint motif feature of the first free DNA as input vectors and whether or not cancer is present as output vectors, to obtain a trained multi-cancer detection model; the multi-cancer detection model is used to detect whether the target object has the target cancer.

[0058] Optionally, a computer-readable medium is provided that can run a computer program used by a multi-cancer detection system based on high-throughput methylation sequencing of free DNA.

[0059] Specifically, the target cancers include lung cancer, stomach cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer.

[0060] Optionally, haploids include MM, MHL, CHALM, PDR, and Entropy.

[0061] Furthermore, the first sample sequencing module includes:

[0062] The first sample collection unit is used to collect 10ml whole blood samples from both the target cancer patient and a healthy person using a cell-free DNA blood collection tube.

[0063] The cell-free DNA extraction unit is used to extract the first cell-free DNA from the whole blood sample plasma using the GENFINE plasma DNA extraction kit;

[0064] The free DNA sequencing unit is used to construct a library and perform degenerate representative bisulfite sequencing on the first free DNA to obtain the sequencing data of the first free DNA.

[0065] Furthermore, the first sample sequencing module also includes:

[0066] The second sample collection unit is used to collect cancer tissue and adjacent normal tissue samples of the target cancer.

[0067] The genomic DNA extraction unit is used to extract the first genomic DNA from cancer tissue and adjacent normal tissue samples of the target cancer using the TIANGEN genomic DNA extraction kit.

[0068] A genomic DNA sequencing unit is used to construct a library and perform degenerate representative bisulfite sequencing on the first genomic DNA to obtain sequencing data of the first genomic DNA.

[0069] Specifically, the first methylation feature set acquisition module includes:

[0070] The control group setting unit is used to set up samples of each type of cancer in the target cancer as positive control groups and samples of adjacent normal tissue corresponding to the target cancer as negative control groups.

[0071] The control group comparison unit is used to compare the methylation levels of the positive control group and the negative control group with respect to the haploid algorithm with p-values ​​after multiple test correction, to obtain the cancer-specific regions of all target cancer samples.

[0072] Specifically, the learning classification module is divided into a first sub-module and a second sub-module connected to the first sub-module. The first sub-module embeds an ensemble algorithm, which includes: logistic regression model algorithm, support vector machine algorithm, random forest algorithm, gradient boosting tree algorithm, Bayesian model algorithm, k-nearest neighbor algorithm, XGBoost algorithm, and CatBoost algorithm. The second sub-module has a built-in logistic regression model. The ensemble algorithm is used to train the free DNA methylation feature set, the free DNA terminal motif features, and the free DNA breakpoint motif features. The logistic regression model is used to integrate and output the training results of the ensemble algorithm.

[0073] Optionally, p is any integer between 4 and 10, and q is any integer between 2 and 5.

[0074] Figure 2 For the process of building a traceability model, such as Figure 2 As shown, this invention provides a multi-cancer tracing system based on high-throughput methylation sequencing of cell-free DNA, comprising:

[0075] The second sample sequencing module is used to collect cancer tissue and adjacent tissue samples of the target cancer and plasma samples of the target cancer patient, extract the second genomic DNA of the cancer tissue and adjacent tissue samples of the target cancer and the second cell-free DNA of the plasma sample of the target cancer patient, and perform methylation sequencing on the second genomic DNA and the second cell-free DNA respectively to obtain the second genomic DNA sequencing data and the second cell-free DNA sequencing data.

[0076] The second methylation feature set acquisition module is used to, after quality control of the second genomic DNA sequencing data and the second cell-free DNA sequencing data, align the second genomic DNA sequencing data and the second cell-free DNA sequencing data to the human reference genome, obtain the second coordinates of the second genomic DNA sequencing data and the second cell-free DNA sequencing data on the human reference genome, divide the human reference genome into multiple CpG Island regions according to the second coordinates, and count the methylation level of the second genomic DNA sequencing data in the CpG Island regions with respect to different haplotype algorithms, compare the differences in the methylation level of the second genomic DNA sequencing data in the CpG Island regions with respect to the preset haplotype algorithm, obtain the tissue-specific regions of the second genomic DNA sequencing data with respect to each haplotype algorithm, count the methylation level of the second cell-free DNA sequencing data in the tissue-specific regions with respect to the preset haplotype algorithm, and obtain the methylation feature set of the second cell-free DNA.

[0077] The second terminal motif feature acquisition module is used to take all p base pairs of the 3' end of the second cell-free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the second cell-free DNA.

[0078] The second breakpoint motif feature acquisition module is used to take all q base pairs upstream and downstream of the 3' end of the second cell-free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the second cell-free DNA.

[0079] The source tracing model construction module is used to train the learning classification module by taking the methylation feature set of the second free DNA, the terminal motif feature of the second free DNA, and the breakpoint motif feature of the second free DNA as input vectors and the probability of cancer as output vectors, to obtain a trained multi-cancer source tracing model; the multi-cancer source tracing model is used to determine the type of cancer detected.

[0080] Specifically, this embodiment also provides a cancer detection method for distinguishing whether a sample has lung cancer, stomach cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer, including the following steps:

[0081] Collect cancer tissue and adjacent normal tissue samples of the target cancer, extract genomic DNA from the cancer tissue and adjacent normal tissue samples of the target cancer, and perform methylation sequencing on the genomic DNA to obtain sequencing data of the cancer tissue and adjacent normal tissue samples.

[0082] After quality control of the sequencing data of the cancer tissue and adjacent normal tissue samples, the sequencing data of the cancer tissue and adjacent normal tissue samples are aligned to the human reference genome to obtain the coordinates of the sequencing data of the cancer tissue and adjacent normal tissue samples on the human reference genome. Based on the coordinates of the sequencing data of the cancer tissue and adjacent normal tissue samples on the human reference genome, the human reference genome is divided into multiple CpG Island regions, and the methylation level of the sequencing data of the cancer tissue and adjacent normal tissue samples on the CpG Island regions with respect to different haplotype algorithms is statistically analyzed.

[0083] By comparing the methylation level differences in the CpG Island region of the sequencing data of the cancer tissue and adjacent normal tissue samples with respect to each haplotype algorithm, the cancer-specific regions of the sequencing data of the cancer tissue and adjacent normal tissue samples with respect to each haplotype algorithm are obtained.

[0084] Plasma samples were collected from the target cancer patients and healthy individuals. Cell-free DNA was extracted from the plasma samples and methylated and sequenced to obtain sequencing data of the plasma samples from the target cancer patients and healthy individuals.

[0085] After quality control of the sequencing data of plasma samples from the target cancer patients and healthy individuals, the plasma samples from the target cancer patients and healthy individuals are aligned to the human reference genome to obtain the coordinates of the sequencing data of the plasma samples from the target cancer patients and healthy individuals on the reference genome. Based on the coordinates of the sequencing data of the plasma samples from the target cancer patients and healthy individuals on the reference genome, the methylation level of the sequencing data of the target cancer patients and healthy individuals in the cancer-specific region with respect to the haploid algorithm is calculated to obtain the methylation feature set of the target cancer patients and healthy individuals.

[0086] The entire 3' end of the p base pairs of sequencing data from plasma samples of the target cancer patients and healthy individuals is taken as the terminal motif fragment set on the reference genome, and the proportion of each base fragment in the terminal motif fragment set in the permutation and combination of the entire p base pairs sequence is taken as the terminal motif feature of the target cancer patients and healthy individuals.

[0087] The sequencing data of all 3' ends of plasma samples from the target cancer patients and healthy individuals, including q base pairs upstream and downstream of the reference genome, are used as a set of breakpoint motif base fragments. The proportion of each base fragment in the set of breakpoint motif base fragments in the permutation and combination of all 2*q base pairs of the sequence is used as the breakpoint motif feature of the target cancer patients and healthy individuals.

[0088] The methylation feature set of the target cancer patient and healthy person, the terminal motif feature of the target cancer patient and healthy person, and the breakpoint motif feature of the target cancer patient and healthy person are used as input vectors, and whether or not cancer is present is used as output vector. The learning classification module is trained to obtain a trained multi-cancer detection model. The multi-cancer detection model is used to detect lung cancer, gastric cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer.

[0089] Furthermore, this embodiment also provides a cancer detection device for distinguishing whether a sample has lung cancer, stomach cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer, comprising:

[0090] The first sample sequencing sub-device includes: a first sample collection unit, a first DNA extraction unit connected to the first sample collection unit, and a first DNA sequencing unit connected to the first DNA extraction unit; the first sample collection unit is used to collect cancer tissue and adjacent normal tissue samples of the target cancer, and plasma samples of the target cancer patient and healthy person; the first DNA extraction unit is used to extract genomic DNA from the cancer tissue and adjacent normal tissue samples of the target cancer, and cell-free DNA from the plasma samples of the target cancer patient and healthy person; the first DNA sequencing unit is used to perform methylation sequencing on the genomic DNA and the cell-free DNA, respectively, to obtain genomic DNA sequencing data and cell-free DNA sequencing data.

[0091] The first methylation feature set acquisition sub-device includes: a first coordinate positioning unit, a first methylation level comparison unit connected to the first coordinate positioning unit, and a first methylation level statistics unit connected to the first methylation level comparison unit; the first coordinate positioning unit is used to, after quality control of the first genomic DNA sequencing data and the first cell-free DNA sequencing data, respectively align the first genomic DNA sequencing data and the first cell-free DNA sequencing data to the human reference genome to obtain the first coordinates of the first genomic DNA sequencing data and the first cell-free DNA sequencing data on the human reference genome, and divide the human reference genome into multiple CpG Island regions according to the first coordinates; the first methylation level comparison unit is used to statistically analyze the methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different preset haplotype algorithms, and compare the differences in methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different haplotype algorithms to obtain the cancer-specific regions of the first genomic DNA sequencing data with respect to each haplotype algorithm; the first methylation level statistics unit is used to statistically analyze the methylation level of the first cell-free DNA sequencing data in the cancer-specific regions with respect to the preset haplotype algorithm to obtain the methylation feature set of the first cell-free DNA;

[0092] The first terminal motif feature acquisition sub-device is used to take all p base pairs of the 3' end of the first free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the first free DNA.

[0093] The first breakpoint motif feature acquisition sub-device is used to take all q base pairs upstream and downstream of the 3' end of the first free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the first free DNA.

[0094] The detection model construction sub-device includes: a first detection model unit and a second detection model unit connected to the first detection model unit. The first detection model unit embeds an ensemble algorithm, which includes: a logistic regression algorithm, a support vector machine algorithm, a random forest algorithm, a gradient boosting tree algorithm, a Bayesian model algorithm, a k-nearest neighbor algorithm, an XGBoost algorithm, and a CatBoost algorithm. The second detection model unit has a built-in logistic regression model. The ensemble algorithm is used to train the free DNA methylation feature set, the free DNA terminal motif features, and the free DNA breakpoint motif features. The logistic regression model is used to integrate and output the training results of the ensemble algorithm.

[0095] Specifically, this embodiment also provides a cancer tracing method for distinguishing whether a sample has lung cancer, stomach cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer, including the following steps:

[0096] Collect cancer tissue and adjacent normal tissue samples of the target cancer, extract genomic DNA from the cancer tissue and adjacent normal tissue samples of the target cancer, and perform methylation sequencing on the genomic DNA to obtain sequencing data of the cancer tissue and adjacent normal tissue samples.

[0097] After quality control of the sequencing data of the cancer tissue and adjacent normal tissue samples, the sequencing data of the cancer tissue and adjacent normal tissue samples are aligned to the human reference genome to obtain the coordinates of the sequencing data of the cancer tissue and adjacent normal tissue samples on the human reference genome. Based on the coordinates of the sequencing data of the cancer tissue and adjacent normal tissue samples on the human reference genome, the human reference genome is divided into multiple CpG Island regions, and the methylation level of the sequencing data of the cancer tissue and adjacent normal tissue samples on the CpG Island regions with respect to different haplotype algorithms is statistically analyzed.

[0098] By comparing the methylation level differences in the CpG Island region of the sequencing data of the cancer tissue and adjacent normal tissue samples with respect to each haplotype algorithm, tissue-specific regions of the sequencing data of the cancer tissue and adjacent normal tissue samples with respect to each haplotype algorithm are obtained.

[0099] Collect plasma samples from the target cancer patient, extract cell-free DNA from the plasma samples of the target cancer patient, and perform methylation sequencing on the cell-free DNA to obtain sequencing data of the plasma samples of the target cancer patient;

[0100] After quality control of the sequencing data of the plasma sample from the target cancer patient, the plasma sample from the target cancer patient is aligned to the human reference genome to obtain the coordinates of the sequencing data of the plasma sample from the target cancer patient on the reference genome. Based on the coordinates of the sequencing data of the plasma sample from the target cancer patient on the reference genome, the methylation level of the sequencing data of the target cancer patient on the tissue-specific region with respect to the haploid algorithm is calculated to obtain the methylation feature set of the target cancer patient.

[0101] The entire 3' end of the plasma sample sequencing data of the target cancer patient is divided into p base pairs on the reference genome as a set of terminal motif base fragments, and the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence is used as the terminal motif feature of the target cancer patient.

[0102] The sequencing data of the plasma sample of the target cancer patient are used as a set of q base pairs upstream and downstream of the 3' end on the reference genome. The proportion of each base fragment in the set of q base pairs in the permutation and combination of the 2*q base pairs is used as the breakpoint motif feature of the target cancer patient.

[0103] The methylation feature set, terminal motif, and breakpoint motif of the target cancer patient are used as input vectors, and the probability of cancer is used as the output vector to train the learning classification module, thereby obtaining a trained multi-cancer tracing model; the multi-cancer tracing model is used to determine the type of detected cancer.

[0104] Furthermore, this embodiment also provides a cancer tracing device for distinguishing whether a sample has lung cancer, stomach cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer, comprising:

[0105] The second sample sequencing sub-device includes: a second sample collection unit, a second DNA extraction unit connected to the second sample collection unit, and a second DNA sequencing unit connected to the second DNA extraction unit; the second sample collection unit is used to collect cancer tissue and adjacent normal tissue samples of the target cancer, and plasma samples of the target cancer patient and healthy person; the second DNA extraction unit is used to extract genomic DNA from the cancer tissue and adjacent normal tissue samples of the target cancer, and cell-free DNA from the plasma samples of the target cancer patient and healthy person; the second DNA sequencing unit is used to perform methylation sequencing on the genomic DNA and the cell-free DNA, respectively, to obtain genomic DNA sequencing data and cell-free DNA sequencing data.

[0106] The second methylation feature set acquisition sub-device includes: a second coordinate positioning unit, a second methylation level comparison unit connected to the second coordinate positioning unit, and a second methylation level statistics unit connected to the second methylation level comparison unit; the second coordinate positioning unit is used to, after quality control of the second genomic DNA sequencing data and the second cell-free DNA sequencing data, align the second genomic DNA sequencing data and the second cell-free DNA sequencing data to the human reference genome to obtain the second coordinates of the second genomic DNA sequencing data and the second cell-free DNA sequencing data on the human reference genome, and divide the human reference genome into multiple CpG island regions according to the second coordinates; the second methylation level comparison unit is used to statistically analyze the methylation level of the second genomic DNA sequencing data in the CpG island regions with respect to different preset haploid algorithms, and compare the differences in methylation level of the second genomic DNA sequencing data in the CpG island regions with respect to different haploid algorithms to obtain the tissue-specific regions of the second genomic DNA sequencing data with respect to each haploid algorithm; the second methylation level statistics unit is used to statistically analyze the methylation level of the second cell-free DNA sequencing data in the tissue-specific regions with respect to the preset haploid algorithm to obtain the methylation feature set of the second cell-free DNA;

[0107] The second terminal motif feature acquisition sub-device is used to take all p base pairs of the 3' end of the second cell-free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and to take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the second cell-free DNA.

[0108] The second breakpoint motif feature acquisition sub-device is used to take all q base pairs upstream and downstream of the 3' end of the second cell-free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the second cell-free DNA.

[0109] The source tracing model construction sub-device includes: a first source tracing model unit and a second source tracing model unit connected to the first source tracing model unit. The first source tracing model unit embeds an ensemble algorithm, which includes: a logistic regression model algorithm, a support vector machine algorithm, a random forest algorithm, a gradient boosting tree algorithm, a Bayesian model algorithm, a k-nearest neighbor algorithm, an XGBoost algorithm, and a CatBoost algorithm. The second source tracing model unit has a built-in logistic regression model. The ensemble algorithm is used to train the free DNA methylation feature set, the free DNA terminal motif features, and the free DNA breakpoint motif features. The logistic regression model is used to integrate and output the training results of the ensemble algorithm.

[0110] refer to Figure 3 The quality control requirements are as follows: the CGG / TGG ratio is greater than 0.6, the alignment rate is greater than 0.4, the number of readings aligned to the reference genome is greater than 10 million, the number of CGI islands covering 10X or higher is greater than 10,000, and the number of CpG sites covering 10X or higher is greater than 1 million.

[0111] Specifically, cell-free DNA is also known as cfDNA, genomic DNA is also known as gDNA, 6-mer end motif is also known as 6-mer breakpoint motif, logistic regression is also known as Logistic Regression (LR), support vector machines are also known as SVM, random forest is also known as Random Forest (RF), gradient boosting machine is also known as Gradient Boosting Machine (GBM), Naive Bayes is also known as Naive Bayes (NB), k-Nearest Neighbor (KNN), gradient bias is also known as radient bias, prediction shift is also known as prediction shift, and meta-learning is also known as meta-learning.

[0112] Furthermore, the cell-free DNA sample library construction adopts a patented technology: a method for rapidly constructing RRBS sequencing libraries using circulating tumor DNA, patent number: ZL 202111060927.0.

[0113] Specifically, each haplotype algorithm calculates the methylation level within a region. Different algorithms for different haplotypes result in different methylation levels from the same sequencing data. A machine learning model is built by analyzing the methylation levels of sequencing data under different haplotype algorithms. This model can be used to distinguish between different cancers. By comparing the haplotype methylation levels of cancerous tissue samples and adjacent normal tissue samples, sequencing data fragments exhibiting significant differences in regions of different CpG islands are identified, and these fragments can serve as distinguishing features.

[0114] Optionally, in tumor tissue, due to changes in chromatin state and abnormal nuclease activity, DNA fragments may be obtained that are cleaved at tumor-specific sites. Because the breakpoints are specific, the proportions of terminal motif sequences vary.

[0115] Furthermore, the calculation method for terminal motif features is as follows: Cell-free DNA high-throughput sequencing data is aligned to the human reference genome, and the 3' end position of each read is extracted, corresponding to a 6-base sequence on the reference genome. The orientation of the base sequence is from the 5' to the 3' end. Each base sequence can be A, T, C, or G, resulting in 4096 possible combinations of the 6 base sequences, thus providing a total of 4096 possible terminal motif percentages.

[0116] Furthermore, the breakpoint motif feature calculation method involves aligning cell-free DNA high-throughput sequencing data to a human reference genome and extracting three base sequences upstream and downstream of the 3' ends of each read from the reference genome. The base sequence orientation is from the 5' to the 3' end. Each base sequence can be A, T, C, or G, resulting in 4096 possible combinations of these six base sequences, thus representing 4096 breakpoint motif percentages.

[0117] Specifically, logistic regression, also known as logistic regression analysis, is a generalized logistic regression analysis model and belongs to supervised learning in machine learning. The logistic regression process is as follows: for a classification problem, a cost function is established, and the optimal model parameters are iteratively solved using optimization methods. Its advantages include speed, suitability for binary classification problems, and ease of understanding.

[0118] Specifically, the Support Vector Machine (SVM) model is a supervised learning algorithm used to predict discrete or continuous variables. It achieves classification or regression by mapping a dataset to a high-dimensional space and finding an optimal hyperplane that maximizes the margin between data points of different classes. During training, the parameters are set as follows: kernel="linear", scale=T, probability=TRUE, type="C-classification".

[0119] Specifically, the Random Forest model is a supervised learning algorithm used to predict discrete or continuous variables. It constructs multiple decision tree models by randomly selecting features and subsets of the dataset, and then averages or votes on their predictions to improve prediction accuracy. During training, the parameters are set as follows: method="rf", prox=TRUE, ntree=500, metric="Accuracy".

[0120] Specifically, the gradient boosting tree model is an ensemble learning algorithm that combines gradient boosting and decision tree techniques. In machine learning, gradient boosting tree models are commonly used to solve regression and classification problems. Its main advantages are its ability to handle mixed data types and its robustness to outliers and noisy data. During training, the parameters are set as follows: method = 'gbm', trControl = trainControl(allowParallel = TRUE, verboseIter = FALSE).

[0121] Specifically, a Bayesian model is a statistical model based on Bayes' theorem, used to handle problems involving uncertainty and probabilistic inference. Bayesian models use probability distributions to describe parameters or unknowns and update these probability distributions based on observed data, making them suitable for problems in various fields such as regression analysis, classification, and clustering. During training, the parameters are set as follows: laplace = 0, type = 'row'.

[0122] Specifically, XGBoost is an additive model based on boosting trees, a type of gradient boosting decision tree. Its basic idea is the same as gradient boosting decision trees, but it incorporates many optimizations. Each base classifier used is a regression tree, and the training method employs a forward stepwise algorithm to progressively optimize each base learner. During training, the parameters are set as follows: nrounds = 1000, params = list(booster = "gbtree", objective = "binary: logistic", eval_metric = "logloss").

[0123] Specifically, the k-nearest neighbors algorithm is a simple and effective supervised learning algorithm used for classification and regression problems. Its basic idea is that if a sample's k nearest neighbors in the feature space mostly belong to either a classification or regression problem, then the sample also belongs to that category or has that average. During training, the parameters are set as follows: k = seq(1, 30, 2), method = 'repeatedcv', number = 10.

[0124] Specifically, CatBoost is a gradient boosting decision tree framework based on symmetric decision trees as base learners. It features fewer parameters, supports categorical variables, and offers high accuracy. Composed of Categorical and Boosting components, it aims to address gradient bias and prediction offset issues, thereby reducing overfitting and improving the algorithm's accuracy and generalization ability. During training, the parameters are set as follows: params = list(loss_function = 'Logloss', iterations = 1000).

[0125] Furthermore, all plasma samples were randomly divided into a training set and a validation set in a 1:1 ratio. The training set was used to train the model, and the validation set was used to evaluate the model.

[0126] Optionally, the features of the multi-cancer early screening model include: MM, MHL, CHALM, PDR, Entropy, terminal motif features, and breakpoint motif features. The samples in the training set are fed into eight machine learning algorithm models for training, and the final output is the probability of developing cancer. Therefore, each feature is trained by eight machine learning models, for a total of 56 machine learning models; this is the training model for the first layer. The results of the 56 machine learning models trained in the first layer are then fed into the logistic regression model in the second layer for integration, and the final multi-cancer early screening model result is obtained using the probability of developing cancer. This method is called a stacked ensemble machine learning model. It uses meta-learning algorithms to learn how best to combine predictions from two or more basic machine learning algorithms. The advantage of this approach is that it can leverage the capabilities of a series of well-performing models on classification or regression tasks and make better predictions than any single model in the ensemble.

[0127] refer to Figure 4 To better train and evaluate the model, this invention collected a large number of plasma samples, totaling 2459. After quality control, 2445 samples were retained. These 2445 samples were randomly divided into a training set and a test set in a 1:1 ratio. The training set was used to train the model, while the test set was used to evaluate the model.

[0128] The results for the training and test sets are shown in Table 1:

[0129] The training set of 1224 cases was used to train the model, and the resulting model had a specificity of 95.18%, a sensitivity of 88.58%, and an accuracy of 90.93%.

[0130] The test set of 1221 cases was used to evaluate the model. The model's specificity was 96.33%, sensitivity was 91.21%, and accuracy was 93.04%.

[0131] The specificity, sensitivity, and accuracy of the training and test sets are almost identical, indicating that the model training is good and there is no obvious overfitting.

[0132] Table 1. Specificity, sensitivity, and accuracy of the training and test sets.

[0133]

[0134]

[0135] refer to Figure 5 The AUC on the test set was 0.983, indicating that this embodiment has good predictive performance.

[0136] refer to Figure 6 and Figure 7 The sensitivity of different cancer stages is shown in Table 2:

[0137] The model's sensitivity across different stages generally follows normal patterns, with lower sensitivity in stages I and II, and higher sensitivity in stages III and IV. Even in early-stage cancer (stage I), the model's sensitivity remains high, reaching 91% for colorectal cancer.

[0138] Table 2 Sensitivity at different cancer stages

[0139]

[0140]

[0141] Specifically, the accuracy rates for different types of cancer are shown in Table 3:

[0142] The model achieved high accuracy across different cancers with no significant bias. The highest accuracy was achieved for colorectal cancer, at 96.92% on the training set and 97.25% on the test set.

[0143] Table 3. Accuracy under different cancers

[0144]

[0145] Specifically, the organization's traceability situation is shown in Table 4:

[0146] The model achieved an organization tracing rate of 71.83% on the training set and 77.23% on the test set, demonstrating good ability to trace back to specific cancer types.

[0147] Table 4. Accuracy of Organizational Tracing

[0148]

[0149] refer to Figure 8 , Figure 9 and Figure 10 The embodiments of the present invention have good accuracy and sensitivity in predicting and tracing the target cancer.

[0150] refer to Figure 11 , Figure 12 and Figure 13 The accuracy of the present invention in predicting and tracing the target cancer is less affected by the amount of cell-free DNA input, age, and gender.

[0151] The beneficial effects of this invention are as follows:

[0152] This invention combines methylation and fragmentomics-related features to train an integrated machine learning classification model, which features non-invasive detection, reduced sequencing costs, and high detection specificity and sensitivity.

[0153] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0154] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA, characterized in that, include: The first sample sequencing module is used to collect cancer tissue and adjacent tissue samples of the target cancer and plasma samples of the target cancer patient and healthy person, extract the first genomic DNA of the cancer tissue and adjacent tissue samples of the target cancer and the first cell-free DNA of the plasma samples of the target cancer patient and healthy person, and perform methylation sequencing on the first genomic DNA and the first cell-free DNA respectively to obtain the first genomic DNA sequencing data and the first cell-free DNA sequencing data. The first methylation feature set acquisition module is used to, after quality control of the first genomic DNA sequencing data and the first cell-free DNA sequencing data, respectively align the first genomic DNA sequencing data and the first cell-free DNA sequencing data to the human reference genome, obtain the first coordinates of the first genomic DNA sequencing data and the first cell-free DNA sequencing data on the human reference genome, divide the human reference genome into multiple CpG Island regions according to the first coordinates, and count the methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different preset haploid algorithms, compare the differences in the methylation level of the first genomic DNA sequencing data in the CpG Island regions with respect to different haploid algorithms, obtain the cancer-specific regions of the first genomic DNA sequencing data with respect to each haploid algorithm, count the methylation level of the first cell-free DNA sequencing data in the cancer-specific regions with respect to the preset haploid algorithm, and obtain the methylation feature set of the first cell-free DNA. The first terminal motif feature acquisition module is used to take all p base pairs of the 3' end of the first free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the first free DNA. The first breakpoint motif feature acquisition module is used to take all q base pairs upstream and downstream of the 3' end of the first free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the first free DNA. The detection model construction module is used to train the learning classification module by taking the methylation feature set of the first free DNA, the terminal motif feature of the first free DNA, and the breakpoint motif feature of the first free DNA as input vectors and whether or not cancer is present as output vectors, to obtain a trained multi-cancer detection model; the multi-cancer detection model is used to detect whether the target object has the target cancer.

2. The multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA provided in claim 1, characterized in that, The target cancers include: lung cancer, stomach cancer, colorectal cancer, liver cancer, breast cancer, and thyroid cancer.

3. The multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA provided in claim 1, characterized in that, The preset haploid algorithms include: MM, MHL, CHALM, PDR, and Entropy.

4. The multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA provided in claim 1, characterized in that, The first sample sequencing module includes: The first sample collection unit is used to collect 10ml whole blood samples from both the target cancer patient and a healthy person using a cell-free DNA blood collection tube. The cell-free DNA extraction unit is used to extract the first cell-free DNA from the whole blood sample plasma using the GENFINE plasma DNA extraction kit; The free DNA sequencing unit is used to construct a library and perform degenerate representative bisulfite sequencing on the first free DNA to obtain the sequencing data of the first free DNA.

5. The multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA provided in claim 1, characterized in that, The first sample sequencing module also includes: The second sample collection unit is used to collect cancer tissue and adjacent normal tissue samples of the target cancer. The genomic DNA extraction unit is used to extract the first genomic DNA from cancer tissue and adjacent normal tissue samples of the target cancer using the TIANGEN genomic DNA extraction kit. A genomic DNA sequencing unit is used to construct a library and perform degenerate representative bisulfite sequencing on the first genomic DNA to obtain sequencing data of the first genomic DNA.

6. The multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA provided in claim 1, characterized in that, The first methylation feature set acquisition module includes: The control group setting unit is used to set up samples of each type of cancer in the target cancer as positive control groups and samples of adjacent normal tissue corresponding to the target cancer as negative control groups. The control group comparison unit is used to compare the methylation levels of the positive control group and the negative control group with respect to the haploid algorithm with p-values ​​after multiple test correction, to obtain the cancer-specific regions of all target cancer samples.

7. The multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA provided in claim 1, characterized in that, The learning and classification module is divided into a first sub-module and a second sub-module connected to the first sub-module. The first sub-module embeds an ensemble algorithm, which includes: logistic regression model algorithm, support vector machine algorithm, random forest algorithm, gradient boosting tree algorithm, Bayesian model algorithm, k-nearest neighbor algorithm, XGBoost algorithm, and CatBoost algorithm. The second sub-module has a built-in logistic regression model. The ensemble algorithm is used to train the free DNA methylation feature set, the free DNA terminal motif features, and the free DNA breakpoint motif features. The logistic regression model is used to integrate and output the training results of the ensemble algorithm.

8. The multi-cancer detection system based on high-throughput methylation sequencing of cell-free DNA provided in claim 1, characterized in that, p is any integer between 4 and 10, and q is any integer between 2 and 5.

9. A multi-cancer tracing system based on high-throughput methylation sequencing of cell-free DNA, characterized in that, include: The second sample sequencing module is used to collect cancer tissue and adjacent tissue samples of the target cancer and plasma samples of the target cancer patient, extract the second genomic DNA of the cancer tissue and adjacent tissue samples of the target cancer and the second cell-free DNA of the plasma sample of the target cancer patient, and perform methylation sequencing on the second genomic DNA and the second cell-free DNA respectively to obtain the second genomic DNA sequencing data and the second cell-free DNA sequencing data. The second methylation feature set acquisition module is used to, after quality control of the second genomic DNA sequencing data and the second cell-free DNA sequencing data, align the second genomic DNA sequencing data and the second cell-free DNA sequencing data to the human reference genome, obtain the second coordinates of the second genomic DNA sequencing data and the second cell-free DNA sequencing data on the human reference genome, divide the human reference genome into multiple CpG island regions according to the second coordinates, and count the methylation level of the second genomic DNA sequencing data in the CpG island regions with respect to different haplotype algorithms, compare the differences in the methylation level of the second genomic DNA sequencing data in the CpG island regions with respect to the preset haplotype algorithm, obtain the tissue-specific regions of the second genomic DNA sequencing data with respect to each haplotype algorithm, count the methylation level of the second cell-free DNA sequencing data in the tissue-specific regions with respect to the preset haplotype algorithm, and obtain the methylation feature set of the second cell-free DNA. The second terminal motif feature acquisition module is used to take all p base pairs of the 3' end of the second cell-free DNA sequencing data on the reference genome as a set of terminal motif base fragments, and take the proportion of each base fragment in the set of terminal motif base fragments in the permutation and combination of all p base pairs of the sequence as the terminal motif feature of the second cell-free DNA. The second breakpoint motif feature acquisition module is used to take all q base pairs upstream and downstream of the 3' end of the second cell-free DNA sequencing data on the reference genome as a breakpoint motif base fragment set, and take the proportion of each base fragment in the breakpoint motif base fragment set in the permutation and combination of all 2*q base pairs as the breakpoint motif feature of the second cell-free DNA. The source tracing model construction module is used to train the learning classification module by taking the methylation feature set of the second free DNA, the terminal motif feature of the second free DNA, and the breakpoint motif feature of the second free DNA as input vectors and the probability of cancer as output vectors, to obtain a trained multi-cancer source tracing model; the multi-cancer source tracing model is used to determine the type of cancer detected.

Citation Information

Patent Citations

  • A method for rapidly constructing RRBS sequencing libraries using circulating tumor DNA

    CN113604540B

  • Individualized management system and method for colorectal cancer risk prediction

    CN120126768A

  • Multi-cancer detection and traceability systems based on DNA high-throughput methylation sequencing

    US12698535B1