Methods of obtaining cancer risk prediction markers and methods of cancer risk assessment

By screening DNA sequences whose cfDNA ends are located within the nucleosome protection region of blood cells, calculating the differences in fragmentation characteristics, and combining machine learning, a highly accurate cancer risk assessment model was developed. This solves the problem of insufficient pan-cancer cfDNA tumor markers in existing technologies and achieves highly sensitive cancer risk assessment.

CN118692649BActive Publication Date: 2025-12-19OMIXSCIENCE (SHENZHEN) CO LTD

Patent Information

Application Number
CN202410378081.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-12-19
Estimated Expiration
2044-03-28

AI Technical Summary

Technical Problem

The lack of universally applicable peripheral blood circulating cell-free DNA (cfDNA) tumor markers in current technologies leads to insufficient accuracy and sensitivity in cancer diagnosis.

Method used

By screening DNA sequences whose cfDNA ends are located within the nucleosome protection region of blood cells, calculating the differences in fragmentation characteristics, and combining machine learning technology to develop a cancer risk assessment model, several novel cancer diagnostic biomarkers based on the distribution patterns of peripheral blood cfDNA ends were identified.

Benefits of technology

It improves the accuracy and sensitivity of cancer prediction, and achieves high-accuracy cancer risk assessment across a wide range of cancer types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118692649B_ABST
    Figure CN118692649B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for obtaining a cancer risk prediction marker, a cancer risk assessment method, an electronic device, and a storage medium performed by a machine. The method for obtaining a cancer risk prediction marker includes: analyzing a plurality of peripheral blood circulating cell-free DNA (cfDNA) fragments of a biological sample of a subject to obtain sequence reads; obtaining data of a distribution of a nucleosome, and determining a predetermined interval upstream and downstream of a nucleosome center position as a nucleosome protection region; screening sequence reads of partial cfDNA fragments whose ends are located in the nucleosome protection region from the sequence reads of the plurality of cfDNA fragments; calculating a first fragmentation feature in the sequence reads of the plurality of cfDNA fragments, and calculating a second fragmentation feature in the sequence reads of the screened partial cfDNA fragments; and calculating one or more difference values of a change level of the second fragmentation feature relative to the first fragmentation feature as a cancer risk prediction marker.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of cancer diagnosis, in particular to a pan-cancer blood diagnostic marker, a method for obtaining a cancer risk prediction marker performed by a machine, a cancer risk assessment method, an electronic device, and a storage medium. BACKGROUND

[0002] Peripheral blood circulating cell-free DNA (cfDNA) is a DNA fragment naturally generated after cell death. Studies have shown that in healthy people, cfDNA is mainly derived from blood cells and the liver. In cancer patients, cfDNA is also released into peripheral blood by dying tumor cells. cfDNA fragments from tumors and cfDNA fragments from non-tumor cells have many differences, such as gene mutations, chromosomal copy number variations, abnormal methylation levels, and so on. Therefore, cancer diagnosis can be performed by detecting cfDNA signals from tumors in peripheral blood, i.e. liquid biopsy technology. Since the collection of peripheral blood is almost non-invasive to the human body and can be repeated, cancer diagnosis based on peripheral blood cfDNA has high application value, and a large number of tumor cfDNA markers have been identified and a large number of cancer diagnosis technologies have been developed, such as high-frequency gene mutation kits, tumor-specific methylation marker kits, etc. In recent years, researchers have found that cfDNA is not randomly fragmented when it is generated, and there are differences between tumor cells and normal cells. Based on this principle, researchers have found some tumor markers related to cfDNA fragmentation patterns and developed multiple cancer diagnosis technologies. However, at present, there are still few verified pan-cancer and widely applicable cfDNA tumor markers. Therefore, a new method for obtaining a cancer risk prediction marker and a cancer risk assessment method are needed to solve the above technical problems. SUMMARY

[0003] In view of the above problems, the present disclosure provides a method for obtaining a cancer risk prediction marker performed by a machine, a cancer risk assessment method, an electronic device, and a storage medium. The present disclosure screens DNA sequences whose cfDNA end positions in the blood of a subject are located in the nucleosome protection region of the cell nucleus, and then quantitatively studies the differences in various types of cfDNA characteristics before and after sequence screening, to obtain multiple new cancer diagnostic markers based on the end distribution pattern of peripheral blood cfDNA. In addition, the present disclosure combines machine learning technology to develop a cancer risk assessment model with high accuracy and high sensitivity.

[0004] The present disclosure provides a method of obtaining a cancer risk prediction marker performed by a machine, comprising: analyzing a plurality of circulating cell-free DNA (cfDNA) fragments of a biological sample of a subject to obtain sequence reads, the sequence reads comprising end-site information indicative of a location of an end of each cfDNA fragment in the plurality of cfDNA fragments on a genome; obtaining data of a distribution of nucleosomes, and determining a predetermined interval upstream and downstream of a nucleosome center position as a nucleosome protection region; performing fragment screening on the plurality of cfDNA fragments so as to screen out sequence reads of partial cfDNA fragments whose ends are located in the nucleosome protection region from the sequence reads of the plurality of cfDNA fragments; calculating a certain fragmentation feature (such as length, end sequence distribution, etc.) in the sequence reads of the un-screened plurality of cfDNA fragments, and calculating a corresponding fragmentation feature in the sequence reads of the screened partial cfDNA fragments; and calculating one or more difference values indicative of a change level of the fragmentation feature of the screened partial cfDNA fragments relative to the fragmentation feature of the un-screened plurality of cfDNA fragments as the cancer risk prediction marker.

[0005] According to some embodiments of the present disclosure, the end-site information indicates a location of a first end or a second end of each cfDNA fragment in the plurality of cfDNA fragments on a genome. Performing fragment screening on the plurality of cfDNA fragments further comprises: performing fragment screening on the plurality of cfDNA fragments so as to screen out sequence reads of partial cfDNA fragments whose first end or second end is located in the nucleosome protection region from the sequence reads of the plurality of cfDNA fragments.

[0006] According to some embodiments of the present disclosure, the end-site information indicates a location of a first end and a second end of each cfDNA fragment in the plurality of cfDNA fragments on a genome. Performing fragment screening on the plurality of cfDNA fragments further comprises: performing fragment screening on the plurality of cfDNA fragments so as to screen out sequence reads of partial cfDNA fragments whose first end and second end are both located in the nucleosome protection region from the sequence reads of the plurality of cfDNA fragments.

[0007] According to some embodiments of the present disclosure, the fragmentation feature of a cfDNA fragment comprises at least one or more of: a first proportion of short fragments with a length less than a predetermined value, a second proportion of end sequence usage, a third proportion of breakpoint sequence usage, a first diversity score of end sequence, a second diversity score of breakpoint sequence, and other biological information present on the cfDNA fragment, including copy number variation, mutation information.

[0008] According to some embodiments of the present disclosure, calculating one or more difference values indicative of a level of change of a second fragmentation feature of the selected portion of cfDNA fragments relative to a first fragmentation feature of the plurality of cfDNA fragments as the cancer risk prediction marker comprises calculating a difference value obtained by subtracting the first proportion in the fragmentation feature of the plurality of cfDNA fragments from the first proportion in the fragmentation feature of the selected portion of cfDNA fragments as a first cancer risk prediction marker As.

[0009] According to some embodiments of the present disclosure, calculating one or more difference values indicative of a level of change of a second fragmentation feature of the selected portion of cfDNA fragments relative to a first fragmentation feature of the plurality of cfDNA fragments as the cancer risk prediction marker comprises calculating a difference value obtained by subtracting the second proportion in the fragmentation feature of the plurality of cfDNA fragments from the second proportion in the fragmentation feature of the selected portion of cfDNA fragments as a second cancer risk prediction marker Am.

[0010] According to some embodiments of the present disclosure, calculating one or more difference values indicative of a level of change of a second fragmentation feature of the selected portion of cfDNA fragments relative to a first fragmentation feature of the plurality of cfDNA fragments as the cancer risk prediction marker comprises calculating a difference value obtained by subtracting the third proportion in the fragmentation feature of the plurality of cfDNA fragments from the third proportion in the fragmentation feature of the selected portion of cfDNA fragments as a third cancer risk prediction marker Am2.

[0011] According to some embodiments of the present disclosure, calculating one or more difference values indicative of a level of change of a second fragmentation feature of the selected portion of cfDNA fragments relative to a first fragmentation feature of the plurality of cfDNA fragments as the cancer risk prediction marker comprises calculating a difference value obtained by subtracting the first diversity score in the fragmentation feature of the plurality of cfDNA fragments from the first diversity score in the fragmentation feature of the selected portion of cfDNA fragments as a fourth cancer risk prediction marker Am3.

[0012] According to some embodiments of the present disclosure, calculating one or more difference values indicative of a level of change of a second fragmentation feature of the selected portion of cfDNA fragments relative to a first fragmentation feature of the plurality of cfDNA fragments as the cancer risk prediction marker comprises calculating a difference value obtained by subtracting the second diversity score in the fragmentation feature of the plurality of cfDNA fragments from the second diversity score in the fragmentation feature of the selected portion of cfDNA fragments as a fifth cancer risk prediction marker Am4.

[0013] According to some embodiments of the present disclosure, the one or more cancer risk prediction markers are used to indicate a non-specific plurality of cancer types.

[0014] According to some embodiments of the present disclosure, the one or more cancer risk prediction markers are used to indicate a non-specific plurality of cancer types, individually or in combination.

[0015] According to another aspect of the present disclosure, there is provided a cancer risk assessment method performed by a machine, comprising: obtaining a first set of biological samples of normal subjects; calculating, for the first set of biological samples, a first set of cancer risk prediction markers associated with the first set of biological samples according to the method as previously described; obtaining a second set of biological samples of a patient to be tested; calculating, for the second set of biological samples, a second set of cancer risk prediction markers associated with the second set of biological samples according to the method as previously described; determining whether a change value of the second set of cancer risk prediction markers relative to the first set of cancer risk prediction markers exceeds a predetermined threshold; and in response to the change value of the second set of cancer risk prediction markers relative to the first set of cancer risk prediction markers exceeding the predetermined threshold, outputting an indication of a risk of the patient to be tested having cancer.

[0016] According to some embodiments of the present disclosure, determining whether the change value of the second set of cancer risk prediction markers relative to the first set of cancer risk prediction markers exceeds a predetermined threshold comprises: for each cancer risk prediction marker in the second set of cancer risk prediction markers, determining whether a change value relative to a corresponding cancer risk prediction marker in the first set of cancer risk prediction markers exceeds a predetermined threshold.

[0017] According to some embodiments of the present disclosure, the predetermined threshold corresponding to each cancer risk prediction marker is different.

[0018] According to some embodiments of the present disclosure, determining whether the change value of the second set of cancer risk prediction markers relative to the first set of cancer risk prediction markers exceeds a predetermined threshold comprises: integrating a plurality of cancer risk prediction markers in the second set of cancer risk prediction markers to obtain a single cancer risk prediction marker, and for the single cancer risk prediction marker obtained from the second set of cancer risk prediction markers, determining whether a change value relative to a corresponding cancer risk prediction marker obtained from the first set of cancer risk prediction markers exceeds a predetermined threshold.

[0019] According to some embodiments of the present disclosure, determining whether the change value of the second set of cancer risk prediction markers relative to the first set of cancer risk prediction markers exceeds a predetermined threshold comprises integrating a plurality of cancer risk prediction markers in the second set of cancer risk prediction markers with other biomarkers to obtain individual cancer risk prediction markers, and determining, for the individual cancer risk prediction markers obtained from the second set of cancer risk prediction markers, whether the change value relative to corresponding cancer risk prediction markers obtained from the first set of cancer risk prediction markers exceeds a predetermined threshold.

[0020] According to some embodiments of the present disclosure, a cancer risk assessment model is generated by machine learning, which integrates a plurality of cancer risk prediction markers in the second set of cancer risk prediction markers to obtain individual cancer risk prediction markers, or integrates a plurality of cancer risk prediction markers in the second set of cancer risk prediction markers with other biomarkers to obtain individual cancer risk prediction markers by the cancer risk assessment model.

[0021] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory, wherein the memory has stored therein computer readable code which, when executed by the processor, implements the method described above.

[0022] According to another aspect of the present disclosure, a non-transitory computer readable storage medium is provided, storing computer readable instructions, wherein the computer readable instructions, when executed by a processor, implement the method described above.

[0023] Therefore, the method for obtaining cancer risk prediction markers, the cancer risk assessment method, the electronic device and the medium executed by a machine according to embodiments of the present disclosure can enrich tumor-derived cfDNA by screening free DNA inconsistent with blood cell nuclei in cancer patients, thereby improving the accuracy and sensitivity of tumor prediction. In addition, by screening the ends of free DNA, and then calculating the difference in cfDNA fragmentation genomics before and after screening, a plurality of novel cancer diagnostic markers based on the end distribution pattern of peripheral blood cfDNA can be effectively identified. In addition, the present disclosure combines machine learning to develop a cancer risk assessment model with high accuracy, which can realize cancer risk assessment for a wide range of cancers. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings needed to be used in the description of the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some of the example embodiments of the present disclosure, and other drawings can also be obtained by those of ordinary skill in the art without any creative effort on the basis of these drawings.

[0025] Figure 1 A schematic diagram of cfDNA end screening is shown;

[0026] Figure 2 A general flowchart of obtaining cancer risk prediction markers and making cancer prediction according to some embodiments of the present disclosure is shown;

[0027] Figure 3 A flowchart of a method of obtaining cancer risk prediction markers according to some embodiments of the present disclosure is shown;

[0028] Figure 4 A flowchart of a cancer risk assessment method performed by a machine according to some embodiments of the present disclosure is shown;

[0029] Figure 5 A comparison of various fragmentation characteristic values such as CCCA before and after fragment screening in liver cancer samples at different stages according to some embodiments of the present disclosure is shown;

[0030] Figure 6 In a data set, the original S150 value (A), CCCA value (B) and CTCC value (C) before fragment screening, the calculated ΔS (D), Δm (E) and Δm2 (F) after fragment screening, and the corresponding ROC curve (G) of liver cancer diagnosis according to some embodiments of the present disclosure are shown;

[0031] Figure 7 In a data set, ΔS (A) and ROC curve (B) in liver cancer group and control group, Δm (C) and Δm3 (D) in lung cancer group and control group, and ROC curve (E) according to some embodiments of the present disclosure are shown;

[0032] Figure 8 In a data set, ΔS (A), Δm (B) and Δm2 (C) in cancer group and control group, and the corresponding ROC curve (D) according to some embodiments of the present disclosure are shown;

[0033] Figure 9 In a data set, ΔS (A), Δm (B) and Δm2 (C) in control group and various types of cancer, and the ROC curves (D-F) in different cancer types according to some embodiments of the present disclosure are shown;

[0034] Figure 10The risk scores predicted by the tumor prediction machine learning models constructed with different feature combinations in one dataset are shown according to some embodiments of the present disclosure.

[0035] Figure 11 The ROC curves in different types of tumors and the overall ROC curve of the tumor prediction machine learning models constructed with different feature combinations in one dataset are shown according to some embodiments of the present disclosure.

[0036] Figure 12 The ΔS (A), Δm (B) and Δm2 (C) in the control group and different stages of colorectal cancer and the corresponding ROC curves (D-F) in one dataset are shown according to some embodiments of the present disclosure.

[0037] Figure 13 The ROC curves in different stages of colorectal cancer and the overall ROC curve of the machine learning models constructed with Δm (A-B) and Δm2 (C-D) of autosomes long arm and short arm in the genome as features in one dataset are shown according to some embodiments of the present disclosure; and

[0038] Figure 14 The structural diagram of an electronic device according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0039] In order to make the purposes, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without any creative effort fall within the protection scope of the present disclosure.

[0040] Unless otherwise defined, technical terms or scientific terms used in the present disclosure shall have the same meaning as those understood by a person of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are used to distinguish different components. The terms "include", "comprise", and similar terms mean that the elements or objects before the term encompass the elements or objects listed after the term and equivalents thereof, and do not exclude other elements or objects. The terms "connected" or "coupled" or similar terms do not mean only physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are used only to indicate relative positions, and when the absolute positions of the described objects are changed, the relative positions can also be changed accordingly. In order to keep the following description of the embodiments of the present disclosure clear and concise, the detailed description of some known functions and known components is omitted.

[0041] Flowcharts are used in the present disclosure to illustrate the steps of the methods according to the embodiments of the present disclosure. It should be understood that the preceding or subsequent steps do not necessarily proceed in sequence. Instead, various steps can be processed in reverse order or simultaneously. Meanwhile, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0042] In the specification and drawings of the present disclosure, elements are described in singular or plural form according to the embodiments. However, the selection of singular and plural forms for the proposed case is only for the convenience of explanation and is not intended to limit the present disclosure to this. Therefore, the singular form can include the plural form, and the plural form can also include the singular form, unless the context clearly indicates otherwise.

[0043] The term

[0044] A“biological sample” refers to any sample taken from a subject (e.g., a human (or other animal), such as an individual with cancer or an individual suspected of having cancer and containing one or more nucleic acid molecules of interest. The biological sample can be a bodily fluid, such as blood, plasma, serum, etc. In various embodiments, a majority of the DNA in a biological sample that has been enriched for cell-free DNA (e.g., a plasma sample obtained by a centrifugation protocol) can be cell-free, e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free. A centrifugation protocol can comprise obtaining a fluid portion at, e.g., 1,500 g x 10 minutes, and centrifuging again at, e.g., 15,000 g for an additional 10 minutes to remove residual cells. At least 1,000 cell-free DNA molecules can be analyzed as part of the analysis of the biological sample. As other embodiments, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 cell-free DNA molecules or more can be analyzed.

[0045] A“sequence read” refers to a string of nucleotides sequenced from any portion or all of a nucleic acid molecule. For example, a sequence read can be a short string of nucleotides sequenced from a nucleic acid fragment (e.g., 20-150 nucleotides), a short string of nucleotides at one or both ends of a nucleic acid fragment, or sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in a variety of ways, e.g., using sequencing technology or using probes, e.g., by hybridization arrays or capture probes or amplification techniques, such as polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. At least 1,000 sequence reads can be analyzed as part of the analysis of the biological sample. As other examples, at least 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000 sequence reads can be analyzed.

[0046] A sequence read can include an“end sequence” associated with an end of a fragment. An end sequence can correspond to the outermost N bases of a fragment, e.g., 2-30 bases of an end of a fragment. If a sequence read corresponds to an entire fragment, the sequence read can comprise two end sequences. When paired-end sequencing provides two sequence reads corresponding to ends of a fragment, each sequence read can comprise an end sequence.

[0047] The term "sequencing depth" refers to the number of times a locus is covered by sequence reads that align to the locus. The locus can be as small as a nucleotide, or as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as 50x, 100x, etc., where "x" refers to the number of times a locus is covered by sequence reads. Sequencing depth can also apply to multiple loci or an entire genome, in which case x can refer to the average number of times a locus or haploid genome or entire genome is sequenced, respectively. Ultra-deep sequencing can refer to sequencing depth of at least 100x.

[0048] The terms "cut-off" and "threshold" refer to a predetermined quantity used in an operation. For example, a cut-off size can refer to a size above which fragments are not included. A threshold can be a value applied above or below a particular classification. Either of these terms can be used in either of these contexts. A cut-off or threshold can be a "reference value" or can be derived from a reference value that represents a particular class or distinguishes between two or more classes. Such reference values can be determined in various ways, as will be appreciated by persons skilled in the art. For example, a metric can be determined for two different subjects having different known classifications, and a reference value can be selected as representative of one classification (e.g., the average) or a value between two clusters of metrics (e.g., selected to obtain a desired sensitivity and specificity). As another example, a reference value can be determined based on a statistical simulation of samples.

[0049] The term "about" or "approximately" can mean within an acceptable range of deviation of a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., limitations of the measurement system. For example, according to the practice in the art, "about" can mean within 1 or more than 1 standard deviation. Alternatively, "about" can mean a range of at most 20%, at most 10%, at most 5%, or at most 1% of a given value. Alternatively, especially for biological systems or processes, the term "about" or "approximately" can mean within an order of magnitude, within 5-fold, and more preferably, within 2-fold of a value. When a particular value is described in the application and claims, unless otherwise indicated, the term "about" is assumed to mean that the particular value is within an acceptable error range for the purposes of the application and claims. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can mean ±10%. The term "about" can mean ±5%.

[0050] The method for obtaining a cancer risk prediction marker, the cancer diagnosis method, the electronic device, and the storage medium provided by the present disclosure will be described in detail below with reference to the accompanying drawings.

[0051] <First Embodiment>

[0052] First, reference will be made to Figure 1The principles of the technology involved in the present disclosure are described.

[0053] Figure 1 A schematic diagram of cfDNA end filtering is shown. The gray shadow represents the nucleosome-protected region of blood cells. The line segments represent cfDNA molecules of background and non-blood cell origin (such as tumor origin), respectively. The cross marks represent cfDNA ends located in the linker DNA and the nucleosome-protected region, respectively. Through fragment filtering, cfDNA molecules with ends located in the nucleosome-protected region are selected, and finally the Δ value and the corresponding diagnostic parameters are calculated.

[0054] Free DNA is mostly generated from chromatin degradation during apoptosis, and its fragmentation pattern is affected by the distribution of nucleosomes: free DNA ends tend to fall on linker DNA between nucleosomes, and less on the inside of nucleosomes. In the peripheral blood of non-cancer patients, cfDNA is mostly from blood cells, and the consistency of its fragmentation pattern with the nucleosome distribution of blood cells is high; in solid tumor cancer patients, tumor and paracancerous tissues do not belong to blood cells, and their nucleosome distribution patterns are different from those of blood cells, so the DNA released by them has a lower consistency with the nucleosome distribution of blood cells. At the same time, the free DNA produced by tumors is shorter, and a higher proportion of its ends are inside the nucleosomes, which leads to a decrease in the consistency of the overall peripheral blood free DNA fragmentation pattern with the nucleosome distribution of blood cells.

[0055] Therefore, by filtering DNA sequences with cfDNA fragment ends located in the nucleosome region of blood cells, tumor-derived cfDNA can be enriched to some extent; changes in various types of cfDNA characteristics (such as fragment length, end motif, diversity score, etc.) before and after enrichment (as shown in curves 101 and 102 in FIG. 1) can be evaluated, such as difference or ratio (hereinafter collectively referred to as Δ value), to perform cancer risk assessment. The present disclosure found that the Δ value indicating the change is more significantly different between the tumor group and the non-tumor group relative to the original characteristics, and the accuracy of the risk assessment model is better. Figure 1

[0056] Based on this principle, the present disclosure performs cancer risk assessment by filtering DNA sequences with cfDNA ends located in the nucleosome region of blood cells in the blood of the subject, and then quantitatively studying the changes in various types of cfDNA characteristics before and after sequence filtering.

[0057] ​For example, after end filtering, the proportion of short free DNA fragments (e.g., less than 150 bp) will increase, and the increase in proportion of short free DNA fragments in solid tumor cancer patients is significantly higher than that in non-cancer patients, which is consistent with the fact that the DNA from tumors is shorter. The proportion of sequences that dominate in end motifs and breakpoint motifs will increase, but the increase in proportion of sequences that dominate in end motifs and breakpoint motifs in solid tumor cancer patients is significantly lower than that in non-cancer patients, which is consistent with the fact that the DNA from tumors has a lower proportion of sequences that dominate in end motifs and breakpoint motifs. Correspondingly, the diversity (or complexity, MotifDiversity) of end motifs and breakpoint motifs will increase, and the increase in diversity of end motifs and breakpoint motifs in solid tumor cancer patients is significantly higher than that in non-cancer patients, which is consistent with the fact that the DNA from tumors has a higher diversity of end motifs and breakpoint motifs. Therefore, the Δ values of various types of cfDNA features such as fragment length, end motif, and breakpoint motif before and after enrichment can be used as potential cancer diagnostic markers.

[0058] Figure 2 A general flowchart of obtaining cancer risk prediction markers and making cancer prediction according to an embodiment of the present disclosure is shown.

[0059] As shown in Figure 2 , first, the collection and processing of peripheral blood are performed. For example, peripheral blood (e.g., 6 ml) is collected from a vein using a blood collection tube containing an anticoagulant (e.g., EDTA), centrifuged at low temperature (e.g., 4°C) and medium speed (e.g., 1600 gravitational acceleration) for a period of time (e.g., 15 minutes) using a centrifuge, and then the upper liquid is aspirated and transferred to a new centrifuge tube. The centrifuge is then used at low temperature (e.g., 4°C) and high speed (e.g., 16000 gravitational acceleration) for a period of time (e.g., 15 minutes), and then the upper liquid (i.e., plasma) is aspirated and transferred to a new centrifuge tube or a cryogenic tube. The plasma can be stored in an ultra-low temperature (e.g., -80°C) refrigerator until use.

[0060] Then, the extraction, sequencing, and basic analysis of cfDNA are performed. For example, a commercial kit and equipment can be used to extract cfDNA from plasma, construct a sequencing library, and perform high-throughput sequencing. The sequencing raw data is first pre-processed to remove sequencing adapters and low-quality cycles (optional step), remove repetitive sequences (optional step), and then aligned to the human reference genome, and then remove repetitive sequences and low-quality alignment results (optional step) to obtain the final alignment results. Considering the influence of gender on the genome, the data information on chromosomes M, X, and Y can be removed. Then, based on the results, the location of one end (for single-end sequencing or double-end sequencing) or both ends (for double-end sequencing data) of each cfDNA fragment on the genome is extracted, i.e., the end site information of the cfDNA fragment.

[0061] Subsequently, the data of the distribution of blood cell nucleosome can be obtained. For example, the consistency of blood cell nucleosome can be obtained by various experimental techniques, such as Mnase-seq, ATAC-seq, ChIP-seq, etc., applied to one or more blood cell types, such as T cells, neutrophils, or peripheral blood mononuclear cells. For example, the Mnase-seq experimental data of the GM12878 cell line can be used. The GM12878 cell line here can be other blood cell lines; it can also be an internal data set constructed by other experiments, algorithms that can reflect the relative chromosomal position information of the blood cell nucleosome of the healthy population, or it can be obtained by using the free DNA of non-cancer patients instead of the Mnase-seq experimental technique.

[0062] After obtaining the data of the distribution of blood cell nucleosome, the upstream and downstream 73bp of the nucleosome center position can be defined as the nucleosome protection region. Here, 73 can be 10, 20, 30, 40, 50, 60, 70, 80, 90, etc. other positive integers (generally less than 100).

[0063] Then, the cfDNA end screening, cfDNA features, and Δ value calculation are performed.

[0064] Next, reference is made to Figure 3 The method for obtaining a cancer risk prediction marker according to the embodiments of the present disclosure is described in detail.

[0065] In step S301, a plurality of circulating cell-free DNA (cfDNA) fragments of a biological sample of a subject are analyzed to obtain sequence reads including end site information indicating a position of an end of each cfDNA fragment in the plurality of cfDNA fragments on a genome.

[0066] In step S302, data of the distribution of blood cell nucleosome is obtained, and a predetermined interval upstream and downstream of the nucleosome center position is determined as a nucleosome protection region.

[0067] In step S303, fragment screening is performed on the plurality of cfDNA fragments so as to screen sequence reads of part of the cfDNA fragments whose ends are located in the nucleosome protection region from the sequence reads of the plurality of cfDNA fragments.

[0068] In step S304, a certain fragmentation feature (including but not limited to S150, end motif, etc.) in the sequence reads of the plurality of cfDNA fragments that have not been screened is calculated, and a corresponding fragmentation feature in the sequence reads of the screened part of the cfDNA fragments is calculated.

[0069] At step S305, one or more difference values indicative of a level of change in the second fragmentation feature of the filtered portion of cfDNA fragments relative to the fragmentation feature of the plurality of cfDNA fragments are calculated as the cancer risk prediction marker.

[0070] In one embodiment, performing fragment filtering on the plurality of cfDNA fragments further comprises performing fragment filtering on the plurality of cfDNA fragments so as to filter out, from the sequence reads of the plurality of cfDNA fragments, sequence reads of cfDNA fragments having a first end or a second end located in the nucleosome-protected region.

[0071] In another embodiment, performing fragment filtering on the plurality of cfDNA fragments further comprises performing fragment filtering on the plurality of cfDNA fragments so as to filter out, from the sequence reads of the plurality of cfDNA fragments, sequence reads of cfDNA fragments having a first end and a second end both located in the nucleosome-protected region.

[0072] For example, first, end filtering is performed on all samples (including the modeling group), i.e., sequence reads having a 5’ end or a 3’ end located in the nucleosome-protected region are selected (in another embodiment, sequence reads having both a 5’ end and a 3’ end located in the nucleosome-protected region can be selected).

[0073] The fragmentation feature of a cfDNA fragment comprises at least one or more of the following: a first proportion of short fragments having a length less than a predetermined value, a second proportion of end sequence usage, a third proportion of breakpoint sequence usage, a first diversity score of end sequence, a second diversity score of breakpoint sequence, and other biological information present on the cfDNA fragment, including copy number variation, mutation information.

[0074] For example, the fragmentation features, such as the proportion of short fragments (in this embodiment, the proportion of cfDNA having a length less than 150 bp S150 is used, but the length cutoff value can also be other numbers, such as 100, 113, 128, 130, 140, 160, etc.), cfDNA end base sequence usage pattern (in this embodiment, the proportion of CCCA end sequence usage and the proportion of CTCC breakpoint sequence usage are used, but other sequences, such as CCTG, can also be used), etc., can be calculated in the raw cfDNA data and the cfDNA data after fragment filtering, respectively.

[0075] Finally, the difference value, i.e., the Δ value, of the same cfDNA feature after fragment filtering is calculated. In this embodiment, subtraction is mainly used to calculate the difference value, but other methods, such as ratio, can also be used as a method for calculating the Δ value.

[0076] In one embodiment, the difference value obtained by subtracting the first proportion in fragmentation characteristics of the plurality of cfDNA fragments from the first proportion in fragmentation characteristics of the filtered portion of cfDNA fragments is calculated as a first cancer risk prediction marker As.

[0077] In another embodiment, the difference value obtained by subtracting the second proportion in fragmentation characteristics of the plurality of cfDNA fragments from the second proportion in fragmentation characteristics of the filtered portion of cfDNA fragments is calculated as a second cancer risk prediction marker Am.

[0078] In another embodiment, the difference value obtained by subtracting the third proportion in fragmentation characteristics of the plurality of cfDNA fragments from the third proportion in fragmentation characteristics of the filtered portion of cfDNA fragments is calculated as a third cancer risk prediction marker Am2.

[0079] In another embodiment, the difference value obtained by subtracting the first diversity score in fragmentation characteristics of the plurality of cfDNA fragments from the first diversity score in fragmentation characteristics of the filtered portion of cfDNA fragments is calculated as a fourth cancer risk prediction marker Am3.

[0080] In another embodiment, the difference value obtained by subtracting the second diversity score in fragmentation characteristics of the plurality of cfDNA fragments from the second diversity score in fragmentation characteristics of the filtered portion of cfDNA fragments is calculated as a fifth cancer risk prediction marker Am4.

[0081] For example, As is obtained by subtracting S150 of original cfDNA from S150 of filtered end, Am is obtained by subtracting the proportion of CCCA end motif of original cfDNA from the proportion of CCCA end motif of filtered end, Am2 is obtained by subtracting the proportion of CTCC breakpoint motif of original cfDNA from the proportion of CTCC breakpoint motif of filtered end, Am3 is obtained by subtracting the end motif diversity score of original cfDNA from the end motif diversity score of filtered end, and Am4 is obtained by subtracting the breakpoint motif diversity score of original cfDNA from the breakpoint motif diversity score of filtered end. Similarly, changes in other types of end motif and breakpoint motif after end filtering, and change values of cfDNA of other lengths can be calculated.

[0082] The one or more cancer risk prediction markers are used to indicate non-specific multiple cancer types. Furthermore, the one or more cancer risk prediction markers are used to indicate non-specific multiple cancer types, individually or in combination.

[0083] In the following, a method for cancer risk assessment based on the one or more cancer risk prediction markers according to embodiments of the present application will be described in detail. Figure 4 A method for cancer risk assessment based on the one or more cancer risk prediction markers according to embodiments of the present application will be described in detail.

[0084] At step 401, a first set of biological samples of normal subjects can be obtained, and for the first set of biological samples, a first set of cancer risk prediction markers associated with the first set of biological samples can be calculated using the method according to the above embodiments.

[0085] At step 402, a second set of biological samples of a patient to be tested can be obtained, and for the second set of biological samples, a second set of cancer risk prediction markers associated with the second set of biological samples can be calculated using the method according to the above embodiments.

[0086] At step 403, it is determined whether the change value of the second set of cancer risk prediction markers relative to the first set of cancer risk prediction markers exceeds a predetermined threshold.

[0087] At step 404, in response to the change value of the second set of cancer risk prediction markers relative to the first set of cancer risk prediction markers exceeding the predetermined threshold, an indication that the patient to be tested has cancer is outputted.

[0088] In one embodiment, for each cancer risk prediction marker in the second set of cancer risk prediction markers, it is determined whether the change value relative to the corresponding cancer risk prediction marker in the first set of cancer risk prediction markers exceeds a predetermined threshold. The predetermined threshold corresponding to each cancer risk prediction marker can be different.

[0089] In another embodiment, a plurality of cancer risk prediction markers in the second set of cancer risk prediction markers can be integrated to obtain a single cancer risk prediction marker. Then, for the single cancer risk prediction marker obtained from the second set of cancer risk prediction markers, it is determined whether the change value relative to the corresponding cancer risk prediction marker obtained from the first set of cancer risk prediction markers exceeds a predetermined threshold.

[0090] In another embodiment, multiple cancer risk prediction markers in the second set of cancer risk prediction markers can be integrated with other biomarkers (e.g., as image data, cfDNA methylation level, length distribution, etc.) to obtain individual cancer risk prediction markers. Then, for the individual cancer risk prediction markers obtained from the second set of cancer risk prediction markers, it is determined whether the change value relative to the corresponding cancer risk prediction markers obtained from the first set of cancer risk prediction markers exceeds a predetermined threshold.

[0091] In one embodiment, a cancer risk assessment model is generated by machine learning, which integrates multiple cancer risk prediction markers in the second set of cancer risk prediction markers to obtain individual cancer risk prediction markers. Then, the cancer risk assessment model is used to integrate multiple cancer risk prediction markers in the second set of cancer risk prediction markers with other biomarkers to obtain individual cancer risk prediction markers.

[0092] For example, a non-cancer (normal subject) control sample set and a cancer patient sample set (which can be a single cancer type or multiple cancer types) can be constructed. Then, according to the method described in the above embodiment, the Δ value of cfDNA features of each sample is calculated respectively. The Δ value can include multiple types of cfDNA features, such as Δs, Δm, etc., and these Δ values can be increased (such as Δs, Δm3, Δm4) or decreased (such as Δm, Δm2) in the cancer sample group compared to the non-cancer sample group, which is due to the biological characteristics of the free DNA itself, so different cutoff values (thresholds) need to be used to distinguish different types of Δ values (i.e., diagnostic parameters) in order to achieve cancer diagnosis.

[0093] Multiple Δ values can also be integrated, and then machine learning can be used to train a set of pre-defined cancer patients and control patients, and the prediction value generated thereby can be used to establish its diagnostic parameters. The accuracy, sensitivity, and specificity of cancer diagnosis using different cutoff values can be evaluated using the receiver operating characteristic curve (ROC), and the overall effect of the diagnostic method can be evaluated using the AUC (Area Under the Curve) value: the higher the AUC (maximum 1) the better the effect. Finally, the cutoff value is determined according to the actual requirements, such as community screening which may require a specificity of > 99%, while high-risk groups may require a sensitivity of > 95%.

[0094] In the following, the present disclosure will be verified by combining Figures 5-12 different data sets.

[0095] The inventors of the present disclosure used the nucleosome position information of the GM12878 cell line in the NucMap database as the blood cell nucleosome position reference data, and used 6 sets of data collections to verify the present application. The 6 sets of data collections are from different sources and have different experimental schemes, but all contain at least one type of cancer patients and non-cancer patients (control group). The information of the 6 data sets is as follows:

[0096] Dataset 1 is shown in Figure 5 A-5X.

[0097] Data set 1 comes from 24 non-tumor control cfDNA samples and 56 liver cancer samples (31 A-stage, 10 B-stage, and 15 C-stage liver cancer cfDNA samples) sequenced by the inventors themselves. In Figure 5 In A-5L, the comparison of the CCCA values before and after fragment screening in liver cancer samples of different stages (A), the original CCCA values (B), Δm (C) and the liver cancer diagnosis ROC curve of Δm (D), the Rm distribution calculated by the ratio method (E) and the liver cancer diagnosis ROC curve of Rm (F); the comparison of CTCC values before and after fragment screening (G), the original CTCC values (H), Δm2 (I) and the liver cancer diagnosis ROC curve of Δm2 (J), the Rm2 calculated by the ratio method (K) and the liver cancer diagnosis ROC curve of Rm2 (L) are shown.

[0098] In Figure 5 In M-5X, the numerical comparison of motif diversity before and after fragment screening (M), the original motif diversity (N), Δm3 (O) and the liver cancer diagnosis ROC curve of Δm3 (P), the Rm3 calculated by the ratio method (Q) and the liver cancer diagnosis ROC curve of Rm3 (R); the numerical comparison of S150 before and after fragment screening (S), the original S150 value (T), ΔS (U) and the liver cancer diagnosis ROC curve of ΔS (V), the Rs calculated by the ratio method (W) and the liver cancer diagnosis ROC curve of Rs (X) are shown.

[0099] Before cfDNA fragment screening, the CCCA values showed no significant difference between the control group and the liver cancer group (P>0.05, Figure 5 B); after fragment screening, Δm was positive in all patients, but lower in the liver cancer group than in the control group, and the difference between the two groups was significant (P<0.05, Figure 5 C). Therefore, Δm can be used for cancer diagnosis, while the original CCCA value cannot. ROC analysis shows that Figure 5 D), Δm can effectively distinguish liver cancer patients from non-liver cancer patients: the AUCs in A, B, and C stage liver cancer are 0.81, 0.79, and 0.87, respectively, and the P values are all less than 0.05. The same phenomenon can be observed in the CTCC values Figure 5G-I), the AUC values of Am2 in A, B, C stage of liver cancer were 0.80, 0.81, 0.92, respectively, and the P values were all less than 0.05 Figure 5 J) S150 in liver cancer group was significantly higher than that in control samples, and the higher the stage, the greater the value; in addition, the diversity score was used to measure the diversity of the motif, that is, where Pi is the frequency of use of 256 4bp end motifs (this definition is mentioned in the previous US patent). After end enrichment, the diversity score decreased, that is, Am3 was negative in all samples Figure 5 O), and Am3 in C stage was significantly higher than that in the control group, so it can be used as a potential diagnostic marker. ROC analysis showed that the AUC of liver cancer in C stage was 0.76 Figure 5 P), and the P values were all less than 0.05.

[0100] S150 in liver cancer group was significantly higher than that in control samples, and the higher the stage, the greater the value; after cfDNA fragment screening, As was positive, but higher in liver cancer group than in control group. Both S150 and As can effectively separate liver cancer patients from control group, such as As in A, B, C stage of liver cancer, AUC was 0.80, 0.82, 0.92, respectively, and the P values were all less than 0.05 Figure 5 V); but there was no significant difference between the performance of S150 and As (P>0.1 compared with AUC).

[0101] In order to illustrate that the calculation method of the difference value can be difference or other types of data calculation method, we calculated the ratios Rm, Rm2, Rm3 and Rs of CCCA, CTCC, diversity score and S150 before and after fragment screening. ROC analysis proved that this kind of difference calculation method can also be used as a biomarker for tumor diagnosis. ROC analysis is as follows Figure 5 F, L, R and X, where the AUC values in liver cancer at different stages were between 0.76 and 0.96, and the P values were all less than 0.05.

[0102] Dataset 2 is shown in Figure 6 Table 2.

[0103] Dataset 2 came from Zhang et al. (BMC Medicine 2020, Aug 3; 18(1): 200), including 37 non-cancer patients and 8 liver cancer. In Figure 6 In Zhang et al. data, the original S150 value before fragment screening (A), CCCA value (B) and CTCC value (C), the calculated AS (D), Am (E) and Am2 (F) after fragment screening, and the corresponding ROC curve for liver cancer diagnosis (G) are shown.

[0104] Without end-stage screening, S150, CCCA, and CTCC values ​​were all significantly different between the control group and the liver cancer patient group. Figure 6 AC); after end-stage screening, Δs, Δm, and Δm2 also showed significant differences between the control group and the liver cancer patient group. Figure 6 The ROC analysis showed that the difference was more significant than that of S150, CCCA, and CTCC values, and the accuracy for diagnosis was also higher. Figure 6 G), these three indicators can effectively distinguish liver cancer patients from non-cancer patients, with AUC values ​​of 0.87, 0.85 and 0.85 respectively, and P values ​​of less than 0.05.

[0105] Dataset 3 is shown in Figure 7 Dataset 3 is shown in

[0106] Dataset 3, from Liang et al. (Clin Transl Med. 2020 Oct; 10(6):e177), includes 10 non-cancer patients, 10 liver cancer patients, and 10 lung cancer patients. Figure 7 In the dataset of Liang et al., ΔS(A) and ROC curves(B) for the liver cancer group and the control group are shown; Δm(C) and Δm3(D) for the lung cancer group and the control group, and ROC curves(E) are shown.

[0107] Δs showed a significant difference between the control group and the liver cancer patient group. Figure 5 A), ROC analysis showed that its AUC value reached 0.92, and the P value was less than 0.05 (A). Figure 7 B). Δm and Δm3 showed significant differences between the control group and the lung cancer patient group. Figure 7 ROC analysis showed that the AUC values ​​reached 0.92 and 1, respectively, and the P value was less than 0.05. Figure 7 E).

[0108] Dataset 4 is shown in Figure 8 Table 4.

[0109] Dataset 4 comes from Song et al. (Cell Res. 2017 Oct; 27(10):1231-1242), and includes 7 non-cancer patients and 39 cancer patients (4 brain tumor patients, 2 breast cancer patients, 4 colorectal cancer patients, 4 gastric cancer patients, 3 liver cancer patients, 15 lung cancer patients, and 7 pancreatic cancer patients). Figure 8 In the dataset of Song et al., ΔS(A), Δm(B) and Δm2(C) for the cancer group and the control group are shown, along with the corresponding ROC curves (D).

[0110] Among these cancer patients, Δs, Δm, and Δm2 were all significantly different between the control group and the cancer patient group. Figure 8A-C). ROC analysis showed that the AUC of these three indexes were 0.75, 0.81 and 0.84, respectively, with P value less than 0.05.

[0111] Dataset 5 is shown in Figures 9-11 Table 1.

[0112] Dataset 5 was from Cristiano et al. (Nature 2019; 570: (385-389)). Only part of the samples in it were used in this patent, including 215 non-cancer patients, 26 cholangiocarcinoma patients, 54 breast cancer patients, 27 colorectal cancer patients, 34 pancreatic cancer patients, 12 lung cancer patients, and 28 ovarian cancer patients. Figure 9 In the Cristiano et al. dataset, ΔS(A), Δm(B) and Δm2(C) in the control group and multiple types of cancer, and their ROC curves in different cancer types (D-F) were shown.

[0113] In these cancer patients, Δs was significantly higher than that in the control group Figure 9 A). ROC analysis showed that the AUC of Δs in these cancer types for cancer diagnosis was between 0.63 and 0.92, with P value less than 0.05 Figure 9 D). Δm and Δm2 also had significant differences between the control group and the cancer patient group Figure 9 B-C); ROC analysis showed that the AUC of Δm and Δm2 in these cancer types for cancer diagnosis was between 0.60 and 0.97, with P value less than 0.05 Figure 9 E-F).

[0114] In Figure 10 In the Cristiano et al. dataset, the risk prediction values in different types of tumors using multiple feature combinations to build tumor prediction machine learning models were shown. In Figure 11 In the Cristiano et al. dataset, the risk prediction values in different types of tumors using different feature combinations to build tumor prediction machine learning models, as well as the ROC curves and overall ROC curves were shown: Δs100, Δs101, and up to Δs200 (A-B) were used as features to build models; the numerical changes of all 4bpend motifs after fragment selection (C-D) were used as features to build models; the Δm and Δm2 of the long arm and short arm of the autosomes in the genome were used as features (E-F) to build models; Δs (G-H), Δm (I-J), and N-index (K-L) based on the 5 Mb long genomic intervals on the autosomes in the genome were used to build models, respectively.

[0115] On this basis, the disclosure extends the value of Δ, which can be calculated for different lengths, end sequences or genomic intervals, and then uses one or several types of Δ values as features to construct a machine learning model for tumors and non-tumors. Next, the Cristiano dataset is used as an example for demonstration. As follows:

[0116] (1) Calculate the proportion of cfDNA of a specific length less than 100bp, 101bp, and up to 200bp after fragment selection, defined as Δs100, Δs101, and up to Δs200, a total of 101 values. Then use these as multi-dimensional features to construct a binary classification model for tumor and non-cancer patients using machine learning algorithms (XGBoost algorithm in this embodiment). 10-fold cross-validation is used, and the process is repeated 10 times to evaluate model performance. ROC analysis shows that the AUC value in these cancers is between 0.89 and 0.95 (A), and the overall AUC for multiple cancers is 0.92 (B), with P values less than 0.05. Figure 11 A), the overall AUC for multiple cancers is 0.92 (B), and the P values are all less than 0.05. Figure 11 B).

[0117] (2) Calculate the numerical change of all 4bp end motifs after fragment selection, a total of 256 values. Then use these as multi-dimensional features to construct a binary classification model for tumor and non-cancer patients using machine learning algorithms. Similarly, 10-fold cross-validation is used, and the process is repeated 10 times. ROC analysis shows that the AUC value in these cancers is between 0.94 and 0.99 (C), and the overall AUC for multiple cancers is 0.97 (D), with P values less than 0.05. Figure 11 C), the overall AUC for multiple cancers is 0.97 (D), and the P values are all less than 0.05. Figure 11 D).

[0118] (3) Calculate the Δ values of different genomic intervals, such as calculating the Δm and Δm2 of the long arm and short arm of the autosomes in the genome, a total of 78 values. Then use these as multi-dimensional features to construct a binary classification model for tumor and non-cancer patients using machine learning algorithms. Similarly, 10-fold cross-validation is used, and the process is repeated 10 times to evaluate model performance. ROC analysis shows that the AUC value in these cancers is between 0.80 and 0.97 (E), and the overall AUC for multiple cancers is 0.86 (F), with P values less than 0.05. Figure 11 E), the overall AUC for multiple cancers is 0.86 (F), and the P values are all less than 0.05. Figure 11 F).

[0119] (4) In addition to the above-mentioned long arm and short arm of the chromosome, the interval of the genome can also be divided into 5Mb long intervals that are continuous and non-overlapping, filter out the regions without DNA sequence reads, leaving 542 intervals, and then calculate the corresponding Δ value of each interval, such as calculating Δs and Δm of each interval in all 542 intervals, and N-index. Then, respectively, using machine learning algorithm to construct the binary classification model of tumor and non-cancer patients, to predict the probability of tumor. ROC analysis shows that the AUC value of the model based on Δm as the feature is between 0.77 and 0.93 in these cancers Figure 11 G), the overall AUC of multiple cancers is 0.85 Figure 11 H), P value is less than 0.05; the AUC value of the model based on Δs as the feature is between 0.88 and 0.98 Figure 11 I) in these cancers, the overall AUC of multiple cancers is 0.90 Figure 11 J), P value is less than 0.05; the AUC value of the model based on N-index as the feature is between 0.94 and 0.99 Figure 11 K) in these cancers, the overall AUC of multiple cancers is 0.96 Figure 11 L), P value is less than 0.05, the sensitivity reaches 83% at 95% specificity, which is better than the DELFI method developed based on the same data, and also better than using the overall N-index for diagnosis.

[0120] In order to predict cancer patients, the present disclosure uses, for example, gradient boosting decision tree (XGBoost) algorithm to construct machine learning model, and then uses cross-validation to estimate the diagnostic effect of the model. For example, the model is trained on 80% of the data, and the remaining 20% of the data is used to test the model. Figure 11As shown, first, all autosomes of human reference genome were divided into consecutive, non-overlapping 5 Mb long windows, and the windows with low coverage were removed (such as telomeres, centromeres), finally 542 genomic regions were obtained for downstream analysis (covering about 2.7 Gb of genome). For each sample, the N-index (or Δ value, different markers can obtain different models) of each genomic window region was calculated, obtaining 542 values. Then, the 542 ΔS values of all test sets and training sets were standardized respectively, with mean value of 0 and standard deviation of 1, and were used as multi-dimensional features to construct the binary classification model of tumor and non-tumor. In the training set, AUC was used as the evaluation index to randomly search and iterate 300 times for hyperparameter optimization and feature selection, and the model constructed on the training set was evaluated on the test set. Finally, in order to average the prediction error caused by the difference of sample grouping, cross-validation was repeated for 10 times again, and the average value of 10 model prediction results of each sample was taken as the final EXCEL cancer risk index of the individual.

[0121] Dataset 6 is shown in Figures 12-13 Table 6.

[0122] Dataset 6 is from Walker et al. (Sci Rep. 2022 Oct 4; 12(1): 16566), containing 274 non-cancer patients and 343 colorectal cancer patients. In Figure 12 In Walker et al. dataset, ΔS (A), Δm (B) and Δm2 (C) in control group and colorectal cancer patients of different stages, and the corresponding ROC curves (D-F) are shown.

[0123] In colorectal cancer patients, Δs is higher than that in non-cancer patients Figure 12 A) and the median value of its number also shows an increasing trend with the increase of stage, while Δm and Δm2 are lower than those in non-cancer patients Figure 12 B-C). ROC analysis shows that the AUC of Δs in colorectal cancer patients of stages I, II, III and IV is 0.61, 0.61, 0.63 and 0.77 respectively Figure 12 D), the AUC of Δm in colorectal cancer patients of stages II, III and IV is 0.56, 0.61 and 0.71 respectively Figure 12 E), and the AUC of Δm2 in colorectal cancer patients of stages II, III and IV is 0.57, 0.60 and 0.68 respectively Figure 12 F), and the P value is less than 0.05.

[0124] The number of samples in this dataset is larger, and a machine learning model can also be established. In Figure 13In Walker et al's data set, the machine learning model was constructed by taking the values of Δm(A-B) and Δm2(C-D) of the long arm and short arm of autosomes in the genome as features, and the ROC curves in different stages of colorectal cancer and the overall ROC curve.

[0125] As Δs on the long arm and short arm of the chromosome in the genome was calculated respectively, 39 values were obtained, and the same as the above case, they were taken as features to further construct a machine algorithm model for binary classification prediction of tumor and non-cancer patients. ROC analysis showed that the AUC of colorectal cancer patients in stages I, II, III and IV was 0.74, 0.78, 0.75 and 0.83 respectively, and the overall AUC value of colorectal cancer patients was 0.77 Figure 13 As Δm on the long arm and short arm of the chromosome in the genome was calculated respectively, 39 Δm values were obtained as features, and a binary classification model of tumor and non-cancer patients was constructed by machine algorithm. ROC analysis showed that the AUC of colorectal cancer patients in stages I, II, III and IV was 0.68, 0.65, 0.64 and 0.81 respectively, and the overall AUC value of colorectal cancer patients was 0.67 Figure 13 C-D), P values were all less than 0.05, and were also higher than the AUC value of the whole genome Δm value for diagnosis.

[0126] The cancer risk assessment method of the present disclosure can be used for various cancer types, has high universality and is widely applicable. At the same time, the diagnostic accuracy of the present disclosure is relatively high, especially in early cancer patients, which is greatly improved compared with the current method. In addition, the data used in the present disclosure can be obtained only by using common whole genome sequencing, and there is no special requirement for experimental scheme and sequencing platform. The operation complexity is low, the cost is low, and it is suitable for clinical promotion.

[0127] <Second embodiment>

[0128] Figure 14 A structural diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0129] Referring to Figure 14 The electronic device 700 can include a processor 701 and a memory 702. The processor 701 and the memory 702 can be connected through a bus 703. The electronic device 700 can be any type of portable device (such as a smart camera, a smart phone, a tablet computer, etc.), any type of fixed device (such as a desktop computer, a server, etc.), or an integrated circuit such as a digital signal processor, a single-chip microcomputer, a system on chip (SoC), etc.

[0130] The processor 701 can perform various actions and processes according to programs stored in the memory 702. Specifically, the processor 701 can be an integrated circuit chip having a processing capability of signals. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like, which can be of X86 architecture or ARM architecture.

[0131] The memory 702 stores computer executable instructions that, when executed by the processor 701, implement the above medical large language model training and logical reasoning method. The memory 702 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DRAM). It should be noted that the memory of the method described herein is intended to include, but not limited to, these and any other suitable types of memory.

[0132] In addition, the computer-implemented method for obtaining a cancer risk prediction marker and the cancer diagnosis method according to the present disclosure can be recorded in a computer-readable recording medium. Specifically, according to the present disclosure, a computer-readable recording medium storing computer executable instructions can be provided, which, when executed by a processor, can cause the processor to perform the method for obtaining a cancer risk prediction marker and the cancer risk assessment method as described above.

[0133] It should be noted that the flowcharts and block diagrams in the drawings are illustrations of the possible architectures, functional processes, and operations of systems, methods, and computer program products in accordance with various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code that comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.

[0134] In general, the various example embodiments of the present disclosure can be implemented in hardware or special-purpose circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software which can be executed by a controller, microprocessor or other computing device, Although the various aspects of embodiments of the present disclosure can be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein can be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controler or other computing devices, or some combination thereof.

[0135] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0136] The foregoing is a summary of the present disclosure, and is not to be considered as limiting its scope. While several example embodiments of the present disclosure have been described, it will be apparent to those of ordinary skill in the art that many modifications are possible without departing from the novel teachings and advantages of the present disclosure. The disclosed embodiments are intended to be merely illustrative of the present disclosure and not limiting thereof. All modifications coming within the meaning and range of equivalency of the claims are to be embraced within the scope of the present disclosure. It is to be understood that all the terms and expressions of necessity used herein are to be interpreted from the context of this disclosure and are intended to be used only in conformity with the present disclosure. The present disclosure is limited only by the claims and their equivalents.

Claims

1. A machine-executed method for obtaining cancer risk predictive biomarkers, comprising: Analyze multiple peripheral blood circulating cell-free DNA (cfDNA) fragments from a subject's biological sample to obtain sequence reads, the sequence reads including end-site information indicating the location of the end of each of the multiple cfDNA fragments on the genome; Using blood cell nucleosome distribution data, and determining the predetermined upstream and downstream intervals of the nucleosome center as the nucleosome protection region; Sequence reads of a subset of cfDNA fragments whose ends are located in the nucleosome protection region are selected from the sequence reads of the plurality of cfDNA fragments to enrich tumor-derived cfDNA fragments; Calculate a first fragmentation feature from the sequence reads of the plurality of cfDNA fragments, and calculate a second fragmentation feature from the sequence reads of a selected subset of cfDNA fragments. The first and second fragmentation features include at least one or more of the following: fragment length, end sequence, and breakpoint sequence of the cfDNA fragment. For the same cfDNA feature in the first fragmentation feature and the second fragmentation feature, one or more difference values ​​indicating the change level of the second fragmentation feature relative to the first fragmentation feature of the plurality of cfDNA fragments are calculated by subtracting the first fragmentation feature from the second fragmentation feature as the cancer risk prediction biomarker.

2. The method of claim 1, wherein, The terminal position information indicates the location of the first and / or second ends of each of the plurality of cfDNA fragments on the genome, and Performing fragment screening on the plurality of cfDNA fragments further includes: Fragment screening is performed on the plurality of cfDNA fragments to select sequence readings of cfDNA fragments whose first and / or second ends are located in the nucleosome protection region from the sequence readings of the plurality of cfDNA fragments.

3. The method according to any one of claims 1-2, wherein, The fragmentation characteristics of cfDNA fragments include at least one or more of the following: Short fragments shorter than the predetermined length constitute the largest proportion of cfDNA fragments. The second proportion used in the terminal sequence, The third ratio used in the breakpoint sequence, The first diversity score of the terminal sequence, The second diversity score of the breakpoint sequence, and Other biological information present on cfDNA fragments includes copy number changes and mutation information.

4. The method according to claim 3, wherein, Calculating one or more difference values ​​indicating the level of change in the second fragmentation feature of the selected partial cfDNA fragments relative to the first fragmentation feature of the plurality of cfDNA fragments as the cancer risk predictor includes: The difference value obtained by subtracting the first proportion of the fragmentation features of the plurality of cfDNA fragments from the first proportion of the fragmentation features of the selected partial cfDNA fragments is used as the first cancer risk prediction biomarker Δs.

5. The method according to claim 3, wherein, Calculating one or more difference values ​​indicating the level of change in the second fragmentation feature of the selected partial cfDNA fragments relative to the first fragmentation feature of the plurality of cfDNA fragments as the cancer risk predictor includes: The difference value obtained by subtracting the second proportion of the fragmentation characteristics of the plurality of cfDNA fragments from the second proportion of the fragmentation characteristics of the selected partial cfDNA fragments is used as the second cancer risk prediction biomarker Δm.

6. The method according to claim 3, wherein, Calculating one or more difference values ​​indicating the level of change in the second fragmentation feature of the selected partial cfDNA fragments relative to the first fragmentation feature of the plurality of cfDNA fragments as the cancer risk predictor includes: The difference value obtained by subtracting the third proportion of the fragmentation characteristics of the plurality of cfDNA fragments from the third proportion of the fragmentation characteristics of the selected partial cfDNA fragments is used as the third cancer risk prediction biomarker Δm2.

7. The method according to claim 3, wherein, Calculating one or more difference values ​​indicating the level of change in the second fragmentation feature of the selected partial cfDNA fragments relative to the first fragmentation feature of the plurality of cfDNA fragments as the cancer risk predictor includes: The difference value obtained by subtracting the first diversity score in the fragmentation features of the selected partial cfDNA fragments from the first diversity score in the fragmentation features of the plurality of cfDNA fragments is used as the fourth cancer risk prediction biomarker Δm3.

8. The method according to claim 3, wherein, Calculating one or more difference values ​​indicating the level of change in the second fragmentation feature of the selected partial cfDNA fragments relative to the first fragmentation feature of the plurality of cfDNA fragments as the cancer risk predictor includes: The difference value obtained by subtracting the second diversity score in the fragmentation features of the selected partial cfDNA fragments from the second diversity score in the fragmentation features of the plurality of cfDNA fragments is used as the fifth cancer risk prediction biomarker Δm4.

9. The method according to claim 1, wherein, The one or more cancer risk prediction biomarkers are used individually or in combination to indicate multiple non-specific cancer types.

10. An apparatus for cancer risk assessment, the apparatus performing a cancer risk assessment method, comprising: Obtain the first biological sample set from normal subjects; For the first biological sample set, a first set of cancer risk prediction biomarkers associated with the first biological sample set are calculated using the method according to any one of claims 1-9; Obtain the second set of biological samples from the patient to be tested; For the second biological sample set, a second set of cancer risk prediction biomarkers associated with the second biological sample set are calculated using the method according to any one of claims 1-9; Determine whether the change in the second group of cancer risk prediction biomarkers relative to the first group of cancer risk prediction biomarkers exceeds a predetermined threshold; as well as In response to the change in the second set of cancer risk prediction biomarkers relative to the first set of cancer risk prediction biomarkers exceeding the predetermined threshold, an indication of the risk that the patient under test has cancer is output.

11. The apparatus according to claim 10, wherein, Determining whether the change in the second group of cancer risk predictor markers relative to the first group of cancer risk predictor markers exceeds a predetermined threshold includes: For each cancer risk predictor in the second group of cancer risk predictors, determine whether the change value of the corresponding cancer risk predictor relative to the first group of cancer risk predictors exceeds a predetermined threshold.

12. The apparatus according to claim 11, wherein, Each cancer risk prediction biomarker has a different predetermined threshold.

13. The apparatus according to claim 10, wherein, Determining whether the change in the second group of cancer risk predictor markers relative to the first group of cancer risk predictor markers exceeds a predetermined threshold includes: Multiple cancer risk predictive biomarkers in the second group are integrated to obtain a single cancer risk predictive biomarker. For a single cancer risk predictor obtained from the second group of cancer risk predictor markers, determine whether the change value relative to the corresponding cancer risk predictor obtained from the first group of cancer risk predictor markers exceeds a predetermined threshold.

14. The apparatus according to claim 13, wherein, Determining whether the change in the second group of cancer risk predictor markers relative to the first group of cancer risk predictor markers exceeds a predetermined threshold includes: Multiple cancer risk predictive biomarkers in the second group are integrated with other biomarkers to obtain individual cancer risk predictive biomarkers. For a single cancer risk predictor obtained from the second group of cancer risk predictor markers, determine whether the change value relative to the corresponding cancer risk predictor obtained from the first group of cancer risk predictor markers exceeds a predetermined threshold.

15. The apparatus according to claim 13 or 14, wherein, A cancer risk assessment model is generated through machine learning. This model integrates multiple cancer risk predictive biomarkers from the second set of cancer risk predictive biomarkers to obtain a single cancer risk predictive biomarker, or The cancer risk assessment model integrates multiple cancer risk predictive biomarkers from the second group of cancer risk predictive biomarkers with other biomarkers to obtain individual cancer risk predictive biomarkers.

16. An electronic device comprising: processor; as well as A memory, wherein the memory stores computer-readable code that, when executed by the processor, implements the method of any one of claims 1-9.

17. A non-transitory computer-readable storage medium storing computer-readable instructions, wherein, When the computer-readable instructions are executed by a processor, they implement the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Cancer diagnosis model based on free DNA and application

    CN115019952A

  • Disease detection in liquid biopsies

    WO2021130356A1

Cited By

  • Cancer diagnosis model based on cfDNA transposon fragmentation mode and application

    CN122038562A