Method for labeling free DNA derived from tumor cell on basis of protection region of nucleosome

Through the cfDNA fragment screening method based on nucleosome protected areas, the terminal location of free DNA from tumor cells was marked, which solved the problem of low signal-to-noise ratio of tumor detection in low-deep whole genome sequencing, and achieved high sensitivity and high specificity detection of early tumors.

WO2025152519A1PCT designated stage expired Publication Date: 2025-07-24SHANGHAI WEIHE MEDICAL LAB CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/124171
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2024-10-11
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

In early tumor detection, the existing low-deep whole genome sequencing methods have a very low signal-to-noise ratio due to the low content of free DNA from tumor cells, which makes it impossible to accurately distinguish between tumor samples and normal samples. The existing fragment enrichment strategies have problems such as large library losses or reduced detection performance.

Method used

By labeling the end positions of the free DNA of the subject to be tested based on the nucleosome protected area of a healthy subject, the DNA that is located in the nucleosome protected area or does not overlap with it is labeled as tumor-derived free DNA, and cfDNA fragment screening is performed using the baseline of the nucleosome protected area to enrich DNA from tumor cell-derived DNA.

Benefits of technology

It significantly improves the sensitivity and specificity of tumor detection, enhances the detection ability of early tumors, reduces interference with normal cellular DNA, and improves the enrichment efficiency of ctDNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124171_24072025_PF_FP_ABST
    Figure CN2024124171_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A method for labeling free DNA derived from a tumor cell on the basis of a protection region of a nucleosome, which method can be used to label free DNA derived from a tumor cell in sequencing data, and can improve the sensitivity and specificity of tumor detection.
Need to check novelty before this filing date? Find Prior Art

Description

A method for labeling free DNA from tumor cells based on nucleosome protection regions Technical Field

[0001] The present application relates to the field of biomedicine and to a method for labeling free DNA derived from tumor cells based on nucleosome protection regions. Background Art

[0002] Low-pass whole genome sequencing (lp-WGS), as a simple and low-cost method, is being increasingly used in tumor detection. This method performs low-depth whole genome sequencing (usually between 0.1X and 10X) on all cell-free DNA in plasma samples, namely cfDNA, and then constructs a machine learning classifier based on the characteristic differences in sequence fragments between cfDNA derived from tumor cells (i.e. ctDNA) and cfDNA derived from normal cells to distinguish tumor samples from normal samples. However, this method of tumor detection based on all sequencing data is naturally limited by the content of ctDNA derived from tumor cells in the sample, especially in the early stages. Since the content of ctDNA in early tumor plasma samples is usually low (<1%), a sufficient amount of ctDNA cannot be released into the blood circulation, resulting in an extremely low signal-to-noise ratio for the tumor signal under the huge background interference of cfDNA from normal cells, making it impossible to accurately detect tumor samples.

[0003] Numerous studies have shown that tumor-derived ctDNA fragments are shorter than normal cell-derived cfDNA fragments. Therefore, a potential strategy for ctDNA enrichment is to select shorter cfDNA fragments to increase the ctDNA content in the sample data, thereby improving detection sensitivity. This short fragment enrichment strategy can generally be implemented at two levels. One is to select for short fragments (e.g., <150 base pairs) during library construction prior to sequencing. The other is to select for short fragments (e.g., <150 base pairs) directly based on the size of the cfDNA fragments in the resulting data after sequencing. While both approaches can achieve ctDNA enrichment, they each have significant drawbacks. For example, the former results in significant library loss, requiring a higher starting amount of plasma DNA and adding additional steps to the process, while the latter results in significant data loss (often exceeding 80%), which can reduce detection performance. Furthermore, many recent studies have shown that ctDNA is more prevalent in longer fragments, such as those between 240 and 330 base pairs. Therefore, simply enriching ctDNA based on the size of cfDNA fragments is not an ideal method.

[0004] There is a constant need in the art for more sensitive and precise detection methods.

[0005] Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method for labeling free DNA derived from tumor cells based on the nucleosome protection regions of healthy subjects, which can enrich free DNA derived from tumor cells in sequencing data and significantly improve the sensitivity and specificity of tumor detection.

[0007] The information generated by the method of the present invention is an intermediate result for diagnosis and cannot be directly used as a basis for determining whether a subject has cancer.

[0008] In one aspect, the present invention provides a method for labeling cell-free DNA derived from tumor cells, comprising the following steps:

[0009] (A) Determine the terminal positions of the free DNA of the test subject based on the nucleosome protection region of the healthy subject;

[0010] (B) Tumor-derived cell-free DNA was labeled based on the terminal positions.

[0011] In some embodiments, step (B) comprises: classifying the terminal positions of the free DNA to be detected into one or more types selected from the following based on the nucleosome protection region:

[0012] a. The free DNA end is located in the nucleosome protection region;

[0013] b. There is no overlap between free DNA and nucleosome-protected regions;

[0014] c. Free DNA covers the nucleosome protection region.

[0015] In some embodiments, the methods of the present invention comprise labeling type a and / or type b cell-free DNA as tumor-derived cell-free DNA.

[0016] In some embodiments, the methods of the present invention further comprise the step of determining nucleosome protection regions in healthy subjects.

[0017] In some embodiments, determining nucleosome protection regions in healthy subjects comprises the following steps:

[0018] (1) Collect DNA from healthy subjects;

[0019] (2) constructing and sequencing a library of the healthy subject's DNA;

[0020] (3) Analyze the sequencing results and determine the nucleosome protection regions in healthy subjects.

[0021] In some embodiments, step (A) comprises the following steps:

[0022] (i) collecting cell-free DNA from the subject to be tested;

[0023] (ii) constructing a library and sequencing the free DNA;

[0024] (iii) Analyze the sequencing results and determine the terminal positions of the free DNA to be tested.

[0025] In some embodiments, the library construction method is a double-stranded DNA library construction method or a single-stranded DNA library construction method.

[0026] In some embodiments, sequencing is sequencing using a WGS method or sequencing using an MNase-seq method.

[0027] In some embodiments, analyzing sequencing results comprises the following steps:

[0028] (i) Sequencing data preprocessing;

[0029] (ii) sequence alignment and quality control;

[0030] (iii) Sequencing fragment screening and analysis.

[0031] In some embodiments, step (iii) comprises determining nucleosome protection regions using a WPS method, using a DANPOS2 method, or using a NucTools method.

[0032] In some embodiments, the DNA of a healthy subject is derived from cell-free DNA in the plasma of a healthy subject or DNA from normal cells of a healthy subject, preferably DNA from peripheral blood cells.

[0033] In some embodiments, the method of the present invention includes mixing DNA samples from multiple healthy subjects and then constructing and sequencing a library, mixing libraries of DNA samples from multiple healthy subjects and then sequencing them, or analyzing the sequencing results of DNA samples from multiple healthy subjects and merging the nucleosome protection regions obtained separately.

[0034] In some embodiments, the analysis results include determining the nucleosome protection region by extending a certain length to both sides with the coordinates of the nucleosome center point obtained by the analysis as the center, wherein the certain length is 80 bp, 75 bp, 70 bp, 65 bp or a length between any of the above values, preferably the certain length is 75 bp.

[0035] In some embodiments, the sequencing depth is between 0.1X and 10X.

[0036] In some embodiments, the method described herein further comprises analyzing total cell-free DNA, type a cell-free DNA, type b cell-free DNA, and / or type c cell-free DNA using ichorCNA to determine the content of tumor-derived cell-free DNA. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1: Whole-genome sequencing data analysis process, which includes preprocessing of raw fastq data, alignment and quality control, construction of nucleosome protection region baseline, and screening and classification of sample cfDNA fragments.

[0038] Figure 2: Estimation of ctDNA enrichment using filtered fragments. HCC: liver cancer sample, control: healthy sample. In / out / span represent the in_nuclMAP, out_nuclMAP, and span_nuclMAP fragment sets, respectively. The y-axis represents the ratio of ctDNA content estimated using the in / out / span fragment set to the ctDNA content estimated using the all_nuclMAP fragment set.

[0039] Figure 3: ROC curves for the tumor detection model on the DELFI data. Top: Average ROC curve of five cross-validation runs on the DELFI training set. Bottom: ROC curve on the DELFI validation set. "All" and "Aberrant" respectively indicate the use of all fragments and aberrant fragments (sequences whose ends are within or do not overlap with nucleosome-protected regions) in the sample. "EDM" indicates the use of the 5'-terminal four-base pair sequence features to construct the machine learning model.

[0040] Figure 4: ROC curves for the tumor detection model on independent cohort data. Top: Average ROC curve from five cross-validation runs on the independent cohort training set; Bottom: ROC curve on the independent cohort validation set. "All" and "Aberrant" indicate the use of all fragments and aberrant fragments (sequences whose ends are within or do not overlap with nucleosome-protected regions) in the sample, respectively. "EDM" indicates the use of the 5'-terminal four-base pair sequence features to construct the machine learning model. DETAILED DESCRIPTION

[0041] The inventors discovered that by pre-constructing a baseline of nucleosome protection regions based on a healthy human sample set, and delineating the regions where normal cell-free DNA fragments are likely to occur, a cell-free DNA fragment from a sample to be analyzed can be mapped to its location on the reference genome and compared with the baseline nucleosome protection region to preliminarily determine whether the fragment is more likely to originate from a normal cell or a tumor cell. For example, if the fragment completely spans the nearest baseline nucleosome protection region, then the fragment is more likely to originate from a normal cell. Conversely, if the fragment is fragmented within the nearest baseline nucleosome protection region or does not overlap with the baseline nucleosome protection region, then it is more likely to originate from a tumor cell. Generally, compared to healthy samples, the ends of cell-free DNA from tumor samples are more likely to be fragmented within the nearest baseline nucleosome protection region or not overlap with the baseline nucleosome protection region. Therefore, by marking these fragments, it is possible to enrich for tumor-derived cell-free DNA. However, for healthy samples, this method does not produce an enrichment effect for tumor-derived cell-free DNA.

[0042] Throughout this specification, reference to "one embodiment," "an embodiment," "some embodiments," "an embodiment," "a specific embodiment," "a related embodiment," "an embodiment," "an additional embodiment," or "a further embodiment," or combinations thereof, means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the aforementioned phrases in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0043] As used herein, the terms "optional" and "optionally" include both selection and non-selection. For example, "optionally modified" includes both modification and non-modification.

[0044] As used herein, the terms "a," "an," and "the" are generally construed to cover both the singular and the plural.

[0045] As used herein, the term "include" means and can be used interchangeably with the phrase "including (but not limited to)". The term "comprising" used herein means and can be used interchangeably with the phrase "including (but not limited to)". The technical solutions using "include" or "comprising" in this patent can be further limited to "consisting of" or "composed of".

[0046] As used herein, the term "healthy subject" refers to any mammal or non-mammal. Mammals include, but are not limited to, humans, vertebrates such as rodents, non-human primates, cattle, horses, dogs, cats, pigs, sheep, goats. Preferably, in some embodiments, the healthy subject is a human healthy subject.

[0047] Nucleosome protection region

[0048] The "nucleosome protection region" described herein refers to a region on a DNA sequence that is highly protected due to DNA wrapping around histones, which can be determined by methods known in the art, such as the WPS (windowed protection score) method (Snyder MW, Kircher M, Hill AJ, et al. Cell-free DNA comprises an in vivo nucleosome footprint that informs its tissues-of-origin. Cell. 2016; 164: 57–68), the DANPOS2 method (Chen K, Xi Y, Pan X, et al. DANPOS: dynamic analysis of nucleosome position and occupancy by sequencing. Genome Res. 2013; 23(2): 341-351), and the NucTools method (Vainshtein Y, Rippe K, Teif VB. NucTools: analysis of chromatin feature occupancy profiles from high-throughput sequencing data. BMC Genomics.2017;18(1):158) or other methods.

[0049] In some embodiments, the nucleosome protection region is determined using the WPS method. In some embodiments, the nucleosome protection region is determined using the DANPOS2 method. In some embodiments, the nucleosome protection region is determined using the NucTools method. In some embodiments, the nucleosome protection region is determined by extending a certain length to both sides with the nucleosome center point coordinates obtained by analysis as the center. In some embodiments, the certain length is 100bp, 95bp, 90bp, 85bp, 80bp, 75bp, 70bp, 65bp or a length between any of the above values, preferably the certain length is 75bp.

[0050] Free DNA

[0051] As used herein, "cell-free DNA" refers to any DNA obtained from a subject that is present outside of cells in the subject.

[0052] As described herein, the "subject to be tested" can be any subject for whom tumor risk assessment is required, and can be a healthy subject, a subject at risk of developing a tumor, or a tumor patient.

[0053] In some embodiments, free DNA is dispersed floating DNA obtained from an organism, generally referring to DNA present in body fluids, free outside cells or outside vesicles. Free DNA is well known in the art and generally refers to DNA that floats freely in body fluids.

[0054] In some embodiments, the cell-free DNA is derived from the blood or plasma of a subject. In some embodiments, the cell-free DNA of a healthy subject is derived from normal cells of a healthy subject, such as peripheral blood cells.

[0055] End position of free DNA

[0056] The inventors discovered that tumor-derived cell-free DNA can be labeled by its terminal positions. The cell-free DNA of the test subject can be classified into the following types based on the terminal positions of the nucleosome-protected regions:

[0057] a. The free DNA end is located in the nucleosome protection region;

[0058] b. There is no overlap between free DNA and nucleosome-protected regions;

[0059] c. Free DNA covers the nucleosome protection region.

[0060] It can be considered that the probability of free DNA that meets type a and type b being tumor-derived is higher than that of free DNA that meets type c.

[0061] The phrase “the free DNA does not overlap with the nucleosome protection region” mentioned herein means that the position of the free DNA on the genome does not overlap with the position of the nucleosome protection region on the genome.

[0062] Those skilled in the art can use conventional methods in the art to determine the position of the free DNA on the genome and its positional relationship with the nucleosome protection region. In some embodiments, the sequence alignment tool bwa-mem or bwa-mem2 is used to determine the position of the free DNA on the genome. In some embodiments, the bedtools method is used to determine the relative relationship between the position of the free DNA on the genome and the position of the nucleosome protection region on the genome.

[0063] In some embodiments, both terminal positions of the free DNA in type a are located in the nucleosome protection region.

[0064] In some embodiments, type a and / or type b cell-free DNA is labeled as tumor-derived cell-free DNA.

[0065] Library construction and sequencing

[0066] Those skilled in the art will appreciate that any suitable method in the art can be used for library construction and sequencing.

[0067] In some embodiments, the library construction method is a double-stranded DNA library construction method.

[0068] In some embodiments, the library construction method is a single-stranded DNA library construction method.

[0069] In some embodiments, the sequencing method is a WGS method.

[0070] In some embodiments, the sequencing method is a MNase-seq method.

[0071] sample

[0072] The DNA (such as free DNA) for sequencing and analysis of the present invention can be extracted from a sample of a subject. In some embodiments, the sample is a body fluid. In some embodiments, the sample is blood. In some embodiments, the sample is peripheral blood. In some embodiments, the sample is plasma, such as the plasma of peripheral blood. In some embodiments, the body fluid is selected from: at least one of blood, serum, gastric juice, intestinal fluid, saliva, bile, tumor fluid, cerebrospinal fluid, breast milk, semen, urine, vaginal fluid, interstitial fluid and feces. In some embodiments, the sample is from the subject. In some embodiments, the sample is from the body fluid of the subject, and more preferably the sample is from the blood (such as: peripheral blood), urine, cerebrospinal fluid, lung lavage fluid and lymph fluid of the subject.

[0073] In some embodiments, the volume of the sample used in the methods of the present invention is 0.005, 0.01, 0.05, 0.1, 0.5, 1.0, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 20, 30, 50 ml or any value therebetween. In some embodiments, the volume of the sample used in the methods of the present invention is at least 0.5 ml. In some embodiments, the volume of the sample used in the methods of the present invention is at most 10 ml.

[0074] Samples of the present invention can all be obtained by standard techniques in this area. For example, a blood sample can be obtained by standard techniques, such as using a needle and syringe. In some embodiments, the blood sample is a peripheral blood sample. Alternatively, the blood sample can be a separated portion of peripheral blood, such as a plasma sample. In some embodiments, from the sample, for example, in the nucleosomes in the sample, free DNA is extracted. In some embodiments, after obtaining the blood sample, total DNA can be extracted from the sample using standard techniques known to those skilled in the art. In some embodiments, intact cells are removed before DNA extraction, thereby only free-floating DNA is extracted. Intact cells can be removed by any method known in the art, such as, as a non-limiting example, by centrifugation or by gradient separation, such as by Ficol gradient separation. A non-limiting example of DNA extraction is the FlexiGene DNA test kit (QIAGEN). Standard techniques for receiving cell-free DNA extraction are known to the skilled person, and a non-limiting example thereof is the QIAamp Circulating Nucleic Acid test kit (QIAGEN).

[0075] Library construction and sequencing

[0076] In some embodiments, the amount of cell-free DNA used to establish the library is 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, or 500 ng, or any value therebetween. In some embodiments, the amount of cell-free DNA used to establish the library is 10 to 100 ng.

[0077] In some embodiments, the sequencing depth is 300X, 250X, 200X, 150X, 100X, 50X, 40X, 30X, 20X, 10X, 9X, 8X, 7X, 6X, 5X, 4X, 3X, 2X, 1X, 0.5X, 0.1X, or any depth therebetween. In some embodiments, the sequencing depth is 10X to 0.1X. In some embodiments, the sequencing depth is 0.1X.

[0078] In some embodiments, sequencing is next generation sequencing.Next generation sequencing, also referred to as high throughput sequencing or large-scale parallel sequencing, is any sequencing method that realizes that the base pairs from DNA or RNA samples are carried out to rapid high throughput sequencing. In some embodiments, sequencing is high throughput sequencing. In some embodiments, sequencing is large-scale parallel sequencing. Such sequencing is well known in the art, and can include using Illumina sequencer, MGI sequencer and nanopore sequencer and ion-based semiconductor sequencer (ion torrent) as non-limiting examples. For non-limiting examples, sequencing machines such as Illumina Nextseq 500 machines can be used, and Illumina 500 / 550V2 test kits can be used to process. In some embodiments, sequencing is whole genome sequencing. In some embodiments, only part of the genome is sequenced.

[0079] Application of the method of the present invention

[0080] In one aspect, the present invention provides uses of the methods of the present invention in one or more of the following aspects: early diagnosis or screening of cancer, evaluation of cancer treatment regimens, or detection of cancer progression.

[0081] The information generated by the method of the present invention is an intermediate result for diagnosis and cannot be directly used as a basis for determining whether a subject has cancer.

[0082] In some embodiments, the methods described herein can be used to assess a subject's cancer risk. In some embodiments, the methods described herein can be used to determine a subject's cancer progression or cancer stage.

[0083] In some embodiments, the methods of the present invention can be used before and after treatment of the same subject to help determine the status of the subject's cancer before and after treatment, thereby evaluating cancer treatment regimens.

[0084] In some embodiments, the method of the present invention can also be used continuously during the subject's cancer treatment to detect the subject's cancer progression.

[0085] In some implementations, the present methods can be used in subjects at high risk for developing cancer to help monitor whether such subjects have cancer.

[0086] In some embodiments, the disease status of a subject can be determined by comparing data obtained using the methods of the present invention with data from healthy subjects. In some embodiments, data from healthy subjects can also be obtained using the methods of the present invention. Analyzing data obtained using the methods of the present invention (e.g., comparing the data with data from healthy subjects) is within the capabilities of those skilled in the art, who can use any algorithm or other analytical method relevant in the art to perform the analysis and draw conclusions.

[0087] In some embodiments, machine learning methods can also be used to analyze the data of the subject obtained using the methods of the present invention to determine the disease status of the subject.

[0088] The following examples illustrate the beneficial effects of the present invention. Those skilled in the art will recognize that these examples are illustrative and non-restrictive. These examples do not limit the scope of the present invention in any way. The experimental methods described in the following examples are conventional methods unless otherwise noted; reagents and materials are commercially available unless otherwise noted.

[0089] The experimental techniques and methods used in this example are conventional unless otherwise specified. For example, in the following examples, experimental methods without specific conditions are generally performed under conventional conditions or according to the conditions recommended by the manufacturer. The materials and reagents used in the examples are all available through regular commercial channels unless otherwise specified.

[0090] Example

[0091] Example 1: Construction of a reference baseline of nucleosome protection regions in healthy individuals

[0092] 1.1 Plasma sample sequencing and basic analysis

[0093] First, plasma samples were collected from 10 healthy volunteers, cfDNA was extracted, and a standard double-stranded WGS library was constructed. The library was then sequenced on a NovaSeq 6000 sequencer, with an average sequencing depth of approximately 30X. The steps involved:

[0094] 1. cfDNA was extracted from plasma separated from 8-10 ml of whole blood using the QIAamp Circulating Nucleic Acid Kit (QIAGEN);

[0095] 2. Quantification was performed using the Qubit dsDNA HS assay (Thermo Fisher Scientific);

[0096] 3. Extract 5 ng of cfDNA and construct a double-stranded WGS library using the IDT xGen Prism DNA Library Prep Kit and refer to its protocol;

[0097] 4. 2x150bp sequencing was performed on the NovaSeq 6000 sequencer, with an average sequencing data volume of 100G base pairs;

[0098] 5. The raw signal data obtained from the sequencing machine is converted into fastq format using the bcl2fastq tool provided with the sequencer;

[0099] 6. Use fastp (https: / / github.com / OpenGene / fastp) to remove the adapters and single molecule identifiers (UMIs) introduced during library construction;

[0100] 7. Use bwa-mem2 (https: / / github.com / bwa-mem2 / bwa-mem2) to align the processed fastq data to the human reference genome hg19;

[0101] 8. Use sambamba (https: / / lomereiter.github.io / sambamba / ) to mark duplicate sequences. Here, duplicate sequences refer to two paired reads whose 5' ends are aligned at exactly the same position on the reference genome.

[0102] 1.2 Construction of a nucleosome protection region reference baseline

[0103] Using plasma cfDNA data obtained through WGS, nucleosome-protected regions on genes in normal cells can be inferred. The analysis strategy is based on the WPS method proposed by Snyder et al. (Snyder MW, Kircher M, Hill AJ, Daza RM, Shendure J. Cell-free DNA comprises an in vivo nucleosome footprint that informs its tissues-of-origin. Cell. 2016; 164:57–68). The specific steps include:

[0104] 1. Select cfDNA fragments located on autosomes 1-22 or XY sex chromosomes on the genome with a length of 120-180bp;

[0105] 2. The analysis region includes all genomic positions on each chromosome starting from 201 to the corresponding chromosome length minus 200;

[0106] 3. The sliding window length is set to 120 base pairs;

[0107] 4. Calculate the WPS score for each position on the chromosome region defined by REGION:

[0108] WPS = the number of cfDNA fragments spanning a 120-base-pair sliding window centered at that position minus the number of cfDNA fragments with fragment ends broken within that sliding window.

[0109] 5. Further, using a 1000 bp moving window, the median value of the WPS value obtained at each position in the genome was subtracted to obtain a new WPS value after local adjustment;

[0110] 6. Use Savitzky-Golay filtering (window size 21, second-order polynomial) for smoothing to remove potential noise;

[0111] 7. Then divide the WPS into areas above zero (allowing up to 5 consecutive locations below zero);

[0112] 8. Calculate the median WPS value in this area and search for WPS values ​​above the median and the largest continuous window;

[0113] 9. Record the start and end positions of this window and define the center point as the "peak" or local maximum of the nucleosome protection region;

[0114] 10. The score of the peak is calculated by the difference between the maximum WPS value in the window and the average of the two minimum WPS values ​​in the adjacent area.

[0115] To obtain conserved and consistent nucleosome protection regions in normal cells, the sequencing data of 10 healthy human plasma samples were subjected to the basic analysis described above (as shown in Figure 1). These BAMs were merged into one file using samtools (https: / / github.com / samtools / samtools), and the above-mentioned nucleosome protection region analysis method was used to obtain the nucleosome protection intervals on the genomes of healthy human samples, which served as the reference baseline for the nucleosome protection regions of healthy humans.

[0116] Example 2: Classification of cfDNA fragments in plasma based on nucleosome protection region baseline

[0117] Since the nucleosome protection reference region used is constructed based on healthy human samples, cfDNA fragments spanning a complete nucleosome protection zone are more likely to be protected by conserved nucleosomes on the normal cell genome and are therefore more likely to originate from normal cells. cfDNA whose ends are located in or do not overlap with the protection zone is more likely to originate from tumor cells. Therefore, all cfDNA fragments sequenced in the sample can be divided into four categories:

[0118] a. Free DNA fragments completely cover the nucleosome protection region: span_nuclMAP

[0119] b. The ends of the free DNA fragments are located in the nucleosome protection region: in_nuclMAP

[0120] c. Free DNA fragments do not overlap with nucleosome protected regions: out_nuclMAP

[0121] d. All free DNA fragments: all_nuclMAP

[0122] In actual analysis, only autosomal fragments were considered, not chromosomes X and Y. Classification of cfDNA fragments was performed using bedtools by comparing the fragment positions on the reference genome with the nucleosome-protected region baseline (https: / / bedtools.readthedocs.io / en / latest / ). The complete steps include (as shown in Figure 1):

[0123] 1. Extract cfDNA from plasma, construct a standard WGS library, and sequence it on the NovaSeq 6000 sequencer. The average sequencing data volume of each sample is 10G base pairs.

[0124] 2. The raw signal data obtained from the sequencing machine is converted into fastq format using the bcl2fastq tool provided with the sequencer;

[0125] 3. Use fastp (https: / / github.com / OpenGene / fastp) to remove the single-molecule sequence tags introduced during the adapter and library construction process;

[0126] 4. Use bwa-mem2 (https: / / github.com / bwa-mem2 / bwa-mem2) to align the processed fastq data to the human reference genome hg19;

[0127] 5. Use sambamba (https: / / lomereiter.github.io / sambamba / ) to mark duplicate sequences. Here, duplicate sequences refer to two paired reads whose 5' ends are aligned at exactly the same position on the reference genome.

[0128] 6. Remove these repeated sequences, sequences that were not successfully aligned or correctly paired, sequences with poor alignment quality, and fragments that intersect with genomic blacklist regions to generate bed format results. Genomic blacklist regions are regions on the genome that are known to cause bias in sequencing data alignment results. Here we use the ENCODE blacklist region as a reference (https: / / github.com / Boyle-Lab / Blacklist)

[0129] 7. Using bedtools, the cfDNA fragments in the bed format of the sample were aligned with the reference baseline of the nucleosome protection region in healthy individuals. Based on the comparison results, the fragments were divided into four categories: span_nuclMAP, in_nuclMAP, out_nuclMAP, and all_nuclMAP.

[0130] Example 3: Evaluation and comparison of sample ctDNA content

[0131] The estimation of ctDNA content in tumor samples is completed using ichorCNA (Adalsteinsson VA, Ha G, Freeman SS, et al. Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat Commun. 2017 Nov 6; 8(1): 1324). ichorCNA estimates the content of tumor ctDNA mainly by analyzing the copy number variation status of fragments on the tumor genome. In simple terms, the whole genome is first divided into multiple non-overlapping windows with a window size of 1Mb. The number of cfDNA fragments aligned to these window intervals on the genome is counted, and the GC content and sequence alignment bias are corrected. Finally, the hidden Markov model is used to simultaneously estimate the results such as the copy number variation of the genomic fragments and the content of tumor ctDNA.

[0132] 3.1 ichorCNA copy number baseline construction

[0133] 1. Collect plasma samples from 20 healthy volunteers, extract cfDNA, construct standard double-stranded WGS libraries, and sequence them on a NovaSeq 6000 sequencer at an average sequencing depth of 30X. Refer to Figure 1 for basic analysis of the off-machine data to generate filtered cfDNA fragment result files.

[0134] 2. Based on the previously constructed nucleosome protection region reference baseline, the cfDNA fragments in these 20 plasma samples were classified into span_nuclMAP, in_nuclMAP, out_nuclMAP fragment sets, and all_nuclMAP fragment set;

[0135] 3. For each sample, randomly downsample 2,000,000 fragments from the four types of fragments.

[0136] The window size was set to 1 Mb, and ichorCNA was used to construct the genomic copy number baselines of span_nuclMAP, in_nuclMAP, out_nuclMAP, and all_nuclMAP using the above 20 samples;

[0137] 3.2 Sample ctDNA content assessment

[0138] For the samples to be analyzed, plasma cfDNA was extracted and subjected to low-depth WGS sequencing (sequencing data volume of approximately 10G). After basic analysis using the method shown in Figure 1, the cfDNA fragments contained therein were classified based on the constructed nucleosome protection region reference baseline into span_nuclMAP, in_nuclMAP, out_nuclMAP fragment sets, and all_nuclMAP fragment sets. To control the impact of the number of cfDNA fragments on the estimation of ctDNA content, the span_nuclMAP, in_nuclMAP, out_nuclMAP, and all_nuclMAP fragment sets for each sample were downsampled to 2,000,000 fragments (approximately corresponding to an average genome sequencing depth of 0.1x. At this sequencing depth, the detection limit of ctDNA content by ichorCNA analysis is 3%). Using ichorCNA, using the corresponding copy number baseline, the ctDNA content of the samples under the three types of cfDNA fragment sets, namely span_nuclMAP, in_nuclMAP, and out_nuclMAP, was estimated. The original ctDNA content of the sample was downsampled to 2,000,000 fragments from the all_nuclMAP fragment set of the entire cfDNA and estimated by ichorCNA based on the corresponding all_nuclMAP copy number baseline.

[0139] 3.3 Evaluation of ctDNA Enrichment Effect

[0140] By comparing the sample ctDNA content with the original ctDNA content under the three types of cfDNA fragment sets, namely span_nuclMAP, in_nuclMAP and out_nuclMAP, the degree of ctDNA enrichment under the fragment screening method based on the reference nucleosome protection region baseline can be evaluated:

[0141] Example 4: Screening plasma for cfDNA fragments with terminal breaks in nucleosome protection regions or no overlap with baseline nucleosome protection regions to enrich ctDNA

[0142] Five hepatocellular carcinoma (HCC) samples and five healthy controls (control) samples were selected for low-depth WGS sequencing, with an average sample sequencing volume of approximately 10G. After data analysis as shown in Figure 1, based on the constructed nucleosome protection region reference baseline, the cfDNA fragments contained therein were divided into span_nuclMAP, in_nuclMAP, out_nuclMAP, and all_nuclMAP fragment sets, and downsampled to obtain 2,000,000 fragments. Using ichorCNA and the corresponding copy number baseline, the ctDNA content of these 10 samples under the four fragment sets was estimated. Among them, based on the all_nuclMAP fragment set, the original ctDNA content of the five HCC samples was 6.2%, 8.8%, 19.1%, 22.2%, and 38.7%, respectively, while the original ctDNA content of the five healthy control samples was 0.6%, 1.6%, 0.8%, 1.1%, and 2%, respectively (as shown in Table 1).

[0143] Table 1. Estimated ctDNA content and enrichment ratio in five HCC and five healthy subjects samples *: TF stands for tumor fraction, which indicates the ctDNA content. "all" indicates that the ctDNA content is estimated using the all_nuclMAP fragment set. Other columns are similar.

[0144] Comparative analysis revealed that compared to the original full-fragment data, the estimated ctDNA content for cfDNA fragments whose ends, after screening for nucleosome-protected regions, fell within or did not overlap with the nucleosome-protected regions (i.e., the reference genomic position to which the fragments were aligned did not overlap with the nucleosome-protected regions) increased in all five HCC samples. This indicates that the proportion of ctDNA derived from tumor cells is higher in the in_nuclMAP and out_nuclMAP cfDNA fragments, resulting in a corresponding enrichment effect. The proportion of ctDNA enriched ranged from 5% to 22% (as shown in Figure 2 and Table 1). However, the estimated ctDNA content for span_nuclMAP cfDNA fragments, which span the nearest nucleosome-protected regions, did not increase but instead decreased, indicating that the proportion of cfDNA derived from normal cells is higher in these fragments. Correspondingly, in healthy human samples, the three types of cfDNA fragment sets screened failed to detect higher ctDNA contents, indicating that there was no relevant ctDNA enrichment effect in healthy human samples based on the screening of nucleosome protection region baselines.

[0145] Example 5: Enriching ctDNA by baseline in nucleosome-protected regions in the DELFI dataset can improve tumor detection sensitivity

[0146] A DELFI sample bed format dataset (Cristiano S, Leal A, Phallen J, et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature. 2019 Jun; 570(7761): 385-389) with low-depth WGS sequencing (average sequencing depth of approximately 2x) was downloaded from the FinaleDB database (http: / / finaledb.research.cchmc.org / ). A machine learning sample dataset was selected, including 204 patients with different cancer stages, such as breast cancer, colorectal cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, and bile duct cancer, as well as 213 healthy subjects. The dataset was divided into a training set (containing 142 cancer samples and 149 healthy subjects) and a validation set (containing 62 cancer samples and 64 tumor samples) in a ratio of 7:3.

[0147] First, as shown in Figure 1, the in / out_nuclMAP fragment sets of the cfDNA fragments in the sample bed format file are extracted based on the nucleosome protection region reference baseline and merged as the fragments with abnormal end break positions (abnormal fragment set).

[0148] Next, using the four-base-pair sequence at the 5' end of the fragment (endpoint motif, EDM), the frequency of this four-base sequence is calculated for all cfDNA fragments or abnormal fragment sets in the DELFI sample bed format file (each base position has four possible bases, so there are 4^4 = 256 possible sequences), and a 256-dimensional feature vector is defined. Here, the four-base-pair sequence at the 5' end of the fragment refers to the four consecutive bases in the 5' to 3' direction on the genome where the 5' end of the cfDNA fragment is located after alignment to the reference genome.

[0149] The machine learning detection model was constructed using a machine learning framework based on Python sklearn. The machine learning algorithms implemented in this framework include xgboost, random forest, support vector machine, and logistic regression. The key steps are: 1) Perform a 5-fold cross-validation split on the training set and repeat it 5 times; 2) Use the boruta feature screening method based on random forest to screen sample features; 3) Use the Bayesian method hyperopt (https: / / hyperopt.github.io / hyperopt / ) to optimize the hyperparameters of the machine learning algorithm; 4) Determine the best machine learning algorithm by comparing the average AUC (Area Under Curve, defined as the area under the ROC curve and the coordinate axis) value of 5 times of 5-fold cross-validation training on the training set; 5) Use the entire training set to fit the best algorithm to construct a tumor detection model; 6) Apply the model to an independent validation set to predict whether the sample is tumor or non-tumor, and further verify the performance of the model, including AUC, sensitivity, and specificity: Sensitivity = true positive / (true positive + false negative) Specificity = true negative / (true negative + false positive)

[0150] Sensitivity refers to the proportion of samples that are actually positive (tumor samples) that are judged as positive, while specificity refers to the proportion of samples that are actually negative (healthy control samples) that are judged as negative.

[0151] As shown in Figure 3, the tumor detection models constructed using both the full and abnormal fragment sets on the DELFI dataset achieved an average AUC of 0.98 across five cross-validations of the training set. Further observation reveals that, at high specificity levels, the abnormal fragment set exhibits higher detection sensitivity than the full fragment set. A similar phenomenon was observed on the DELFI validation set.

[0152] To fully compare the sensitivity of the two detection models at a given specificity level, we first determined prediction score thresholds for healthy control samples at specific levels (e.g., 95%, 98%, and 99%) based on the prediction scores obtained in five cross-validation runs. These thresholds indicate that 95%, 98%, and 99% of the prediction scores for healthy control samples were below the threshold. The model was then applied to the validation set to obtain prediction scores for independent test samples. Samples with scores greater than or equal to the threshold were classified as tumorous; otherwise, they were classified as non-tumorous. As shown in Table 2, on the DELFI validation set, the model constructed using the aberrant segment set achieved sensitivities of 88.7% and 85.5% at specificity levels of 98% and 99%, respectively, significantly higher than the 32.3% and 32.3% achieved using the full segment set. Specific to different tumor types and stages, the sensitivity of the model constructed using the aberrant segment set was significantly higher than that achieved using the full segment set at both specificity levels of 98% and 99%.

[0153] Table 2. Tumor detection model performance on the DELFI data validation set *: All / Aberrant indicates the use of all fragment sets and aberrant fragment sets (sequences whose ends are within or do not overlap with nucleosome-protected regions) in the sample, respectively. EDM indicates the use of the 5'-terminal four-base pair sequence features to construct the machine learning model. **: The numerator in parentheses indicates the number of samples correctly detected at the specified specificity, while the denominator indicates the total number of samples tested.

[0154] Example 6: In another independent cohort, baseline enrichment of ctDNA by nucleosome-protected regions can improve the sensitivity and accuracy of tumor detection

[0155] To further validate the performance of screening and enriching ctDNA fragments based on a nucleosome-protected region reference baseline, two tumor detection models were constructed using an independent cohort, using the same processing methods as the DELFI dataset. This independent cohort consisted of 519 training samples and 226 validation samples. The training cohort included 84 colorectal cancer (CRC), 87 hepatocellular carcinoma (HCC), 53 lung adenocarcinoma (LUAD), 26 lung squamous cell carcinoma (LUSC), and 269 healthy controls; the validation cohort included 53 CRC, 41 HCC, 27 LUAD, 12 LUSC, and 93 healthy controls. First, 5 ng of cfDNA was extracted from all plasma samples in this cohort. Double-stranded WGS libraries were constructed and subjected to low-depth whole-genome sequencing at 2x150 bp on a NovaSeq 6000 sequencer, resulting in an average sequencing data size of approximately 10 Gbp. After analysis using the workflow shown in Figure 1, the in / out_nuclMAP fragment set was obtained. Similar to the DELFI dataset, these two datasets were merged and used as the fragments with abnormal end break positions (abnormal fragment set) for subsequent modeling analysis.

[0156] As shown in Figure 4, the models constructed using all cfDNA fragments and those constructed using the screened abnormal cfDNA fragments showed similar performance on the training set, with an AUC of approximately 0.9. However, on the independent validation set, the use of abnormal cfDNA fragments screened based on nucleosome-protected regions significantly improved the final AUC (0.91 vs. 0.87). Furthermore, it can be seen that the performance of the model using the cfDNA fragments screened based on nucleosome-protected regions on the training and validation sets is very similar, suggesting that this model has better generalization than the model constructed using all cfDNA fragments.

[0157] Comparing sensitivities at 95%, 98%, and 99% specificity levels, we can see that the sensitivities using the abnormal fragment set were 53.4%, 36.1%, and 21.8%, respectively, compared to 47.4%, 32.3%, and 16.5% using the full fragment set. Across different tumor types and stages, the abnormal fragment set achieved higher sensitivity than the full fragment set in most cases. Combined with its higher AUC value, these results suggest that pre-screening sample cfDNA fragments based on a reference baseline of nucleosome-protected regions can enrich ctDNA, translating into better detection sensitivity and accuracy.

[0158] Table 3. Tumor detection model performance on the independent cohort data validation set *: All / Aberrant indicates the use of all fragment sets and aberrant fragment sets (sequences whose ends are within or do not overlap with nucleosome-protected regions) in the sample, respectively. EDM indicates the use of the 5'-terminal four-base pair sequence features to construct the machine learning model. **: The numerator in parentheses indicates the number of samples correctly detected at the specified specificity, while the denominator indicates the total number of samples tested.

Claims

1. A method for labeling cell-free DNA derived from tumor cells, comprising the following steps: (A) Determining the end positions of the cell-free DNA of a subject to be tested based on the nucleosome protection regions of healthy subjects; (B) Labeling the cell-free DNA derived from tumors based on the end positions.

2. The method according to claim 1, wherein step (B) comprises: Based on the nucleosome protection regions, classifying the end positions of the cell-free DNA to be tested into one or more types selected from the following: a. The end position of the cell-free DNA is located within the nucleosome protection region; b. The cell-free DNA does not overlap with the nucleosome protection region; c. The cell-free DNA covers the nucleosome protection region.

3. The method according to claim 2, comprising labeling the cell-free DNA of type a and / or type b as the cell-free DNA derived from tumors.

4. The method according to any one of claims 1-3, further comprising the step of determining the nucleosome protection regions of healthy subjects.

5. The method according to claim 4, wherein the determining the nucleosome protection regions of healthy subjects comprises the following steps: (1) Collecting the DNA of a healthy subject; (2) Constructing a library and sequencing the DNA of the healthy subject; (3) Analyzing the sequencing results and determining the nucleosome protection regions of the healthy subject.

6. The method according to any one of claims 1-5, wherein step (A) comprises the following steps: (i) Collecting the cell-free DNA of the subject to be tested; (ii) Constructing a library and sequencing the cell-free DNA; (iii) Analyzing the sequencing results and determining the end positions of the cell-free DNA to be tested.

7. The method according to claim 5 or 6, wherein the method for library construction is a double-stranded DNA library construction method or a single-stranded DNA library construction method.

8. The method according to any one of claims 5-7, wherein the sequencing is performed using the WGS method or the MNase-seq method.

9. The method according to any one of claims 5-8, wherein the analyzing the sequencing results comprises the following steps: (i) Preprocessing the sequencing data; (ii) Sequence alignment and quality control; (iii) Screening and analyzing the sequencing fragments.

10. The method according to claim 9, wherein step (iii) comprises using the WPS method, using the DANPOS2 method, or using the NucTools method to determine the nucleosome protection regions.

11. The method according to claim 5, wherein the DNA of the healthy subject is derived from the cell-free DNA in the plasma of the healthy subject or the DNA of the cells of the healthy subject, preferably the DNA of peripheral blood cells.

12. The method according to claim 5, comprising mixing the DNA samples of multiple healthy subjects, constructing a library and sequencing, mixing the libraries of the DNA samples of multiple healthy subjects and then sequencing, or analyzing the sequencing results of the DNA samples of multiple healthy subjects and combining the respectively obtained nucleosome protection regions.

13. The method according to claim 5, wherein the analysis result includes determining a nucleosome protection region by extending a certain length to both sides centered on the coordinates of the nucleosome center point obtained by analysis, and the certain length is 80 bp, 75 bp, 70 bp, 65 bp or a length between any of the above values, and preferably the certain length is 75 bp.

14. The method according to any one of the preceding claims, wherein the sequencing depth is between 0.1X and 10X.

15. The method according to any one of the preceding claims, further comprising using ichorCNA to analyze all cell-free DNA, type a cell-free DNA, type b cell-free DNA, and / or type c cell-free DNA to determine the content of tumor-derived cell-free DNA.

Citation Information

Patent Citations

  • Mutation analysis method and device for cell-free DNA (cfDNA) sequencing data

    CN112111565A

  • Method and product for predicting curative effect of immune checkpoint inhibitor combined targeting therapy of hepatobiliary tumor patient

    CN112961920A

  • Quantitative detection method for methylation of free DNA in blood and application of quantitative detection method

    CN114703284A

  • Cancer diagnosis model based on free DNA and application

    CN115019952A

  • Method for marking free DNA (deoxyribonucleic acid) derived from tumor cells based on nucleosome protection region

    CN117887810A