Analysis of nucleic acids in dried blood
Patent Information
- Application Number
- PCT/US2025/017909
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-03
- Filing Date
- 2025-02-28
- Publication Date
- 2025-10-02
AI Technical Summary
Current multi-cancer early detection technologies using genomics-based liquid biopsies are costly and require phlebotomy, making them inaccessible in low-resourced and rural settings, and there is a need for methods that allow for at-home or distributed sample collection that preserve diagnostic characteristics of nucleic acids in blood samples.
The method involves analyzing nucleic acids extracted from solid substrates, such as dried blood samples, using a variety of materials including nylon, polypropylene, and glass microfiber filters, which can absorb blood samples, and extracting nucleic acids using solvents to preserve diagnostic information for cancer detection.
This method allows for the preservation of diagnostic characteristics of nucleic acids in dried blood samples, enabling effective cancer detection through fragmentation pattern analysis, even in low-resource settings, without the need for large blood volumes or controlled storage conditions.
Abstract
Description
[0001] ANALYSIS OF NUCLEIC ACIDS IN DRIED BLOOD
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] Priority is hereby claimed to US Provisional Application 63 / 560,703, filed March 3, 2024, which is incorporated herein by reference in its entirety.
[0004] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0005] This invention was made with government support under CA243078 awarded by the National Institutes of Health. The government has certain rights in the invention.
[0006] FIELD OF THE INVENTION
[0007] The invention is directed to the analysis of nucleic acids in dried blood, such as the detection of fragmentation patterns and / or sequence characteristics of cfDNA for detecting cancer.
[0008] BACKGROUND
[0009] Recent advances in multi-cancer early detection (MCED) using genomics-based liquid biopsies have yielded promising results. However, most current technologies remain costly and require access to phlebotomy to collect 1-2 blood tubes. For early detection or screening, these factors make MCED less accessible in low-resourced and rural settings, and can reduce compliance with early detection tests. As telehealth continues to develop and is implemented, there is a growing need to develop technologies that provide at home or distributed sample / data collection.
[0010] Blood sampling technologies that avoid the need for large, liquid blood volumes, are easy to obtain, can be stored and shipped under ambient conditions, and still preserve the diagnostic characteristics of the nucleic acids contained in the blood are needed.
[0011] SUMMARY OF THE INVENTION
[0012] One aspect of the invention is directed to methods of analyzing nucleic acids extracted from solid substrates. In some versions, the methods comprise providing a solid substrate comprising a blood sample comprising nucleic acid, extracting the nucleic acid from the solid substrate to obtain extracted nucleic acid, and analyzing the extracted nucleic acid.
[0013] In some versions, the blood sample is a dried blood sample.
[0014] In some versions, the blood sample is a cell-removed blood sample. In some versions, the cell-removed blood sample comprises at least one of dried plasma and dried serum.
[0015] In some versions, the methods comprise obtaining a whole blood sample from a subject, removing cells from the whole blood to generate a liquid cell-removed blood sample, applying either the whole blood or the liquid cell-removed blood sample to the solid substrate, and then drying the liquid cell-removed blood sample to obtain the cell-removed blood sample.
[0016] In some versions, the whole blood or the liquid cell-removed blood sample is applied to the solid substrate in an amount from 10 pl to 600 pl.
[0017] In some versions, the solid substrate comprises the liquid cell-removed blood sample, prior to the drying, in an amount from 10 pl to 600 pl.
[0018] In some versions, the extracted nucleic acid comprises cell-free DNA.
[0019] In some versions, the extracted nucleic acid is substantially devoid of intact genomic DNA.
[0020] In some versions, the solid substrate comprises one or more of nylon, polypropylene, polyester, rayon, cellulose, cellulose acetate, nitrocellulose, mixed cellulose ester, glass microfiber filters, cotton, quartz microfiber, polytetrafluoroethylene, wax-patterned cellulose, polyethersulfone, and polyvinylidene fluoride.
[0021] In some versions, extracting the nucleic acid from the solid substrate comprises combining the solid substrate with a liquid solvent, incubating the solid substrate in the liquid solvent to extract the nucleic acid from the solid substrate into the liquid solvent and thereby obtain the extracted nucleic acid in the liquid solvent, and removing the liquid solvent from the solid substrate.
[0022] In some versions, the extracted nucleic acid is in an amount from 10 pg to 25 ng.
[0023] In some versions, the extracted nucleic acid has a modal size from 160 to 174 bp.
[0024] In some versions, analyzing the extracted nucleic acid comprises analyzing a sequence of the extracted nucleic acid.
[0025] In some versions, analyzing the extracted nucleic acid comprises determining a fragmentation pattern of the extracted nucleic acid.
[0026] In some versions, determining the fragmentation pattern of the extracted nucleic acid comprises non-targeted sequencing of the nucleic acid.
[0027] In some versions, determining the fragmentation pattern of the extracted nucleic acid comprises whole-genome sequencing.
[0028] In some versions, determining the fragmentation pattern of the extracted nucleic acid comprises genome -wide fragmentation analysis. In some versions, determining the fragmentation pattern of the extracted nucleic acid comprises targeted sequencing of the nucleic acid.
[0029] In some versions, determining the fragmentation pattern of the extracted nucleic acid comprises targeted amplification of the nucleic acid.
[0030] In some versions, determining the fragmentation pattern of the extracted nucleic acid comprises targeted amplification and sequencing of the nucleic acid. In some versions, determining the fragmentation pattern of the extracted nucleic acid comprises determining a fragment size distribution, information-weighted fraction of aberrant fragments analysis, determining incorporation of nucleotide frequencies at fragment ends, determining incorporation of nucleotide frequencies surrounding fragment ends, determining ratios of fragment sizes across one or more portions of a genome, motif frequencies at fragment ends, among others.
[0031] In some versions, the blood sample is from a subject suspected of having cancer.
[0032] In some versions, determining the fragmentation pattern of the extracted nucleic acid detects an aberrant fragmentation pattern relative to control nucleic acid from a subject that does not have cancer. In some versions, the fragmentation pattern is determined using a machine learning classifier. See, e.g., US 2024 / 0209455 and Budhraja et al. 2023 (Budhraja KK, McDonald BR, Stephens MD, et al. Genome-wide analysis of aberrant position and sequence of plasma DNA fragment ends in patients with cancer. Science Translational Medicine 2023;15(678):eabm6863), which are incorporated herein by reference in their entireties.
[0033] The objects and advantages of the invention will appear more fully from the following detailed description of the preferred embodiment of the invention made in conjunction with the accompanying drawings.
[0034] BRIEF DESCRIPTION OF THE DRAWINGS
[0035] FIG. 1. Example overview. We collected matched plasma samples from blood tubes, plasma-separating dried blood spots, and conventional dried blood spots from 75 individuals. Samples from 45 healthy individuals, 25 patients with cancer, and 5 patients with non- malignant conditions including pancreatic cysts and pelvic mass were included in this study (Table 2). After DNA extraction from each sample type, we performed whole genome sequencing and measured multiple features related to plasma DNA fragmentation. We compared these features between paired samples to assess similarity and correlation. We also compared performance of machine learning classifiers built to distinguish cancer from healthy samples, across paired sample sets. Image is adapted from original image Created in BioRender. Murtaza, M. (2025) BioRender.com v97I923.
[0036] FIG. 2. Electropherogram of DNA isolated from representative samples of plasma from blood tubes, psDBS, and eDBS. psDBS and eDBS samples chosen here showed the highest concentration amongst this sample set.
[0037] FIG. 3. Electropherogram of whole genome libraries generated from representative samples of plasma from blood tubes, psDBS, and eDBS. Plasma library was diluted 1 : 10 prior to analysis.
[0038] FIGS. 4A-4F. Comparison of plasma DNA whole genome sequencing libraries prepared from blood tubes, plasma-separating dried blood spots and conventional dried blood spots. FIG. 4A. Comparison of plasma DNA concentration in blood tubes with library yield from matched psDBS (one outlier is excluded in this analysis, with a plasma concentration of 127.2 ng / ml and a psDBS library yield of 450 ng). FIG. 4B. Comparison of plasma DNA concentration in blood tubes with library yield from matched eDBS (one outlier is excluded in this analysis, with a plasma concentration of 127.2 ng / ml and eDBS library yield of 87.6 ng). FIG. 4C. Average fragment size distribution observed in plasma DNA whole genome libraries from blood tubes, psDBS, and eDBS from healthy individuals. FIG. 4D. Fragment size distribution of plasma DNA whole genome libraries from blood tubes from healthy individuals. Each individual sample is one grey line, with the average plotted with the thick black line. FIG. 4E. Fragment size distribution of whole genome libraries from psDBS samples from healthy individuals. Each individual sample is one grey line, with the average plotted with the thick black line. FIG. 4F. Fragment size distribution of whole genome libraries from eDBS samples from healthy individuals. Each individual sample is one grey line, with the average plotted with the thick black line.
[0039] FIGS. 5A-5J. Comparison of plasma DNA fragmentation features between matched blood tubes, plasma-separating dried blood spots and conventional dried blood spots. FIG. 5A. Correlation in aberrant fragmentation score (AFS) between matched blood tubes and eDBS samples. FIG. 5B. Correlation in AFS between matched blood tubes and psDBS samples. The color scale indicates proportion of fragments in the psDBS sample of less than or equal to 120 bp. FIG. 5C. Correlation in AFS between matched blood tubes and psDBS samples, excluding samples with short fragment proportion of greater than 20%. FIG. 5D. Boxplot of AFS in healthy samples, excluding matched samples with a psDBS short fragment proportion of greater than 20%. FIG. 5E. Principal component analysis for nine selected single nucleotide frequencies from 5’ fragment ends, previously described in a cancer detection machine learning model. FIG. 5F. Box plot comparing principal component 1 for single nucleotide frequencies from 5’ fragment ends across all three sample types. FIG. 5G. Principal component analysis for dinucleotide frequencies from 5’ fragment ends. FIG. 5H. Box plot comparing principal component 1 for dinucleotide frequencies from 5 ’ fragment ends across all three sample types. FIG. 51. Principal component analysis for the ratio of short to long fragments in 544 bins of 5000 kbp each from across the genome. FIG. 5J. Box plot comparing principal component 1 for ratio of short to long fragments across all three sample types.
[0040] FIGS. 6A-6C. Correlation in individual nucleotide frequencies surrounding the 5’ fragment end between plasma-separating DBS samples (y-axis) and plasma samples (x-axis). Each plot represents frequency of a single nucleotide at that position relative to the 5’ end. OFOl represents the first position outside the fragment. IF02 represents the second position inside the fragment. IF03 represents the third position inside the fragment.
[0041] FIGS. 7A-7C. Correlation in individual nucleotide frequencies surrounding the 5’ fragment end between conventional DBS samples (y-axis) and plasma samples (x-axis). Each plot represents frequency of a single nucleotide at that position relative to the 5 ’ end. OFO 1 represents the first position outside the fragment. IF02 represents the second position inside the fragment. IF03 represents the third position inside the fragment.
[0042] FIGS. 8A-8B. Comparison of 4-mer frequencies observed at 5’ ends of cell-free DNA fragments. FIG. 8 A. Principal component analysis for dinucleotide frequencies from 5’ fragment ends. FIG. 8B. Box plot comparing principal component 1 for dinucleotide frequencies from 5’ fragment ends across all three sample types.
[0043] FIGS. 9A-9D. Comparison of plasma DNA between healthy individuals and patients with cancer using blood tubes and psDBS. FIG. 9A. Comparison of copy-number aberrations and inferred tumor fraction in plasma DNA between matched blood tubes (top) and psDBS (bottom) from a patient with Stage IV pancreatic cancer. FIG. 9B. ROC curves evaluating accuracy of the genome-wide analysis of fragment ends (GALYFRE) classifier based on 10 fragmentomic features including aberrant fragmentation scores and frequencies of 9 single nucleotides at loci surrounding 5’ fragment ends, using blood tubes (solid) and psDBS (dashed). FIG. 9C. ROC curves evaluating accuracy of a random forest machine learning model based on 4-mer nucleotide motifs at 5 ’ fragment ends, using blood tubes (solid) and psDBS (dashed). FIG. 9D. ROC curves evaluating accuracy of a random forest machine learning model based on the ratio between short and long plasma DNA fragments in 544 bins of 5000 kbp each across the genome, using blood tubes (solid) and psDBS (dashed). FIG. 9E. ROC curves evaluating accuracy of an ensemble model that combines the GALYFRE classifier (FIG. 9B), and models based on nucleotide motifs (FIG. 9C) and short-to-long fragment ratios (FIG. 9D), using blood tubes (solid) and psDBS (dashed). FIG. 10. Comparison of copy-number aberrations and inferred tumor fraction between matched plasma (top) and psDBS (bottom) samples from a patient with Stage IV breast cancer.
[0044] FIGS. 11A-11C. Comparison of fragment size distribution of sequenced DNA. FIG. 11A. Plasma from blood tubes. FIG. 11B. ADX100 device. FIG. 11C. HemaSep™ Punch device. FIG. HD. HemaSep™ Strip device. FIG. HE. Telimmune™ device.
[0045] FIGS. 12A-12C. Comparison of aberrant fragmentation score (AFS) between matched plasma samples and psDBS devices. FIG. 12A. Correlation in AFS between matched plasma and ADX100 devices. N = 10, r2= 0.93, p = 8.7e-06. FIG. 12B. Correlation in AFS between matched plasma and HemaSep™ Punch devices. N = 6, r2= 0.98, p = 2.0e-04. FIG. 12C. Correlation in AFS between matched plasma and HemaSep™ Strip devices. N = 5, r2= 0.99, p = 4.3e-04.
[0046] FIGS. 13A-13F. Ratio of short to long fragments in 5 megabase bins across the genome.
[0047] DETAILED DESCRIPTION OF THE INVENTION
[0048] One aspect of the invention is directed to providing a solid substrate comprising a blood sample.
[0049] The solid substrate can be any substrate capable of at least partially absorbing a volume of blood. In some versions, of the invention, the solid substrate comprises one or more of nylon, polypropylene, polyester, rayon, cellulose, cellulose acetate, nitrocellulose, mixed cellulose ester, glass microfiber filters, cotton, quartz microfiber, polytetrafluoroethylene, wax-patterned cellulose, polyethersulfone, and polyvinylidene fluoride. In some versions of the invention, the solid substrate comprises nylon, polypropylene, polyester, rayon, cellulose, cellulose acetate, nitrocellulose, mixed cellulose ester, glass microfiber filters, cotton, quartz microfiber, polytetrafluoroethylene, wax-patterned cellulose (Baillargeon et al. 2022), polyethersulfone (Morbioli et al. 2025), polyvinylidene fluoride, or any combination thereof in an amount of at least 10% w / w, at least 20% w / w, at least 30% w / w, at least 40% w / w, at least 50% w / w, at least 60% w / w, at least 70% w / w, at least 80% w / w, at least 90 w / w, or at least 99% w / w of the solid substrate. Exemplary solid substrates include Whatman® cellulose filter paper (Cytiva, Marlborough, MA), Telimmune™ DUO Plasma Separation Card (formerly Noviplex™ DUO Plasma Separation Card (Telimmune™ LLC, North Webster, IN), HemaSpot™ SE (SpotOnSciences, San Francisco, CA), AdvanceDx 100™ (ADX100) serum collection card (Advance Dx, Inc., Scottsdale, AZ), HemaSep™ Punch card (Ahlstrom, Helsinki, Finland), HemaSep™ Strip card (Ahlstrom, Helsinki, Finland), and substrates such as those described in US Patent 10,883,977, US Patent 8,062,608, among others. “Blood” as used herein refers to whole blood of processed forms thereof in which at least some of the components present in whole blood are removed. An exemplary form of processed whole blood includes cell-removed blood. “Cell-removed blood” is a form of blood from which at least some of the cells initially contained in the whole blood are removed. Exemplary forms of cell-removed blood include plasma and serum. “Plasma” as used herein is cell-removed blood that still includes the clotting factors present in whole blood. Plasma can be generated by centrifuging whole blood in the presence of anticoagulants or by separating cells from whole blood from the remaining components on a solid substrate, among other methods. “Serum” is cell-removed blood from which at least some of the clotting factors originally present in the whole blood have been removed. Serum can be generated by coagulating whole blood and removing the coagulant from the remaining portion, among other methods. The general terms “blood,” “cell-removed blood,” “plasma,” and “serum” refer to both liquid and dried forms. The dried forms include all the components of the liquid forms except that the majority of the water present in the liquid forms have been removed. The liquid forms of blood, cell-removed blood, plasma, and serum are referred to herein as “liquid blood,” “liquid cell-removed blood,” “liquid plasma,” and “liquid serum.” The dried forms of blood, cell-removed blood, plasma, and serum are referred to herein as “dried blood,” “dried cell-removed blood,” “dried plasma,” and “dried serum.”
[0050] “Sample” as used with reference to whole blood, cell-removed blood, plasma, or serum, whether in liquid or dried form, refers to a portion of whole blood, cell-removed blood, plasma, or serum.
[0051] In some versions, the solid substrate comprising the blood sample can be provided by obtaining a liquid blood sample from a subject and applying the liquid blood sample to the solid substrate. In some versions, the liquid blood sample is applied directly to the solid substrate without processing. In some versions, the liquid blood sample is processed after obtaining it from the subject and prior to applying it to the solid substrate. In some versions, the liquid blood sample is obtained from the subject in the form of liquid whole blood. In some versions, the liquid whole blood is processed to remove cells, such as red blood cells. In some versions, the liquid whole blood is processed to remove cells, such as red blood cells, prior to applying it to the solid substrate. In such cases, a liquid cell-removed blood sample is applied to the solid substrate. The blood cells in such situations can be removed by filtration, centriiugation, and / or coagulation, among other methods, as outlined above. In some versions, liquid whole blood is applied to the solid substrate, and the liquid whole blood is processed to remove cells, such as red blood cells, after applying it to the solid substrate. Solid substrates capable of removing cells, such as red blood cells, from liquid whole blood applied thereto are known in the art and include, without limitation, HemaSpot™ SE (SpotOnSciences, San Francisco, CA) and substrates such as those described in US Patent 10,883,977. After the liquid blood sample is applied to the solid substrate, the liquid blood sample can be dried to obtain a dried blood sample, such as dried whole blood, dried cell-removed blood, dried, plasma, or dried serum.
[0052] In various versions of the invention, the liquid blood sample is applied to the solid substrate in an amount of at least 0.1 pl, at least 0.5 pl, at least 1 pl, at least 5 pl, at least 10 pl, at least 15 pl, at least 20 pl, at least 25 pl, at least 30 pl, at least 35 pl, at least 40 pl, at least 45 pl, at least 50 pl, at least 75 pl, at least 100 pl, at least 125 pl, at least 150 pl, at least 175 pl, at least 200 pl, at least 225 pl, at least 250 pl, at least 275 pl, at least 300 pl, at least 325 pl, at least 350 pl, at least 375 pl, at least 400 pl, at least 425 pl, at least 450 pl, at least 475 pl, at least 500 pl, or more. In various versions of the invention, the liquid blood sample is applied to the solid substrate in an amount of up to 25 pl, up to 50 pl, up to 75 pl, up to 100 pl, up to 125 pl, up to 150 pl, up to 175 pl, up to 200 pl, up to 225 pl, up to 250 pl, up to 275 pl, up to 300 pl, up to 325 pl, up to 350 pl, up to 375 pl, up to 400 pl, up to 425 pl, up to 450 pl, up to 475 pl, up to 500 pl, up to up to 525 pl, up to 550 pl, up to 575 pl, up to 600 pl, up to up to 625 pl, up to 650 pl, up to 675 pl, up to 700 pl, up to up to 725 pl, up to 750 pl, up to 775 pl, up to 800 pl, up to up to 825 pl, up to 850 pl, up to 875 pl, up to 900 pl, up to up to 925 pl, up to 950 pl, up to 975 pl, up to 1000 pl or more. The form of the liquid blood sample in such cases can be liquid whole blood, liquid cell-removed blood, liquid plasma, or liquid serum.
[0053] In various versions of the invention, the solid substrate comprises the liquid blood sample, prior to drying, in an amount of at least 0.1 pl, at least 0.5 pl, at least 1 pl, at least 5 pl, at least 10 pl, at least 15 pl, at least 20 pl, at least 25 pl, at least 30 pl, at least 35 pl, at least 40 pl, at least 45 pl, at least 50 pl, at least 75 pl, at least 100 pl, at least 125 pl, at least
[0054] 150 pl, at least 175 pl, at least 200 pl, at least 225 pl, at least 250 pl, at least 275 pl, at least
[0055] 300 pl, at least 325 pl, at least 350 pl, at least 375 pl, at least 400 pl, at least 425 pl, at least
[0056] 450 pl, at least 475 pl, at least 500 pl, or more. In various versions of the invention, the solid substrate comprises the liquid blood sample, prior to drying, in an amount of up to 25 pl, up to 50 pl, up to 75 pl, up to 100 pl, up to 125 pl, up to 150 pl, up to 175 pl, up to 200 pl, up to 225 pl, up to 250 pl, up to 275 pl, up to 300 pl, up to 325 pl, up to 350 pl, up to 375 pl, up to 400 pl, up to 425 pl, up to 450 pl, up to 475 pl, up to 500 pl, up to up to 525 pl, up to 550 pl, up to 575 pl, up to 600 pl, up to up to 625 pl, up to 650 pl, up to 675 pl, up to 700 pl, up to up to 725 pl, up to 750 pl, up to 775 pl, up to 800 pl, up to up to 825 pl, up to 850 pl, up to 875 pl, up to 900 pl, up to up to 925 pl, up to 950 pl, up to 975 pl, up to 1000 pl or more. The form of the liquid blood sample in such cases can be liquid whole blood, liquid cell-removed blood, liquid plasma, or liquid serum.
[0057] For purposes of downstream analysis, the blood sample comprised by the solid substrate preferably comprises nucleic acid. The nucleic acid can be any type of nucleic acid, in any form, capable of yielding diagnostic information about an individual from which the blood sample was obtained. In some versions, the nucleic acid comprises DNA. In some versions, the nucleic acid comprises RNA. In some versions, the nucleic acid comprises intact genomic DNA. In some versions, the nucleic acid comprises fragmented genomic DNA. In some versions, the nucleic acid comprises cell-free DNA (cfDNA), In some versions, the nucleic acid comprises circulating tumor DNA (ctDNA). In some versions, the nucleic acid is substantially devoid of intact genomic DNA. “Substantially devoid” as used in the context of the nucleic acid being substantially devoid of intact genomic DNA means that the total nucleic acid comprised by the solid substrate is completely devoid of intact genomic DNA or contains intact genomic DNA in amount less than 5% w / w of the total nucleic acid comprised by the solid substrate.
[0058] Another aspect of the invention is directed to extracting the nucleic acid from the solid substrate to obtain extracted nucleic acid. As used herein, “extracting” used in the context of extracting the nucleic acid from the solid substrate means separating at least a portion of the nucleic acid from the solid substrate. In some versions, the nucleic acid can be extracted from the solid substrate by combining the solid substrate with a liquid solvent, incubating the solid substrate in the liquid solvent to extract the nucleic acid from the solid substrate into the liquid solvent and thereby obtain the extracted nucleic acid in the liquid solvent, and removing the liquid solvent from the solid substrate.
[0059] The liquid solvent used to extract the nucleic acid from the solid substrate can comprise any solvent suitable for solubilizing nucleic acid. In some versions, the liquid solvent is an aqueous solvent. “Aqueous solvent” refers to a solvent comprising at least 50% w / w water, such as 60% w / w water, 70% w / w water, 80% w / w water, 90% w / w water, 99% w / w water, or more.
[0060] In some versions, the liquid solvent can be buffered. The buffer can comprise a phosphate buffer such as Na2HPO4, KH2PO4, or any other buffer suitable for maintaining a pH effective to solubilize nucleic acid. In some versions, the liquid solvent can comprise a detergent. Exemplary detergents include sodium dodecyl sulfate, among others. In some versions, the liquid solvent can comprise a protease, such as proteinase K.
[0061] For purposes of downstream analysis, the extracted nucleic acid can be any type of nucleic acid, in any form, capable of yielding diagnostic information about an individual from which the blood sample was obtained. In some versions, the extracted nucleic acid comprises DNA. In some versions, the extracted nucleic acid comprises RNA. In some versions, the extracted nucleic acid comprises intact genomic DNA. In some versions, the extracted nucleic acid comprises fragmented genomic DNA. In some versions, the extracted nucleic acid comprises cell-free DNA (cfDNA), In some versions, the extracted nucleic acid comprises circulating tumor DNA (ctDNA). In some versions, the extracted nucleic acid is substantially devoid of intact genomic DNA. “Substantially devoid” as used in the context of the extracted nucleic acid being substantially devoid of intact genomic DNA means that the total extracted nucleic acid is completely devoid of intact genomic DNA or contains intact genomic DNA in amount less than 5% w / w of the total extracted nucleic acid present.
[0062] In various versions of the invention, the amount of nucleic acid extracted from the solid substrate is at least 1 pg, at least 5 pg, at least 10 pg, at least 25 pg, at least 50 pg, at least 100 pg, at least 150 pg, at least 200 pg, at least 250 pg, at least 300 pg, at least 350 pg, at least 400 pg, at least 450 pg, at least 500 pg, at least 550 pg, at least 600 pg, at least 650 pg, at least 700 pg, at least 750 pg, at least 800 pg, at least 850 pg, at least 900 pg, at least 950 pg, at least 1 ng, at least 50 ng, at least 100 ng, at least 150 ng, at least 200 ng, at least 250 ng, at least 300 ng, at least 350 ng, at least 400 ng, at least 450 ng, or at least 500 ng. In various versions of the invention, the amount of nucleic acid extracted from the solid substrate is up to 100 pg, up to 150 pg, up to 200 pg, up to 250 pg, up to 300 pg, up to 350 pg, up to 400 pg, up to 450 pg, up to 500 pg, up to 550 pg, up to 600 pg, up to 650 pg, up to 700 pg, up to 750 pg, up to 800 pg, up to 850 pg, up to 900 pg, up to 950 pg, up to 1 ng, up to 25 ng, up to 50 ng, up to 75 ng, up to 100 ng, up to 150 ng, up to 200 ng, up to 250 ng, up to 300 ng, up to 350 ng, up to 400 ng, up to 450 ng, up to 500 ng, up to 550 ng, up to 600 ng, up to 650 ng, up to 700 ng, up to 750 ng, up to 800 ng, up to 850 ng, up to 900 ng, up to 950 ng, up to 1,000 ng, or more.
[0063] In some versions, the extracted nucleic acid has a modal size from 156 bp and 177 bp, such as from 157 bp to 176 bp, from 158 bp to 175 bp, from 159 bp to 174 bp, from 160 bp to 173 bp, from 161 bp to 172 bp, from 162 bp to 171 bp, from 163 bp to 170 bp, from 164 bp to 169 bp, from 165 bp to 168 bp, or about 166-167. The modal size of nucleic acids in a sample can be determined by generating a quantitative fragment size distribution, methods for which are readily known in the art. In some versions, the extracted nucleic acid has a modal size of at least 147 bp.
[0064] In some versions of the invention, the solid substrate comprising the blood sample is stored prior to extracting the nucleic acid therefrom. “Stored” is used in this context is used merely to indicate a period of time, and not necessarily being located in a single location. For example, a sample can be stored for a week by being mailed from one location to another location over the course of a week.
[0065] In various versions of the invention, the substrate comprising the blood sample is stored prior to extracting the nucleic acid therefrom for a period of time of at least 1 hour, at least 12 hours, at least 1 day, at least 3 days, at least 1 week, at least 2 weeks, at least 1 month, at least 2 months, at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 7 months, at least 8 months, at least 9 months, at least 10 months, at least 11 months, at least 1 year, at least 2 years, or at least 3 years. In various versions of the invention, the substrate comprising the blood sample is stored prior to extracting the nucleic acid therefrom for a period of time of up to 1 hour, up to 12 hours, up to 1 day, up to 3 days, up to 1 week, up to 2 weeks, up to 1 month, up to 2 months, up to 3 months, up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, up to 5 years, up to 6 years, up to 7 years, up to 8 years, up to 9 years, up to 10 years, or more.
[0066] In various versions of the invention, the substrate comprising the blood sample is stored at a temperature prior to extracting the nucleic acid therefrom for a period of time of at least 1 hour, at least 12 hours, at least 1 day, at least 3 days, at least 1 week, at least 2 weeks, at least 1 month, at least 2 months, at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 7 months, at least 8 months, at least 9 months, at least 10 months, at least 11 months, at least 1 year, at least 2 years, or at least 3 years. In various versions of the invention, the substrate comprising the blood sample is stored at a temperature prior to extracting the nucleic acid therefrom for a period of time of up to 1 hour, up to 12 hours, up to 1 day, up to 3 days, up to 1 week, up to 2 weeks, up to 1 month, up to 2 months, up to 3 months, up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, up to 5 years, up to 6 years, up to 7 years, up to 8 years, up to 9 years, up to 10 years, or more. The temperature in some versions is a temperature of at least 0°C, at least 5°C, at least 10°C, at least 15°C, at least 20°C, at least 25°C, at least 30°C, or at least 35°C. The temperature in some versions is a temperature of up to 0°C, up to 5°C, up to 10°C, up to 15°C, up to 20°C, up to 25°C, up to 30°C, up to 35°C, up to 40°C, up to 45°C, up to 50°C, up to 55°C, up to 60°C, or more.
[0067] In various versions of the invention, the substrate comprising the blood sample is ambiently stored prior to extracting the nucleic acid therefrom for a period of time of at least 1 hour, at least 12 hours, at least 1 day, at least 3 days, at least 1 week, at least 2 weeks, at least 1 month, at least 2 months, at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 7 months, at least 8 months, at least 9 months, at least 10 months, at least 11 months, at least 1 year, at least 2 years, or at least 3 years. In various versions of the invention, the substrate comprising the blood sample is ambiently stored at room temperature prior to extracting the nucleic acid therefrom for a period of time of up to 1 hour, up to 12 hours, up to 1 day, up to 3 days, up to 1 week, up to 2 weeks, up to 1 month, up to 2 months, up to 3 months, up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, up to 5 years, up to 6 years, up to 7 years, up to 8 years, up to 9 years, up to 10 years, or more. “Ambiently stored” in this context refers to being stored in an environment in which the temperature is either not controlled or is controlled but is permitted to fluctuate beyond a defined temperature range, wherein the defined temperature range spans 20°C, 15°C, 10°C, 5°C, or 1°C depending on the particular embodiment.
[0068] Another aspect of the invention is directed to analyzing the extracted nucleic acid. The analyzing preferably comprises physically analyzing the extracted nucleic acid. “Physically analyzing” as used herein refers to performing at least one physical step on or with the extracted nucleic acid, such as performing a physical manipulation, interrogation, change of form, etc., on or with the nucleic acid. Examples include amplifying, sequencing, purifying (e.g., sequence-specific target capture), separating, and / or binding (e.g, sequence-specific binding with a probe) the extracted nucleic acid.
[0069] In some versions, the analysis of the extracted nucleic acid comprises analyzing a sequence of the extracted nucleic acid. The nucleic acid can be analyzed, for example, for the presence of characteristic sequences. The characteristic sequences can be sequences characteristic of a disease state, such as cancer or other disease states. The characteristic sequences can comprise structural variations, sequences of alleles of disease-associated genes, or other sequences associated with a disease state. In some versions, the sequence analysis comprises targeted amplification of the extracted nucleic acid. In some versions, the sequence analysis comprises sequencing the extracted nucleic acid. In some versions, the sequence analysis comprises targeted sequencing of the extracted nucleic acid. In some versions, the sequence analysis comprises non-targeted sequencing of the extracted nucleic acid. Some versions, the sequence analysis comprises whole-genome sequencing of the extracted nucleic acid. Extracted nucleic acid is this context refers to the nucleic acid directly extracted from the solid substrate and any nucleic acid generated therefrom in which original sequences of the nucleic acid directly extracted from the solid substrate are preserved, such as any libraries generated from the nucleic acid directly extracted from the solid substrate. In some versions, the analysis of the extracted nucleic acid compnses amplifying the extracted nucleic acid. In some versions, the analysis of the extracted nucleic acid comprises targeted amplification of the extracted nucleic acid. In some versions, the analysis of the extracted nucleic acid comprises non-targeted amplification of the extracted nucleic acid. The targeted amplification can comprise amplifying nucleic acid corresponding to only a portion of the genome. Such amplification can be employed, for example, when amplifying nucleic acid corresponding to only certain portions of the genome, such as genes, gene portions, or other portions of the genome associated with cancer or any particular type of cancer. The nontargeted amplification can comprise sequencing all nucleic acid regardless of the portion of the genome to which it corresponds. An example of such amplification is shotgun amplification.
[0070] A surprising aspect of the invention was the discovery that the fragmentation patterns of nucleic acid in a blood sample is preserved after applying the blood sample to a solid substrate, drying the blood sample on the solid substrate, and extracting the nucleic acid from the solid substrate.
[0071] Accordingly, in some versions of the invention, the analysis of the extracted nucleic acid comprises determining a fragmentation pattern of the extracted nucleic acid. The fragmentations patterns can include any quantifiable fragmentation characteristic of the extracted nucleic acid. Nonlimiting examples of such characteristics include the length of cfDNA fragments that align with one or more regions of a genome, a number of cfDNA fragments that align with one or more regions of a genome, a number of cfDNA fragments that start or end at each of one or more regions of a genome, a number of cfDNA fragments outside a nucleosome region, a number of cfDNA fragments within a nucleosome region, a size peak distribution of cfDNA fragments relative to a mappable genomic location, a particular location of a size peak of cfDNA fragments, a particular range of cfDNA fragment sizes associated with a size peak, or any combination thereof. Exemplary' methods of determining such characteristics are described in following examples or are otherwise known in the art.
[0072] An exemplary fragmentation pattern that can be used for analysis is a fragment size distribution. “Fragment size distribution'’ as used herein refers to a quantitation of the number of nucleic acids within each of one or more different size intervals. The quantitation can be an absolute or relative quantitation. The size of the nucleic acid is the length of the nucleic acid, and each size interval can be a single value (a single length) or range of values (a range of lengths). In some versions, the determining the fragmentation pattern of the extracted nucleic acid comprises determining a fragment size distribution, wherein the extracted nucleic acid has a modal size from 156 bp and 177 bp, such as from 157 bp to 176 bp, from 158 bp to 175 bp, from 159 bp to 174 bp, from 160 bp to 173 bp, from 161 bp to 172 bp, from 162 bp to 171 bp, from 163 bp to 170 bp, from 164 bp to 169 bp, from 165 bp to 168 bp, or about 166-167. In some versions, the determining the fragmentation pattern of the extracted nucleic acid comprises determining a fragment size distribution, wherein the extracted nucleic acid has a modal size of at least 147 bp.
[0073] Other methods of determining fragmentation patterns are well-known in the art (Liu 2022). These include information- weighted fraction of aberrant fragments analysis (Budhraja et al. 2023), determining incorporation of nucleotide frequencies at fragment ends (Budhraja et al. 2023), determining incorporation of nucleotide frequencies surrounding fragment ends (Budhraja et al. 2023), determining ratios of fragment sizes across one or more portions of a genome (Cristiano et al. 2019), determining motif frequencies at fragment ends, and determining any other fragmentation patterns as described in Budhraja et al. 2023, US 2024 / 0209455, Budhraja et al. 2023, Cristiano et al. 2019, Markus et al. 2021, Adalsteinsson et al. 2017, and Jiang et al. 2020.
[0074] In some versions, the determining fragmentation pattern comprises targeted sequencing of the extracted nucleic acid. In some versions, the determining fragmentation pattern comprises non-targeted sequencing of the extracted nucleic acid. In some versions, the determining fragmentation pattern comprises whole-genome sequencing of the extracted nucleic acid. Extracted nucleic acid is this context refers to the nucleic acid directly extracted from the solid substrate and any nucleic acid generated therefrom in which original sequences of the nucleic acid directly extracted from the solid substrate are preserved, such as any libraries generated from the nucleic acid directly extracted from the solid substrate. Such sequencing can be employed, for example, in a process of quantitating extracted nucleic acid fragments and / or mapping the extracted nucleic acid fragments to a genome. The targeted sequencing can comprise sequencing fragments corresponding to only a portion of the genome. Such sequencing can be employed, for example, when determining fragmentation patterns of nucleic acid fragments corresponding to only certain portions of the genome, such as genes, gene portions, or other portions of the genome associated with cancer or any particular type of cancer. The non-targeted sequencing can comprise sequencing all nucleic acid fragments regardless of the portion of the genome to which they correspond. Such sequencing can be employed, for example, in genome-wide fragmentation analysis. In some versions, the determining the fragmentation pattern comprises amplification of the extracted nucleic acid. In some versions, the determining fragmentation pattern comprises targeted amplification of the extracted nucleic acid. In some versions, the determining fragmentation pattern comprises non-targeted amplification of the extracted nucleic acid. The targeted amplification can comprise amplifying nucleic acid corresponding to only a portion of the genome. Such amplification can be employed, for example, when determining fragmentation patterns of nucleic acid corresponding to only certain portions of the genome, such as genes, gene portions, or other portions of the genome associated with cancer or any particular type of cancer. The non-targeted amplification can comprise sequencing all nucleic acid regardless of the portion of the genome to which it corresponds. An example of such amplification is shotgun amplification. In some versions, the amplification comprises quantitative PCR.
[0075] Other methods of determining fragmentation patterns and measuring fragmentation lengths are known in the art. Such methods may include purifying (e.g, sequence-specific target capture), separating, and / or binding (e.g, sequence-specific binding with a probe), electrophoresis, mass spectrometry, atomic force microscopy, electron microscopy, nuclear magnetic resonance, or other methods.
[0076] In some versions, the blood sample is from a subject that has a disease. In some versions, the blood sample is from a subject that has cancer. In some versions, the blood sample is from a subject suspected of having a disease. In some versions, the blood sample is from a subject suspected of having cancer. As shown in the following examples, the fragmentation patterns determined using the methods of the invention are able to distinguish subjects who have cancer from subjects who do not have cancer and therefore can be used in cancer diagnosis.
[0077] In some versions, the determining the fragmentation pattern of the extracted nucleic acid detects an aberrant fragmentation pattern relative to control nucleic acid. “Aberrant fragmentation pattern” in this context refers to a fragmentation pattern that differs from the control. The difference in the fragmentation patterns can be any characteristic of the patterns. In some versions, the difference in the fragmentation patterns is a different number of fragments of a certain size. In some versions, the difference in the fragmentation patterns is an increase in the number of fragments having an end in a recurrently protected region (RPR) of the genome (see, e.g., Budhraja et al. 2023). In some versions, the control nucleic acid is from a subject that does not have cancer.
[0078] The elements and method steps described herein can be used in any combination whether explicitly described or not. All combinations of method steps as used herein can be performed in any order, unless otherwise specified or clearly implied to the contrary by the context in which the referenced combination is made.
[0079] As used herein, the singular forms “a,” "an." and “the” include plural referents unless the content clearly dictates otherwise.
[0080] Numerical ranges as used herein are intended to include every number and subset of numbers contained within that range, whether specifically disclosed or not. Further, these numerical ranges should be construed as providing support for a claim directed to any number or subset of numbers in that range. For example, a disclosure of from 1 to 10 should be construed as supporting a range of from 2 to 8, from 3 to 7, from 5 to 6, from 1 to 9, from 3.6 to 4.6, from 3.5 to 9.9, and so forth.
[0081] All patents, patent publications, and peer- reviewed publications (i.e., “references”) cited herein are expressly incorporated by reference to the same extent as if each individual reference were specifically and individually indicated as being incorporated by reference. In case of conflict between the present disclosure and the incorporated references, the present disclosure controls.
[0082] It is understood that the invention is not confined to the particular construction and arrangement of parts herein illustrated and described, but embraces such modified forms thereof as come within the scope of the claims.
[0083] EXAMPLES
[0084] EXAMPLE 1. ANALYSIS OF PLASMA DNA FRAGMENTATION PATTERNS FROM DRIED BLOOD SPOTS
[0085] Summary
[0086] Circulating tumor DNA analysis holds promise for early detection of cancer. However, costs and complexity of genomic methods, and logistical requirements for blood collection and processing limit large scale implementation. Here, we evaluated the feasibility of genomewide analysis of fragmentation patterns in plasma DNA obtained from dried blood spots (DBS). Across 75 individuals including 25 patients with cancer, we prepared conventional DBS, and plasma-separating DBS (psDBS) from whole blood samples, prior to plasma separation. In paired comparisons of plasma DNA between psDBS and blood tubes, we observed correlated concentrations, similar distributions of fragment lengths, correlated genome-wide fragmentation features, and equivalent performance of machine learning algorithms for cancer detection. Our results provide proof-of-principle that plasma DNA fragmentation analysis can be performed from small amounts of DNA obtained from psDBS. With further development and validation, this approach can expand the reach and impact of blood-based early cancer detection, particularly for resource constrained environments such as low- and middle-income countries and rural healthcare systems.
[0087] Introduction
[0088] Circulating tumor DNA analysis has shown promising results for cancer detection and response monitoring during treatment (Wan et al. 2017). Early work in this field relied on deep molecular analysis of a few genomic loci with known or recurrent somatic genomic alterations and methylation changes (McDonald et al. 2019, Cohen et al. 2018, Schrag et al. 2023). More recently, new approaches have traded high sequencing depth at a few’ loci with wide breadth of genomic coverage at low depth (Budhraja et al. 2023, Cristiano et al. 2019). Low-depth whole genome sequencing methods analyze variation in genome-wide features such as sequencing coverage (driven by copy number alterations), mutation signatures, fragmentomic features, and nucleotide frequencies surrounding fragment ends (Budhraja et al. 2023, Cristiano et al. 2019, Markus et al. 2021, Adalsteinsson et al. 2017, Jiang et al. 2020). Incorporating several thousand or more loci from across the genome into these features enables high accuracy while reducing requirements of sequencing depth, amount of input DNA for sequencing library preparation, and sequencing costs (Markus et al. 2022). However, unlike somatic single nucleotide variants, differences in fragmentation features are not inherently tumor specific, making them more susceptible to pre-analytical sources of variation (Moser et al. 2023, van der Pol et al. 2019, Markus et al. 2018). Wide adoption of plasma DNA fragmentation analysis for cancer detection and monitoring is hindered by stringent upstream requirements for collection, processing, and storage of blood samples. In resourcelimited environments, particularly low- and middle-income countries (LMICs), the costs associated with phlebotomy, sample processing to isolate plasma, storage, and shipping while maintaining an effective cold chain can be prohibitive (Radich et al. 2022). These challenges reduce the likelihood that patients will receive molecular testing and perpetuate disparities in cancer detection and outcomes (Drake et al. 2018). An inexpensive, and logistically simpler alternative has immense potential to bring blood-based early detection of cancer to resource limited environments.
[0089] We have recently show i that robust fragmentation analysis is feasible from sequencing data representing 1 to 10 million plasma DNA fragments from each sample (equivalent to genomic coverage of 0.05 to 0.5x) (Budhraja et al. 2023). At the reported mean concentration of 4 ng / mL or 1200 haploid genome equivalents / mL of plasma DNA in healthy individuals, adequate amount of plasma DNA for shallow whole genome sequencing should be available in as little as 100 to 150 pL of blood (3-5 drops of blood). Here, we test this hypothesis and evaluate whether plasma DNA fragmentation patterns can be analyzed from a few drops of blood, when collected as dried blood spots (DBS).
[0090] Materials and Methods
[0091] Enrollment and sample collection
[0092] Healthy volunteers and cancer patients were enrolled at the Translational Genomics Research Institute and by the University of Wisconsin-Madison Biobank. Samples were collected under predetermined protocols. Samples from cancer patients were collected prior to surgery. No patients had received neoadjuvant chemotherapy.
[0093] Sample processing
[0094] All plasma samples were collected in EDTA tubes and centrifuged twice at room temperature. Plasma aliquots were stored at -80° C prior to DNA extraction.
[0095] For conventional dried blood spots (eDBS) and plasma-separating dried blood spots (psDBS), whole blood was collected in EDTA or preservative-free tubes. All samples collected in preservative-free tubes were spotted within 20 minutes. Samples collected in EDTA tubes were spotted within 1 hour. For all spots, whole blood was pipetted onto the respective devices. eDBS were prepared on QIAcard FTA micro cards (Qiagen) from 125 pl whole blood; psDBS were prepared on the HemaSpot™ SE device from 125 - 150 pl whole blood. Spots were fully dried before storage at room temperature in a zippered plastic bag.
[0096] DNA extraction
[0097] Plasma DNA was extracted from 1 mL plasma with the MagMAX Cell-Free DNA Isolation Kit (ThermoFisher) according to manufacturer instructions, using the EDTA tube workflow.
[0098] For eDBS and psDBS, DNA was extracted with the MagMAX Cell-Free DNA Isolation Kit (ThermoFisher) according to a modified protocol. For the eDBS, one-half to the full spot was excised and cut into quarters. For the psDBS, the entirety of the plasmacontaining portion of the membrane was excised and cut into pieces with a maximum length of approximately 10 millimeters. Seven hundred fifty microliters of phosphate buffered saline was added to the sample, and the sample was lysed with proteinase K and sodium dodecyl sulfate according to manufacturer instructions. The supernatant, excluding any portions of the membrane, was collected and added to the binding solution / beads mix. All remaining aspects of the extraction protocol were unchanged from manufacturer instructions.
[0099] Extracted DNA was assessed and quantified with the cfDNA assay on the TapeStation system (Agilent).
[0100] Library preparation & sequencing
[0101] Whole genome libraries were prepared from plasma, eDBS, and psDBS DNA. For plasma DNA, input was limited to a maximum of 2 ng, and 10 cycles of amplification were performed. For eDBS and psDBS DNA, a fixed input of 12 pl (of a total elution volume of 30 pl) of the sample was used, and 14 cycles of amplification were performed. All libraries were prepared using the ThruPLEX DNA-Seq HV kit (TakaraBio). Libraries were quantified using the Tapestation D1000 HS assay (Agilent). Sequencing was performed on the NextSeq 1000 / 2000 system (Illumina) to generate 100 base pair paired-end reads. Demultiplexing, trimming, and FASTQ conversion was performed on-board using DRAGEN (version).
[0102] Data processing
[0103] All sequencing data was aligned to human genome build hgl9 (pl3_105) using BWA- MEM. Files were converted to BAM format using SAMtools, and duplicate reads were marked. Analysis was limited to non-duplicate reads marked as properly paired with a minimum mapping quality of 60. The fragment size of all aligned reads was calculated using bedtools (Quinlan et al. 2010). Fragments between 30 and 1000 base pairs were analyzed.
[0104] Copy number analysis
[0105] Copy number analysis was performed on aligned sequencing data with ichorCNA_U (Favaro et al. 2022), a fork of ichorCNA (Adalsteinsson et al. 2017). Only somatic chromosomes were analyzed, and a bin size of 500 kb was used.
[0106] Measurement of aberrant fragmentation
[0107] Aberrant fragmentation scores (AFS) were calculated as previously described (Budhraja et al. 2023). To calculate AFS, the same map of recurrently protected regions (RPRs) as published was used. Fragments fully spanning an RPR were considered nonaberrant, while fragments with one or more ends falling within an RPR were considered aberrant. AFS was calculated as a proportion of aberrant fragments to all fragments intersecting an RPR, with adjustments for fragment length and GC content as previously described. Only fragments between 140 and 220 bp were considered for calculation of AFS. Single nucleotide frequencies
[0108] Nucleotide frequencies were calculated as previously described (Budhrajaet al. 2023). For all fragments between 50 and 1000 base pairs, the reference sequence from 10 base pairs outside the fragment end to the first 10 base pairs of the fragment end was collected with Homertools (Heinz et al. 2010). Average frequencies at all positions were calculated based on the reference sequence.
[0109] End motifs
[0110] Counts of four-mer and two-mer end motifs were calculated for the primary alignment of all non-duplicate fragments. To assess similarity between sample types, hierarchical clustering was performed with Seaborn (Waskom et al. 2021), using average' as the linkage method and ’ Euclidean' as the distance metric.
[0111] Short / long fragments
[0112] Counts of short (100 to 150 bp) and long (151 - 220 bp) fragments were measured in 5000 kb bins across the genome. From 590 initial bins, only bins with at least 2000 long fragments were included in analysis, resulting in 544 bins. The ratio of short to long fragments in each bin was calculated for these 544 bins, and then each sample was normalized to a mean bin fragment ratio of 0, with a standard deviation of 1.
[0113] Modeling
[0114] Classification performance was assessed using four machine learning models. For all models, an ensemble model of 100 random forest models was generated using 100 iterations of cross-validation, and the score for each sample was taken as the average of all scores when the sample was in the validation split. For each iteration, the data is split into an 80 / 20 train / validation split, and a random forest model is generated with a maximum depth of 5 and a minimum samples per leaf of 5.
[0115] The genome-wide analysis of fragment ends (GALYFRE) model was implemented with features and hyperparameters as described previously (Budhraja et al. 2023). Briefly, 10 features are used, including AFS, and 9 nucleotide frequency features at positions 2 and 3 base pairs within the fragment and at 1 base pair outside the fragment. Statistical analysis
[0116] All statistical analysis were conducted using SciPy (version 1.13) in Python. Linear correlation was measured using the Pearson correlation coefficient. Statistical differences for continuous variables and receiver operator characteristic curves were calculated using the Mann- Whitney U test. All reported P-values are two-sided.
[0117] Results
[0118] Samples and processing
[0119] From a total of 75 participants including 45 healthy individuals, 25 patients with cancer and 5 patients with non-malignant disease, we collected blood samples and spotted conventional DBS (eDBS) as well as plasma-separating DBS (psDBS) from whole blood prior to further processing and isolation of plasma (FIG. 1 and Table 1). All paired samples underwent DNA extraction and were analyzed using shallow whole genome sequencing (WGS) performed without any further DNA fragmentation. Mean coverage was 0.27x for plasma DNA from blood tubes (SD 0.11), 0.24x for eDBS (SD 0.11), and 0.14x for psDBS (SD 0.07).
[0120] Table 1.
[0121] Demographic characteristics of healthy individuals and patients included in this study.
[0122] Table 2.
[0123] Clinical characteristics of patients with cancer included in this study
[0124] Evaluation of plasma DNA concentration and fragment lengths
[0125] Following DNA extraction and elution, plasma DNA concentration was too low in DBS samples for accurate quantification using fluorometric or PCR-based methods. However, for one set of outlier samples where plasma DNA concentration in the blood tube was 127 ng / ml, electrophoresis of extracted DNA showed a fragment length distribution consistent with plasma cell-free DNA in psDBS and a large fraction of DNA >1000 bp in eDBS (FIG. 2). In addition, whole genome libraries prepared using extracted DNA from psDBS and eDBS showed length distributions consistent with plasma DNA libraries prepared from blood tubes (FIG. 3). To infer whether DNA concentration was preserved in DBS, we measured DNA concentration in plasma from blood tubes and compared it with the yield of WGS libraries from each set of DBS samples (such that all compared samples underwent the same number of amplification cycles). Library yield was correlated with plasma DNA concentration in blood tubes when WGS libraries were prepared from psDBS (^=0.46; n=74; p=2.6 x 10'11, FIG. 4 A), but not when libraries were prepared from eDBS (r2=(). 14, n=43, p=l .4 x 1 O'2; FIG. 4B). One sample with very7high plasma DNA concentration in the blood tube (127 ng / ml) was excluded from this analysis.
[0126] To evaluate whether plasma DNA fragment lengths were preserved in DBS, we compared insert size distributions measured using WGS data in samples obtained from healthy individuals. The modal size of all three sample sets was 165 - 166 bp, similar to prior reports of plasma DNA fragment lengths associated with mono-nucleosomes (FIG. 4C) (Lo et al. 2010). Plasma DNA libraries from blood tubes and psDBS showed a similar and large proportion of fragments at this size, compared to eDBS samples which showed a smaller proportion of mono-nucleosome length fragments. Blood tubes and psDBS samples also showed a 10-bp periodicity below7130 bp, suggestive of enzymatic degradation of nucleosome-associated DNA as previously described (Markus et al. 2021). In addition, both sample types showed very few fragments between 230 and 300 bp, and a small second modal peak at approximately 340 bp, a fragment length associated with DNA preserved in dinucleosomes. In comparison, eDBS libraries showed a far higher proportion of fragments between 40 and 150 bp without 10 bp periodicity and did not show’ a distinct peak of di- nucleosomal length. For plasma samples from blood tubes, fragment length distributions were highly conserved across samples (FIG. 4D). For psDBS samples, most samples had a similar distribution, but a subset had very high contribution of fragments under 120 bp, displaying strong periodicity (FIG. 4E). For eDBS samples, fragment size distributions were highly variable (FIG. 4F).
[0127] Evaluation of genome-wide plasma DNA fragmentation features
[0128] To evaluate whether genome-wide plasma DNA fragmentation patterns w’ere preserved in DBS, we measured multiple fragmentation features using WGS and compared them across the paired sample sets.
[0129] In recent work, we have developed a cancer detection assay called genome-wide analysis of fragment ends (GALYFRE), that incorporates multiple fragmentomic features measured using WGS into a machine learning model (Budhraja et al. 2023). One feature included in GALYFRE is aberrant fragmentation score (AFS), a measurement based on the fraction of fragment ends observed within nucleosome-associated recurrently protected genomic regions in each sample. In the current study, we found no correlation in AFS between blood tubes and eDBS (FIG. 5A, r^O.O, n=44, pH).92). but observed a positive correlation between blood tubes and psDBS (FIG. 5B, r2=0.24, n=75, p=9.5 x 10'6). We observed that in a subset of psDBS where the AFS was higher than the corresponding plasma samples from blood tubes, the psDBS samples also had disproportionate contributions of fragments <=120 bp. Once psDBS samples with >20% of fragments <=120 bp were excluded (n=17, 23% of total samples), the coefficient of determination (r2) for the linear relationship between blood tubes and psDBS samples improved to 0.66 (n=58, p=5.6 x 1016; FIG. 5C). In healthy samples, AFS in eDBS samples were much higher than in plasma from blood tubes or psDBS (p = 9.3x IO’7; FIG. 5D).
[0130] Another set of fragmentomic features included in GALYFRE was nucleotide frequencies surrounding fragment ends (Budhraja et al. 2023). In the current study, we compared mean fragment end nucleotide frequencies, measured for each position surrounding the 5’end of plasma DNA fragments, averaged across all fragments in a sample. Using principal components to analyze single nucleotide frequencies at plasma DNA fragment ends, we found that blood tubes overlapped with psDBS samples, but eDBS samples were clearly separated in the first principal component (FIGS. 5E and 5F). Individual nucleotide frequencies at each position surrounding the 5’end of DNA fragments showed better correlation between psDBS and blood tubes than between eDBS and blood tubes (FIGS. 6A- 6C and FIGS. 7A-7C). Similarly, when considering dinucleotide frequencies at the 5’ fragment end (Jiang et al. 2020, Wong et al. 2024), we found that blood tubes overlapped together with psDBS, while eDBS was separated from the two in the first principal component (FIGS. 5G and 5H). Using hierarchical clustering based on dinucleotide frequencies, psDBS and plasma samples from blood tubes clustered together while eDBS showed clear differences and clustered separately. The most common dinucleotide observed at the 5’ fragment ends in blood tubes and psDBS was CC. Using a similar approach to evaluate 4-mer motifs at fragment ends, we found psDBS and plasma from blood tubes were more similar to each other, compared to eDBS (FIG. 8).
[0131] One recent study showed that genome-wide variation in the ratio of short fragments (100 to 150 bp) to long fragments (151 to 220 bp) (Cristiano et al. 2019) can be informative for cancer detection. Here, we measured the ratio of short to long fragments for 544 bins of 5000 kbp across the genome. Using principal components analysis, we found that blood tubes and psDBS samples overlapped in the first principal component while eDBS samples showed a clear difference (FIGS. 51 and 5J). Classification between cancer and healthy individuals using plasma DNA fragmentation analysis
[0132] Encouraged by results from comparisons of fragmentation features, we evaluated the potential to detect cancer-derived DNA using psDBS samples compared to corresponding plasma samples from blood tubes. Blood tube and psDBS samples were collected from 25 patients with cancer across multiple cancer types (breast, head and neck, pancreas, endometrium, kidney, prostate, and gall bladder). 21 / 25 samples (84%) were collected from patients with Stage I-III cancer (Table 2). We first evaluated whether copy number aberrations were detectable in plasma DNA from blood tubes. Of 25 patients with cancer, we found copy number aberrations detectable in plasma from 2 patients. Copy number aberrations detected in corresponding psDBS samples were nearly identical (FIG. 9A and FIG. 10).
[0133] Next, we evaluated performance of multiple machine learning models based on different sets of fragmentation features to classify between cancer patients and healthy individuals. For each model, we compared its performance using plasma DNA WGS from blood tubes and corresponding psDBS samples. Using GALYFRE, a random forest model that incorporates AFS and nucleotide frequencies surrounding fragment ends (Budhraja et al. 2023), we observed an area under the receiver operating characteristic curve (AUROC) of 0.73 and 0.83 for blood tubes and psDBS samples, respectively (FIG. 9B, p=0.27). We also evaluated the performance of a random forest model based on the frequency of 256 4-mer motifs observed at the 5 ’end of DNA fragment ends, as recently described (Jiang et al. 2020). We observed an AUROC of 0.79 and 0.89 for blood tubes and psDBS samples, respectively (FIG. 9C, p=0. 11). We further evaluated the performance of a random forest model incorporating short-to-long DNA fragment length ratios measured in each of 544 bins of 5000 kbp across the genome, as recently described (Cristiano et al. 2019). We observed an AUROC of 0.71 and 0.73 for blood tubes and psDBS samples, respectively (FIG. 9D, p=0.84). Finally, we built an ensemble model incorporating all these features together. Using this approach, we observed an AUROC of 0.81 and 0.90 for blood tubes and psDBS samples, respectively (FIG. 9E, p=0.06).
[0134] Discussion
[0135] Detection and treatment of cancer at earlier stages of disease dramatically increases the likelihood of achieving cure. Several studies have demonstrated the relevance of circulating tumor DNA analysis for cancer detection, either as a multicancer test (Schrag et al. 2023) or for individual cancer types (Mazzone et al. 2024, Medina et al. 2024). Methods for circulating tumor DNA analysis have evolved from targeted molecular assays to genome-wide approaches that integrate hundreds to thousands of genomic loci to generate features for machine learning models and achieve high classification accuracy (Moser et la. 2023). Such genome-wide approaches often rely on shallow sequencing coverage and require limited amounts of input DNA. In earlier work, we had performed in silico analysis to show that plasma DNA fragmentation patterns could be reliably measured using data representing only 1-2 million fragments (Budhraja et al. 2023). Here, we have tested this hypothesis and found that plasma DNA fragmentation patterns are largely preserved in plasma-separating dried blood spots. In addition, we found that psDBS can achieve equivalent performance for cancer detection when compared with corresponding, paired plasma samples obtained from blood tubes.
[0136] Since concentration of DNA in plasma is generally much lower than the peripheral blood cells compartment, the field of circulating tumor DNA analysis has long insisted on stringent conditions for sample processing and storage to ensure accurate results (Markus et al. 2018, Greytak et al. 2020). For large multisite clinical trials, this increases costs associated with sample collection for ctDNA analysis because it requires either trained local biospecimen processing teams at each site or real-time shipping of samples collected in specialized cell- free DNA blood tubes to a central biorepository. Here, we have found that when plasma is separated from the blood cell fraction at the time of blood spot collection, plasma DNA fragmentation patterns are preserved in psDBS and can generate insights for early cancer detection that are similar to corresponding blood tubes.
[0137] Collection of psDBS is straightforward and could be performed in ambulatory care or self-collected at home, without requiring a trained phlebotomist. Once the blood spot has dried, there is no additional processing or cold-chain storage required, and blood spots can be shipped at ambient temperature to a central lab for processing (Gyanchandani et al. 2018). These improvements in logistics of sample collection can enable future clinical studies and applications such as early detection of cancer in primary care patients and distributed monitoring of treatment response. In addition, this approach can enable the expansion of current advances in cancer detection and monitoring to underserved populations in resourcelimited settings, including LMICs and rural healthcare systems. Addressing this challenge is urgent and critical, since the burden of cancer is rapidly rising in LMICs where nearly two- thirds of all new cancer cases and 70% of all cancer deaths are projected to occur by 2030 (Global Burden of Disease Cancer Collaboration, 2017).
[0138] To enable evaluation of multiple blood spots from each participant and comparison with blood tubes obtained at the same time, we collected blood tubes from patients and then prepared eDBS and psDBS from whole blood. This allowed us to minimize participant discomfort while evaluating different sample collection devices, and to establish the minimum volume of blood sample needed for DBS. As a result, our analysis of psDBS represents DNA fragmentation patterns measured from venous blood rather than capillary blood collected from fingersticks. One study using PCR amplification of a repetitive genomic locus using multiple amplicon sizes showed no difference in DNA integrity between venous and capillary cfDNA (Breitbach et al. 2014). We predict capillary blood can be used with the methods of the present invention.
[0139] In summary, we report that analysis of DNA fragmentation patterns from plasmaseparating dried blood spots is feasible and can overcome the current logistical barriers that limit scalability of plasma DNA testing. This approach can vastly improve access to early cancer detection as well as other diagnostic applications of plasma DNA fragmentation analysis beyond oncology (Sun et al. 2019).
[0140] EXAMPLE 2. COMPARISON OF DRIED BLOOD SPOT DEVICES
[0141] Methods
[0142] Samples and processing
[0143] Ten healthy individuals were recruited by the UW Biobank and one tube of whole blood (EDTA tube) was collected from each. Within 3 hours, up to five different plasmaseparating dried blood spot devices (psDBS) were prepared by applying whole blood with a pipette. The remaining whole blood was processed by two rounds of centrifugation at 1500 x g for 15 minutes, and the resulting plasma was stored at -80°C prior to extraction. Devices were stored at room temperature prior to extraction for a period of 9 - 12 days.
[0144] Device overview
[0145] • AdvanceDx 100™ serum collection card (“ADX100”), manufactured by Advance Dx, Inc. (Scottsdale, AZ, USA). Between 125 - 150 pl was applied to the application surface of the card. The card was dried for a period of one hour, before being placed in the provided zipped storage bag. This storage bag contained a desiccant packet and an oxygen absorption packet. For extraction, the plasma portion of a single device was extracted.
[0146] • HemaSep™ Punch card, manufactured by Ahlstrom (Helsinki, Finland). Two spots per punch card were prepared, each with a volume of 100 pl whole blood. After drying for one hour, cards were placed in an empty zippered plastic bag. For extraction, the plasma portion from a single spot was extracted.
[0147] • HemaSep™ Strip card, manufactured by Ahlstrom (Helsinki, Finland). For each card, 50 pl of whole blood was applied to each of 4 strips. After drying for one hour, cards were placed in an empty zippered plastic bag. For extraction, the four strips from each card were pooled into a single extraction.
[0148] • Telimmune™ DUO Plasma Separation Card (“Telimmune™”) - Note this device was previously called the Noviplex Duo. Manufacturer is Telimmune, LLC (Telimmune™ LLC, North Webster, IN). For two individuals, nine Telimmune™ cards were prepared from each by applying either 60 or 75 pl whole blood per card. For extraction, plasma discs from all nine cards were pooled into a single extraction.
[0149] Data generation
[0150] Cell-free DNA was extracted from each device from the plasma-containing portion of the device (excluding any areas with visible red blood cells) plus matched plasma and whole genome sequencing libraries were prepared. Sequencing and bioinformatics analyses were performed. All methods (extraction, library preparation, sequencing, and analysis) were conducted as described in main methods unless described otherwise.
[0151] Results
[0152] A total of 23 devices were prepared from ten healthy individuals: 10 ADX100, 6 HemaSep™ Punch, 5 HemaSep™ Strip, and 2 Telimmune™. Matched plasma from all ten healthy individuals was analyzed.
[0153] Fragment size distributions of DNA from the three psDBS devices were similar to DNA from blood tubes. For all libraries, the modal fragment size was 165 - 167 base pairs, a fragment size associated with a mononucleosome. In addition, all libraries had few fragments from 230 to 300 base pairs, but a visible peak in the 320 - 360 base pair range associated with the dinucleosome (FIGS. 11 A-l IE).
[0154] Aberrant fragmentation scores (AFS) were calculated for all fragments in the range of 50 - 1000 base pairs. Telimmune™ devices were excluded from AFS analysis as too few samples were available for meaningful correlations. AFS were highly correlated between matched plasma and all devices (FIGS. 12A-12C).
[0155] For 36 I 48 inside fragment nucleotide frequencies across all devices, the coefficient of determination (r2) was at least 0.5 (Table 3). For all three devices, the majority of positions had an r2of at least 0.5. For the 48 possible outside fragment nucleotide frequencies across all devices, 46 had an r2of at least 0.5 (Table 4). Telimmune™ devices were excluded from nucleotide frequency analysis as too few samples were available for meaningful correlations.
[0156] Table 3. Coefficient of determination (r2) between matched plasma samples and three psDBS devices for 12 nucleotide frequencies inside fragment ends.
[0157] IF01 = Inside Fragment position 1, the first base sequenced. IF02 = Inside Fragment position
[0158] 2, the second base sequenced. IF03 = Inside fragment position 3, the third base sequenced. IF04 = Inside Fragment position 4, the fourth base sequenced.
[0159] Table 4. Coefficient of determination (r2) between matched plasma samples and three psDBS devices for 12 nucleotide frequencies outside fragment ends.
[0160]
[0161] OFOl = Outside Fragment position 1, the first base outside the sequenced fragment. OF02 = Outside Fragment position 2, the second base outside the sequenced fragment. OF03 = Outside fragment position 3, the third base outside the sequenced fragment. OF04 = Outside Fragment position 4, the fourth base outside the sequenced fragment.
[0162] For all samples (blood tube plasma, ADX100, HemaSep™ Punch, HemaSep™ strip, and Telimmune™) the most common dinucleotide motif at the fragment end was CC, with the second most common being CA. For all samples except one HemaSep™ punch sample, the most common 4-mer end motif was CCCA.
[0163] The ratio of short to long fragments in 5 megabase bins across the genome was conserved in all 3 psDBS devices (FIG. 13A-13F). Note consistent elevations in chromosomes 3, 11, 19, and 20.
[0164] REFERENCES
[0165] V. A. Adalsteinsson, G. Ha, S. S. Freeman, A. D. Choudhury, D. G. Stover, H. A. Parsons, G. Gydush, S. C. Reed, D. Rotem, J. Rhoades, D. Loginov, D. Livitz, D. Rosebrock, I. Leshchiner, J. Kim, C. Stewart, M. Rosenberg, J. M. Francis, C. Z. Zhang, O. Cohen, C. Oh, H. Ding, P. Polak, M. Lloyd, S. Mahmud, K. Helvie, M. S. Merrill, R. A. Santiago, E. P. O'Connor, S. H. Jeong, R. Leeson, R. M. Barry, J. F. Kramkowski, Z. Zhang, L. Polacek, J. G. Lohr, M. Schleicher, E. Lipscomb, A. Saltzman, N. M. Oliver, L. Marini, A. G. Waks, L. C. Harshman, S. M. Tolaney, E. M. Van Allen, E. P. Winer, N. U. Lin, M. Nakabayashi, M. E. Taplin, C. M. Johannessen, L. A. Garraway, T. R. Golub, J. S. Boehm, N. Wagle, G. Getz, J. C. Love, M. Meyerson, Scalable whole- exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat Commun 8, 1324 (2017).
[0166] Baillargeon KR, Morbioli GG, Brooks JC, Miljanic PR, Mace CR. Direct Processing and Storage of Cell-Free Plasma Using Dried Plasma Spot Cards. ACS Meas Sci Au. 2022 Oct 19;2(5):457-465
[0167] S. Breitbach, B. Sterzing, C. Magallanes, S. Tug, P. Simon, Direct measurement of cell-free DNA from serially collected capillary plasma during incremental exercise. J Appl Physiol (1985) 117, 119-130 (2014).
[0168] K. K. Budhraja, B. R. McDonald, M. D. Stephens, T. Contente-Cuomo, H. Markus, M. Farooq, P. F. Favaro, S. Connor, S. A. Byron, J. B. Egan, B. Ernst, T. K. McDaniel, A. Sekulic, N. L. Tran, M. D. Prados, M. J. Borad, M. E. Berens, B. A. Pockaj, P. M. LoRusso, A. Bryce. J. M. Trent, M. Murtaza, Genome-wide analysis of aberrant position and sequence of plasma DNA fragment ends in patients with cancer. Sci Transl Med 15, eabm6863 (2023).
[0169] J. D. Cohen, L. Li, Y. Wang, C. Thobum, B. Afsari, L. Danilova, C. Douville, A. A. Javed, F. Wong, A. Mattox, R. H. Hruban, C. L. Wolfgang, M. G. Goggins, M. Dal Molin, T. L. Wang, R. Roden, A. P. Klein, J. Ptak, L. Dobbyn, J. Schaefer, N. Silliman, M. Popoli, J. T. Vogelstein, J. D. Browne. R. E. Schoen, R. E. Brand, J. Tie, P. Gibbs, H. L. Wong, A. S. Mansfield, J. Jen, S. M. Hanash, M. Falconi, P. J. Allen, S. Zhou, C. Bettegowda, L. A. Diaz, Jr., C. Tomasetti, K. W. Kinzler, B. Vogelstein, A. M. Lennon, N. Papadopoulos, Detection and localization of surgically resectable cancers with a multi-analyte blood test. Science 359, 926-930 (2018).
[0170] S. Cristiano, A. Leal, J. Phallen, J. Fiksel, V. Adleff, D. C. Bruhm, S. O. Jensen, J. E. Medina,
[0171] C. Hruban, J. R. White, D. N. Palsgrove, N. Niknafs, V. Anagnostou, P. Forde, J. Naidoo, K. Marrone, J. Brahmer, B. D. Woodward, H. Husain, K. L. van Rooijen, M. W. Omtoft, A. H. Madsen, C. J. H. van de Velde, M. Verheij, A. Cats, C. J. A. Punt, G. R. Vink, N. C. T. van Grieken, M. Koopman, R. J. A. Fijneman, J. S. Johansen, H. J. Nielsen, G. A. Meijer, C. L. Andersen, R. B. Scharpf, V. E. Velculescu, Genomewide cell-free DNA fragmentation in patients with cancer. Nature 570, 385-389 (2019).
[0172] T. M. Drake, S. R. Knight, E. M. Harrison, K. Soreide, Global Inequities in Precision Medicine and Molecular Cancer Research. Front Oncol 8, 346 (2018). P. F. Favaro, S. D. Stewart, B. R. McDonald, J. Cawley, T. Contente-Cuomo, S. Wong, W. P. D. Hendricks, J. M. Trent, C. Khanna, M. Murtaza, Feasibility of circulating tumor DNA analysis in dogs with naturally occurring malignant and benign splenic lesions. Sci Rep-Uk 12, (2022).
[0173] Global Burden of Disease Cancer Collaboration; Fitzmaurice C, Allen C, Barber RM, Barregard L, Bhutta ZA, Brenner H, Dicker DJ, Chimed-Orchir O, Dandona R, Dandona L, Fleming T, Forouzanfar MH, Hancock J, Hay RJ, Hunter-Merrill R, Huynh C, HosgoodHD, Johnson CO, Jonas JB, Khubchandani J, Kumar GA, Kutz M, Lan Q, Larson HJ, Liang X, Lim SS, Lopez AD, MacIntyre MF, Marczak L, Marquez N, Mokdad AH, Pinho C, Pourmalek F, Salomon JA, Sanabria JR, Sandar L, Sartorius B, Schwartz SM, Shackelford KA, Shibuya K, Stanaway J, Steiner C, Sun J, Takahashi K, Vollset SE, Vos T, Wagner JA, Wang H, Westerman R, Zeeb H, Zoeckler L, Abd- Allah F, Ahmed MB, Alabed S, Alam NK, Aldhahri SF, Alem G, Alemayohu MA, Ah
[0174] R, Al-Raddadi R, Amare A, Amoako Y, Artaman A, Asayesh H, Atnafu N, Awasthi A, Saleem HB, Barac A, Bedi N, Bensenor I, Berhane A, Bernabe E, Betsu B, Binagwaho A, Boneya D, Campos-Nonato I, Castaneda-Oijuela C, Catala-Lopez F, Chiang P, Chibueze C, Chitheer A, Choi JY, Cowie B, Damtew S, das Neves J, Dey
[0175] S, Dharmaratne S, Dhillon P, Ding E, Driscoll T, Ekwueme D, Endries AY, Farvid M, Farzadfar F, Fernandes J, Fischer F, G / Hiwot TT, Gebru A, Gopalani S, Hailu A, Horino M, Horita N, Husseini A, Huybrechts I, Inoue M, Islami F, Jakovljevic M, James S, Javanbakht M, Jee SH, Kasaeian A, Kedir MS, Khader YS, Khang YH, Kim D, Leigh J, Linn S, Lunevicius R, El Razek HMA, Malekzadeh R, Malta DC, Marcenes W, Markos D, Melaku YA, Meles KG, Mendoza W, Mengiste DT, MeretojaTJ, Miller TR, Mohammad KA, Mohammadi A, Mohammed S, Moradi-Lakeh M, Nagel G, Nand D, Le Nguyen Q, Nolte S, Ogbo FA, Oladimeji KE, Oren E, Pa M, Park EK, Pereira DM, Plass D, Qorbani M, Radfar A, Rafay A, Rahman M, Rana SM, Soreide K, Satpathy M, Sawhney M, Sepanlou SG, Shaikh MA, She J, Shiue I, Shore HR, Shrime MG, So S, Soneji S, Stathopoulou V, Stroumpoulis K, Sufiyan MB, Sykes BL, Tabares-Seisdedos R, Tadese F, Tedla BA, Tessema GA, Thakur JS, Tran BX, Ukwaja KN, Uzochukwu BSC, Vlassov VV, Weiderpass E, Wubshet Terefe M, Yebyo HG, YimamHH, Yonemoto N, Younis MZ, Yu C, Zaidi Z, Zaki MES, Zenebe ZM, Murray CJL, Naghavi M. Global, Regional, and National Cancer Incidence, Mortality, Years of Life Lost, Years Lived With Disability, and Disability-Adjusted Life-years for 32 Cancer Groups, 1990 to 2015: A Systematic Analysis for the Global Burden of Disease Study. JAMA Oncol. 2017 Apr l;3(4):524-548. Erratum in: JAMA Oncol. 2017 Mar 1;3(3):418.
[0176] S. R. Greytak, K. B. Engel, S. Parpart-Li, M. Murtaza, A. J. Bronkhorst, M. D. Pertile, H. M. Moore, Harmonizing Cell-Free DNA Collection and Processing Practices through Evidence-Based Guidance. Clin Cancer Res 26, 3104-3109 (2020).
[0177] R. Gyanchandani, E. Kvam, R. Heller, E. Finehout, N. Smith, K. Kota, J. R. Nelson, W.
[0178] Griffin, S. Puhalla, A. M. Brufsky, N. E. Davidson, A. V. Lee, Whole genome amplification of cell-free DNA enables detection of circulating tumor DNA mutations from fingerstick capillary blood. Sci Rep 8, 17313 (2018).
[0179] K. Heider, J. C. M. Wan, J. Hall, J. Belie, S. Boyle, I. Hudecova, D. Gale, W. N. Cooper, P. G. Corrie, J. D. Brenton, C. G. Smith, N. Rosenfeld, Detection of ctDNA from Dried Blood Spots after DNA Size Selection. Clin Chem 66, 697-705 (2020).
[0180] S. Heinz, C. Benner, N. Spann, E. Bertolino, Y. C. Lin, P. Laslo, J. X. Cheng, C. Murre, H.
[0181] Singh, C. K. Glass, Simple Combinations of Lineage-Determining Transcription Factors Prime-Regulatory Elements Required for Macrophage and B Cell Identities. Mol Cell 38, 576-589 (2010).
[0182] P. Jiang, K. Sun, W. Peng, S. H. Cheng, M. Ni, P. C. Yeung, M. M. S. Heung, T. Xie, H. Shang, Z. Zhou, R. W. Y. Chan, J. Wong, V. W. S. Wong, L. C. Poon, T. Y. Leung, W. K. J. Lam, J. Y. K. Chan, H. L. Y. Chan, K. C. A. Chan, R. W. K. Chiu, Y. M. D. Lo, Plasma DNA End-Motif Profiling as a Fragmentomic Marker in Cancer, Pregnancy, and Transplantation. Cancer Discov 10, 664-673 (2020).
[0183] Liu Y. At the dawn: cell-free DNA fragmentomics and gene regulation. Br J Cancer. 2022 Feb;126(3):379-390.
[0184] Y. M. Lo, K. C. Chan, H. Sun, E. Z. Chen, P. Jiang, F. M. Lun, Y. W. Zheng, T. Y. Leung, T. K. Lau, C. R. Cantor, R. W. Chiu, Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the fetus. Sci Transl Med 2, 61ra91 (2010).
[0185] H. Markus, D. Chandrananda, E. Moore, F. Mouliere, J. Morris, J. D. Brenton, C. G. Smith, N. Rosenfeld, Refined characterization of circulating tumor DNA through biological feature integration. Sci Rep 12, 1928 (2022).
[0186] H. Markus, T. Contente-Cuomo, M. Farooq, W. S. Liang, M. J. Borad, S. Sivakumar, S. Gollins, N. L. Tran, H. D. Dhruv, M. E. Berens, A. Bryce, A. Sekulic, A. Ribas, J. M. Trent, P. M. LoRusso, M. Murtaza, Evaluation of pre-analytical factors affecting plasma DNA analysis. Sci Rep 8, 7375 (2018). H. Markus, J. Zhao, T. Contente-Cuomo, M. D. Stephens, E. Raupach, A. Odenheimer- Bergman, S. Connor, B. R. McDonald, B. Moore, E. Hutchins, M. McGilvrey, M. C. de la Maza, K. Van Keuren-Jensen, P. Pirrotte, A. Goel, C. Becerra, D. D. Von Hoff,
[0187] S. A. Celinski, P. Hingorani, M. Murtaza, Analysis of recurrently protected genomic regions in cell-free DNA found in urine. Sci Transl Med 13, (2021).
[0188] P. J. Mazzone, P. B. Bach, J. Carey, C. A. Schonewolf, K. Bognar, M. S. Ahluwalia, M. Cruz- Correa, D. Gierada, S. Kotagiri, K. Lloyd, F. Maldonado, J. D. Ortendahl, L. V. Sequist, G. A. Silvestri, N. Tanner, J. C. Thompson, A. Vachani, K. K. Wong, A. H. Zaidi, J. Catallini, A. Gershman, K. Lumbard, L. K. Millberg, J. Nawrocki, C. Portwood, A. Rangnekar, C. C. Sheridan, N. Trivedi, T. Wu, Y. Zong, L. Cotton, A. Ryan, C. Cisar, A. Leal, N. Dracopoli, R. B. Scharpf, V. E. Velculescu, L. R. G. Pike, Clinical Validation of a Cell-Free DNA Fragmentome Assay for Augmentation of Lung Cancer Early Detection. Cancer Discov 14, 2224-2242 (2024).
[0189] B. R. McDonald, T. Contente-Cuomo, S. J. Sammut, A. Odenheimer-Bergman, B. Ernst, N. Perdigones, S. F. Chin, M. Farooq, R. Mejia, P. A. Cronin, K. S. Anderson, H. E. Kosiorek, D. W. Northfelt, A. E. McCullough, B. K. Patel, J. N. Weitzel, T. P. Slavin, C. Cal das, B. A. Pockaj, M. Murtaza, Personalized circulating tumor DNA analysis to detect residual disease after neoadjuvant therapy in breast cancer. Sci Transl Med 11, (2019).
[0190] J. E. Medina, A. V. Annapragada, P. Lof, S. Short, A. L. Bartolomucci, D. Mathios, S. Koul, N. Niknafs, M. Noe, Z. H. Foda, D. C. Bruhm, C. Hruban, N. A. Vulpescu, E. Jung, R. Dua, J. V. Canzoniero, S. Cristiano, V. Adleff, H. Symecko, D. van den Broek, L. J. Sokoll, S. B. Baylin, M. F. Press, D. J. Slamon, G. E. Konecny, C. Therkildsen, B. Carvalho, G. A. Meijer, C. L. Andersen, S. M. Domchek, R. Drapkin, R. B. Scharpf, J. Phallen, C. A. R. Lok, V. E. Velculescu, Early detection of ovarian cancer using cell-free DNA fragmentomes and protein biomarkers. Cancer Discov, (2024).
[0191] Morbioli GG, Baillargeon KR, Kalimashe MN, Kana V, Zwane H, van der Walt C, Tierney
[0192] AJ, Mora AC, Goosen M, Jagaroo R, Brooks JC, Cutler E, Hunt G, Jordan MR, Tang A, Mace CR. Clinical evaluation of patterned dried plasma spot cards to support quantification of HIV viral load and reflexive genotyping. Proc Natl Acad Sci U S A. 2025 Feb 18; 122(7): e2419160122.
[0193] T. Moser, S. Kuhberger, I. Lazzeri, G. Vlachos, E. Heitzer, Bridging biological cfDNA features and machine learning approaches. Trends Genet 39, 285-307 (2023).
[0194] A. R. Quinlan, I. M. Hall, BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842 (2010). J. P. Radich, E. Briercheck, D. T. Chiu, M. P. Menon, O. Sala Torra, C. C. S. Yeung, E. H.
[0195] Warren, Precision Medicine in Low- and Middle-Income Countries. Annu Rev Pathol 17, 387-402 (2022).
[0196] C. M. Sauer, K. Heider, J. Belie, S. E. Boyle, J. A. Hall, D. L. Couturier, A. An, A.
[0197] Vijayaraghavan, M. A. Reinius, K. Hosking, M. Vias, N. Rosenfeld, J. D. Brenton, Longitudinal monitoring of disease burden and response using ctDNA from dried blood spots in xenograft models. EMBO Mol Med 14, el 5729 (2022).
[0198] D. Schrag, T. M. Beer, C. H. McDonnell, 3rd, L. Nadauld, C. A. Dilaveri, R. Reid, C. R.
[0199] Marinac, K. C. Chung, M. Lopatin, E. T. Fung, E. A. Klein, Blood-based tests for multicancer early detection (PATHFINDER): a prospective cohort study. Lancet 402, 1251-1260 (2023).
[0200] K. Sun, P. Jiang, S. H. Cheng, T. H. T. Cheng, J. Wong, V. W. S. Wong, S. S. M. Ng, B. B.
[0201] Y. Ma, T. Y. Leung, S. L. Chan, T. S. K. Mok, P. B. S. Lai, H. L. Y. Chan, H. Sun, K. C. A. Chan, R. W. K. Chiu, Y. M. D. Lo, Orientation-aware plasma cell-free DNA fragmentation analysis in open chromatin regions informs tissue of origin. Genome Res 29, 418-427 (2019).
[0202] Y. van der Pol, F. Mouliere, Toward the Early Detection of Cancer by Decoding the Epigenetic and Environmental Fingerprints of Cell-Free DNA. Cancer Cell 36, 350-368 (2019).
[0203] J. C. M. Wan, C. Massie, J. Garcia-Corbacho, F. Mouliere, J. D. Brenton, C. Caldas, S. Pacey, R. Baird, N. Rosenfeld. Liquid biopsies come of age: towards implementation of circulating tumour DNA. Nat Rev Cancer 17, 223-238 (2017).
[0204] M. L. Waskom, seaborn: statistical data visualization. Journal of Open Source Software 6, 3021 (2021).
[0205] D. Wong, M. Tageldein, P. Luo, E. Ensminger, J. Bruce, L. Oldfield, H. Gong, N. W. Fischer, B. Laverty7, V. Subasri, S. Davidson, R. Khan, A. Villani, A. Shlien, R. H. Kim, D. Malkin, T. J. Pugh, Cell-free DNA from germline TP53 mutation carriers reflect cancer-like fragmentation patterns. Nat Commun 15, 7386 (2024).
Claims
CLAIMSWhat is claimed is:
1. A method comprising: providing a solid substrate compnsing a blood sample comprising nucleic acid; extracting the nucleic acid from the solid substrate to obtain extracted nucleic acid; and analyzing the extracted nucleic acid.
2. The method of claim 1, wherein the blood sample is a dried blood sample.
3. The method of any prior claim, wherein the blood sample is a cell-removed blood sample.
4. The method of claim 3, wherein the cell-removed blood sample comprises at least one of dried plasma and dried serum.
5. The method of any one of claims 3-4, comprising: obtaining a whole blood sample from a subject; removing cells from the whole blood to generate a liquid cell-removed blood sample; applying either the whole blood or the liquid cell-removed blood sample to the solid substrate; and then drying the liquid cell-removed blood sample to obtain the cell-removed blood sample.
6. The method of claim 5, wherein the whole blood or the liquid cell-removed blood sample is applied to the solid substrate in an amount from 10 pl to 600 pl.
7. The method of any one of claims 5-6, wherein the solid substrate comprises the liquid cell- removed blood sample, prior to the drying, in an amount from 10 pl to 600 pl.
8. The method of any prior claim, wherein the extracted nucleic acid comprises cell-free DNA.
9. The method of any pnor claim, wherein the extracted nucleic acid is substantially devoid of intact genomic DNA.
10. The method of any prior claim, wherein the solid substrate comprises one or more of nylon, polypropylene, polyester, rayon, cellulose, cellulose acetate, nitrocellulose, mixed cellulose ester, glass microfiber filters, cotton, quartz microfiber, polytetrafluoroethylene, wax-patterned cellulose, poly ethersulfone, and poly vinylidene fluoride.1 1. The method of any prior claim, wherein the extracting the nucleic acid from the solid substrate comprises: combining the solid substrate with a liquid solvent; incubating the solid substrate in the liquid solvent to extract the nucleic acid from the solid substrate into the liquid solvent and thereby obtain the extracted nucleic acid in the liquid solvent; and removing the liquid solvent from the solid substrate.
12. The method of any prior claim, wherein the extracted nucleic acid is in an amount from 10 pg to 25 ng.
13. The method of any prior claim, wherein the extracted nucleic acid has a modal size from 160 to 174 bp.
14. The method of any prior claim, wherein the analyzing the extracted nucleic acid comprises analyzing a sequence of the extracted nucleic acid.
15. The method of any prior claim, wherein the analyzing the extracted nucleic acid compnses determining a fragmentation pattern of the extracted nucleic acid.
16. The method of claim 15, wherein the determining the fragmentation pattern of the extracted nucleic acid comprises non-targeted sequencing of the nucleic acid.
17. The method of any one of claims 15-16, wherein the determining the fragmentation pattern of the extracted nucleic acid comprises whole-genome sequencing.
18. The method of any one of claims 15-17, wherein the determining the fragmentation pattern of the extracted nucleic acid comprises genome-wide fragmentation analysis.
19. The method of claim 15, wherein the determining the fragmentation pattern of the extracted nucleic acid comprises targeted amplification and / or sequencing of the nucleic acid.
20. The method of any one of claims 15-19, wherein the determining the fragmentation pattern of the extracted nucleic acid comprises determining a fragment size distribution, information- weighted fraction of aberrant fragments analysis, determining incorporation of nucleotide frequencies at fragment ends, determining incorporation of nucleotide frequencies surrounding fragment ends, determining ratios of fragment sizes across one or more portions of a genome, and determining motif frequencies at fragment ends.
21. The method of any prior claim, wherein the blood sample is from a subject suspected of having cancer.
22. The method of claim 21, wherein the determining the fragmentation pattern of the extracted nucleic acid detects an aberrant fragmentation pattern relative to control nucleic acid from a subject that does not have cancer.