Library preparation and analysis methods for preserving topological information of cell-free DNA
By extending cfDNA ends with inosine bases and soft-clipping sequencing adaptors, the method addresses the loss of topological information in conventional sequencing, enabling precise cancer detection.
Patent Information
- Application Number
- JP2025531212
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-16
- Filing Date
- 2023-12-15
- Publication Date
- 2026-01-30
AI Technical Summary
Conventional DNA sequencing methods fail to capture the topological information of cell-free DNA (cfDNA) due to the removal of uneven DNA ends during the end repair step, leading to a loss of critical information for cancer diagnosis and treatment.
A method for determining cfDNA topology by extending the ends of DNA duplexes with inosine bases to fill in overhangs, attaching sequencing adaptors, and soft-clipping the inosine extensions to determine the length and sequence of overhangs, enabling detection of diseases like cancer based on these features.
Preserves and analyzes the topological information of cfDNA, allowing for accurate detection of cancer through sequencing, thereby enhancing diagnostic capabilities.
Smart Images

Figure 2026503833000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 433,348, filed December 16, 2022, the entire contents of which are incorporated herein by reference for all purposes.
[0002] Disclosed herein are methods and systems for determining and analyzing the topological information of DNA duplexes, such as cell-free DNA (cfDNA), and constructing sequencing libraries for obtaining such topological information.Some aspects of the present disclosure more particularly relate to methods and systems for constructing DNA libraries for determining and / or analyzing the 5' and 3' overhangs from DNA duplexes (such as cfDNA). [Background technology]
[0003] Cell-free DNA (cfDNA) molecules are free-flowing double-stranded DNA molecules (dsDNA or duplex DNA) found in the bloodstream, typically as a result of cellular apoptosis or necrosis, particularly in disease settings such as cancer. These degraded linear DNA fragments are often approximately 50-300 base pairs in length. Most commonly, cfDNA is assayed to screen for cancer early in disease progression by analyzing cfDNA sequences to identify cancer-associated mutations.
[0004] Natural cfDNA often has DNA ends, either the 5'-end or the 3'-end protruding from the complementary strand, thereby making the ends uneven.Because standard DNA sequencing methods can only sequence DNA with blunt ends, conventional sequencing techniques involve the removal of these natural topologies during the end repair step of conventional double-stranded library construction.As a result, the opportunity to identify all the information about the length and sequence of the uneven ends present in cfDNA, natural end sequences, gaps, and nicks, and any possible relevance of this information for cancer diagnosis and cancer cure is lost. Summary of the Invention
[0005] Described herein are methods for determining cell-free DNA topology and methods for generating nucleic acid constructs for determining cell-free DNA topology. Also described are methods for detecting disease based at least in part on the determined cell-free DNA topology.
[0006] In some implementations, a method for determining cell-free DNA topology includes extending the 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill in a 5' overhang of a second strand of the cell-free DNA duplex, attaching a sequencing adaptor to the cell-free DNA duplex, sequencing the first strand of the cell-free DNA duplex to generate first sequence reads, determining, by one or more processors, one or more bases in the first sequence reads to be soft-clipped, soft-clipping, by the one or more processors, bases corresponding to the 3' inosine extension of the first strand of the cell-free DNA duplex from the first sequence reads, and determining, by the one or more processors, the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex. The method may further include detecting, by the one or more processors, the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex.
[0007] A method for determining cell-free DNA topology includes attaching a second sequencing adaptor to a 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to a 5' end of a first strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the first strand cell-free DNA duplex. and determining, by the one or more processors, one or more bases attached to the 5' end of the first sequence read to be soft-clipped, soft-clipping, by the one or more processors, bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence read, and determining, by the one or more processors, the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. In some embodiments, the method may further include detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex. In some embodiments, attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises extending the 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex, and attaching the second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0008] In some embodiments of the above method, determining one or more bases in the first sequence read to be soft-clipped includes aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read.
[0009] In some embodiments, the method for determining cell-free DNA topology further includes sequencing a second strand of the cell-free DNA duplex to generate second sequence reads and determining, by one or more processors, one or more bases in the soft-clipped second sequence reads. In some embodiments, determining one or more bases in the soft-clipped second sequence reads includes aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. In some embodiments, the first sequence read and the second sequence read are related via a unique molecular identifier (UMI). In some embodiments, determining one or more bases in the soft-clipped first sequence read includes aligning the first sequence to the second sequence read to identify unaligned portions of the first sequence read. In some embodiments, determining one or more bases in the soft-clipped second sequence read includes aligning the second sequence to the first sequence read to identify unaligned portions of the second sequence read.
[0010] In some embodiments of the above methods, the method further includes extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex, attaching a second sequencing adaptor to the cell-free DNA duplex, soft-clipping, by one or more processors, bases corresponding to the 3' inosine extension of the second strand of the cell-free DNA duplex from the second sequence reads, and determining, by the one or more processors, the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the second strand of the cell-free DNA duplex. In some embodiments, the method further includes detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex.
[0011] In some embodiments of the above method, extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a 3' single inosine overhang. In some embodiments, the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the 3' inosine overhang.
[0012] In some embodiments, the method for determining cell-free DNA topology further comprises amplifying the cell-free DNA duplex and binding the sample index to the cell-free DNA duplex.
[0013] In some embodiments of the above methods, the sequencing adaptor or the second sequencing adaptor comprises a sample index.
[0014] In some embodiments of the above method, the sequencing adaptor or the second sequencing adaptor comprises a UMI.
[0015] In some embodiments of the above methods, the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. In some embodiments, the adaptor can be a Y-full-length adaptor (e.g., a Y-shaped adaptor including an index sequence), a stubby adaptor, or a hairpin adaptor.
[0016] Also provided herein are methods of making a sequencing construct, the methods comprising: attaching a sequencing adaptor to a 3' overhang of a first strand of a cell-free DNA duplex; extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the sequencing adaptor to the 5' end of a second strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the second strand of the cell-free DNA duplex. In some embodiments, attaching a sequencing adaptor to the 3' overhang of the first strand of the cell-free DNA duplex comprises extending the 3' overhang of the first strand of the cell-free DNA duplex to provide a 3' extension, wherein the sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the first strand of the cell-free DNA duplex, and attaching the sequencing adaptor to the 3' extension of the first strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the first strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0017] Further described herein are methods for determining cell-free DNA topology, the methods comprising: generating a sequencing construct according to the methods described above; sequencing the second strand of the cell-free DNA duplex to generate first sequence reads; determining, by one or more processors, one or more bases in the first sequence reads to be soft-clipped; soft-clipping, by the one or more processors, bases corresponding to the 5' inosine stretch of the second strand of the cell-free DNA duplex from the first sequence reads; and determining, by the one or more processors, the length or sequence of a 3' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the second strand of the cell-free DNA duplex. In some embodiments, the methods further comprise detecting, by the one or more processors, the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex. In some embodiments, determining one or more bases in the first sequence read to be soft-clipped includes aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read.
[0018] In some embodiments of the method for determining cell-free DNA topology, the method further includes sequencing a first strand of the cell-free DNA duplex to generate second sequence reads and determining, by one or more processors, one or more bases in the soft-clipped second sequence reads. In some embodiments, determining one or more bases in the soft-clipped second sequence reads includes aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. In some embodiments, the first sequence read and the second sequence read are related via a unique molecular identifier (UMI). In some embodiments, determining one or more bases in the soft-clipped first sequence read includes aligning the first sequence to the second sequence read and identifying unaligned portions of the first sequence read. In some embodiments, determining one or more bases in the soft-clipped second sequence read includes aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read.
[0019] In some embodiments of the methods for determining cell-free DNA topology, the method further comprises attaching a second sequencing adaptor to a 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to a 5' end of the first strand of the cell-free DNA duplex, wherein the 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, the method further includes sequencing a first strand of the cell-free DNA duplex to generate second sequence reads, determining, by one or more processors, one or more bases attached to the 5' end of the second sequence reads to be soft-clipped, soft-clipping, by one or more processors, bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the second sequence reads, and determining, by one or more processors, the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. In some embodiments, the method further includes detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex. In some embodiments, attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises attaching a single inosine to the 3' overhang of the second strand of the cell-free DNA duplex and attaching a second sequencing adaptor to the single inosine. In some embodiments, the sequencing adaptor comprises a 3' cytosine overhang complementary to the inosine attached to the 3' overhang of the second strand of the cell-free DNA duplex. In some embodiments, determining one or more bases attached to the 5' end of the soft-clipped second sequence read comprises aligning the second sequence read to a reference sequence to identify unaligned portions of the second sequence read. In some embodiments, the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI).In some embodiments, determining one or more bases attached to the 5' end of the second sequence read to be soft clipped comprises aligning the second sequence read to the first sequence read and identifying an unaligned portion of the second sequence read. In some embodiments, determining one or more bases attached to the 5' end of the first sequence read to be soft clipped comprises aligning the first sequence read to the second sequence read and identifying an unaligned portion of the first sequence read.
[0020] In some embodiments of the above methods, the cell-free DNA duplexes are amplified and the sample index is attached to the cell-free DNA duplexes.
[0021] In some embodiments of the above methods, the sequencing adaptor or the second sequencing adaptor comprises a sample index.
[0022] In some embodiments of the above method, the sequencing adaptor or the second sequencing adaptor comprises a UMI.
[0023] In some embodiments of the above methods, the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. In some embodiments, the adaptor can be a Y-full-length adaptor (e.g., a Y-shaped adaptor including an index sequence), a stubby adaptor, or a hairpin adaptor.
[0024] In some embodiments of the above methods, the cell-free DNA duplexes are obtained from a subject suspected of having or determined to have cancer.
[0025] In some embodiments of any of the above methods, the method further comprises treating the subject with an anti-cancer therapy.
[0026] In some embodiments of any of the above methods, the method further comprises obtaining the cell-free DNA duplex from the subject.
[0027] In some embodiments of any of the above methods, the method further comprises obtaining cell-free DNA duplexes from a liquid biopsy sample. In some embodiments, the sample is a liquid biopsy sample, including blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the cell-free DNA duplexes are circulating tumor DNA (ctDNA) duplexes.
[0028] In some embodiments of any of the above methods, the sequencing adaptor or the second sequencing adaptor comprises an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence.
[0029] In some embodiments of any of the above methods, the method further comprises amplifying the first strand and the second strand of the cell-free DNA duplex, hi some embodiments, amplifying comprises performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0030] In some embodiments of any of the above methods, the sequencing comprises the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. In some embodiments, the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technology comprises next-generation sequencing (NGS). In some embodiments, the sequencing is performed using a next-generation sequencer.
[0031] In some embodiments of any of the above methods, the method further includes generating, by the one or more processors, a report indicating the length of the 3' overhang of the first strand of the cell-free DNA duplex, the length of the 3' overhang of the second strand of the cell-free DNA duplex, the length of the 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of the 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the method further includes transmitting the report to a health care provider. In some embodiments, the report is transmitted over a computer network or a peer-to-peer connection.
[0032] In some embodiments of any of the above methods, the method further comprises generating a genomic profile of the subject, the genomic profile comprising the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of a 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of a 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the subject's genomic profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the subject's genomic profile further comprises results from a nucleic acid sequencing-based test.
[0033] Also described herein is a system that includes one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, direct the system to: receive a first sequence read obtained by extending a 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill a 5' overhang of a second strand of the cell-free DNA duplex; attach a sequencing adaptor to the cell-free DNA duplex; sequence the first strand of the cell-free DNA duplex to generate a first sequence read; determine one or more bases in the first sequence read to be soft-clipped; soft-clip a base from the first sequence read that corresponds to the 3' inosine extension of the first strand of the cell-free DNA duplex; and determine the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, the instructions further instruct the system to detect the presence or absence of a disease (e.g., cancer) based on the length or sequence of a 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the first sequence read is further obtained by attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex, extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex, and attaching the 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex.
[0034] In some embodiments of the above system, the instructions, when executed by the one or more processors, further direct the system to: soft-clip bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence read; and determine the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. In some embodiments, the instructions further direct the system to detect the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex. In some embodiments, attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises extending the 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex, and attaching the second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0035] In some embodiments of the above system, one or more bases in the first sequence read are determined to be soft-clipped by a method that includes aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read.
[0036] In some embodiments of the above system, the instructions, when executed by one or more processors, further direct the system to receive a second sequence read obtained by sequencing a second strand of the cell-free DNA duplex and determine one or more bases in the second sequence read to be soft-clipped. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence and identifying unaligned portions of the second sequence read. In some embodiments, the first sequence read and the second sequence read are related via a unique molecular identifier (UMI). In some embodiments, one or more bases in the first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to the second sequence read and identifying unaligned portions of the first sequence read. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. In some embodiments, the second sequence read is further obtained by extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex and ligating a second sequencing adaptor to the cell-free DNA duplex. In some embodiments, the instructions, when executed by the one or more processors, further direct the system to: soft-clip bases corresponding to the 3' inosine stretch of the second strand of the cell-free DNA duplex from the second sequence read; and determine the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine stretch of the second strand of the cell-free DNA duplex.
[0037] In some embodiments of the above system, extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a 3' single inosine overhang. In some embodiments, the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the 3' single inosine overhang.
[0038] Further described herein is a system comprising one or more processors and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive a first sequence read obtained by ligating a sequencing adaptor to a 3' overhang of a first strand of a cell-free DNA duplex; extend the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and ligate the 3' inosine-extended end of the sequencing adaptor to the cell-free DNA duplex. and memory that directs the system to: bind to the 5'-end of the second strand of the heavy strand, wherein the bound 3' inosine extension provides a 5' inosine extension of the second strand of the cell-free DNA duplex; determine one or more bases in the first sequence read to be soft-clipped; soft-clip bases from the first sequence read that correspond to the 5' inosine extension of the second strand of the cell-free DNA duplex; and determine a length or sequence of a 3' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine extension of the second strand of the cell-free DNA duplex. In some embodiments, the instructions further direct the system to detect the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex.
[0039] In some embodiments of the above system, attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises extending the 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; and attaching the second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0040] In some embodiments of the above system, one or more bases in the first sequence read are determined to be soft-clipped by a method that includes aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read.
[0041] In some embodiments of the above system, the instructions, when executed by one or more processors, further direct the system to receive a second sequence read obtained by sequencing a second strand of the cell-free DNA duplex and determine one or more bases in the second sequence read to be soft-clipped. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence and identifying unaligned portions of the second sequence read. In some embodiments, the first sequence read and the second sequence read are related via a unique molecular identifier (UMI). In some embodiments, one or more bases in the first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to the second sequence read and identifying unaligned portions of the first sequence read. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. In some embodiments, the second sequence read is further obtained by attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex, extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex, and attaching the 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, the instructions, when executed by the one or more processors, further direct the system to soft-clip bases corresponding to the 3' inosine extension of the first strand of the cell-free DNA duplex from the second sequence read, and determine the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex.In some embodiments, attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises attaching a single inosine to the 3' overhang of the second strand of the cell-free DNA duplex and attaching the second sequencing adaptor to the single inosine. In some embodiments, the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the single inosine attached to the 3' overhang of the second strand of the cell-free DNA duplex.
[0042] In some embodiments of any of the above systems, the sequencing adaptor or the second sequencing adaptor comprises a UMI.
[0043] In some embodiments of any of the above systems, the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. In some embodiments, the adaptor can be a Y full-length adaptor (e.g., a Y-shaped adaptor including an index sequence), a stubby adaptor, or a hairpin adaptor.
[0044] In some embodiments of any of the above systems, the system further comprises a nucleic acid amplifier configured to amplify the cell-free DNA duplexes and bind the sample index to the cell-free DNA duplexes. In some embodiments, the nucleic acid amplifier is a thermal cycler.
[0045] In some embodiments of any of the above systems, the cell-free DNA duplexes are obtained from a subject suspected of having or determined to have cancer.
[0046] In some embodiments of any of the above systems, the cell-free DNA duplexes are obtained from a subject.
[0047] In some embodiments of any of the above systems, the cell-free DNA duplexes are obtained from a liquid biopsy sample.
[0048] In some embodiments of any of the above systems, the sample is a liquid biopsy sample and includes blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva.
[0049] In some embodiments of any of the above systems, the cell-free DNA duplexes are circulating tumor DNA (ctDNA) duplexes.
[0050] In some embodiments of any of the above systems, the sequencing adaptor or the second sequencing adaptor comprises an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence.
[0051] In some embodiments of any of the above systems, the system further includes a nucleic acid amplifier configured to amplify the first strand and the second strand of the cell-free DNA duplex. In some embodiments, amplifying includes performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0052] In some embodiments of any of the above systems, the system further comprises a sequencer configured to sequence the first strand of the cell-free DNA duplex and / or the second strand of the cell-free DNA duplex. In some embodiments, the sequencer is configured for massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. In some embodiments, the sequencer is configured for massively parallel sequencing, and the massively parallel sequencing technology comprises next-generation sequencing (NGS). In some embodiments, the sequencer is a next-generation sequencer.
[0053] In some embodiments of any of the above systems, the instructions, when executed by the one or more processors, further direct the system to generate a report indicating the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of the 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of the 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the instructions, when executed by the one or more processors, further direct the system to transmit the report to a healthcare provider. In some embodiments, the report is transmitted over a computer network or a peer-to-peer connection.
[0054] In some embodiments of any of the above systems, the instructions, when executed by the one or more processors, further direct the system to generate, by the one or more processors, a genomic profile of the subject comprising the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of a 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of a 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the subject's genomic profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the subject's genomic profile further comprises results from a nucleic acid sequencing-based test.
[0055] Also described herein are non-transitory computer-readable storage media storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the system, direct the system to: receive a first sequence read obtained by extending the 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill a 5' overhang of a second strand of the cell-free DNA duplex; attach a sequencing adapter to the cell-free DNA duplex; sequence the first strand of the cell-free DNA duplex to generate a first sequence read; determine one or more bases in the first sequence read to be soft-clipped; soft-clip a base from the first sequence read that corresponds to the 3' inosine extension of the first strand of the cell-free DNA duplex; and determine the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, the instructions further instruct the system to detect the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex.
[0056] In some embodiments of the above storage medium, the first sequence read is further obtained by attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, the instructions, when executed by the one or more processors, further direct the system to: soft-clip bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence read; and determine the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. In some embodiments, the instructions further direct the system to detect the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex.
[0057] In some embodiments of the above-described storage medium, attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises extending the 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; and attaching the second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0058] In some embodiments of the above storage medium, one or more bases in the first sequence read are determined to be soft-clipped by a method including aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read.
[0059] In some embodiments of the above storage medium, the instructions, when executed by one or more processors, further direct the system to receive a second sequence read obtained by sequencing a second strand of the cell-free DNA duplex and determine one or more bases in the second sequence read to be soft-clipped. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence and identifying unaligned portions of the second sequence read. In some embodiments, the first sequence read and the second sequence read are related via a unique molecular identifier (UMI). In some embodiments, one or more bases in the first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to the second sequence read and identifying unaligned portions of the first sequence read. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. In some embodiments, the second sequence read is further obtained by extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex and ligating a second sequencing adaptor to the cell-free DNA duplex. In some embodiments, the instructions, when executed by the one or more processors, further direct the system to: soft-clip bases corresponding to the 3' inosine stretch of the second strand of the cell-free DNA duplex from the second sequence read; and determine the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine stretch of the second strand of the cell-free DNA duplex.
[0060] In some embodiments of the above storage medium, extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a 3' single inosine overhang. In some embodiments, the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the 3' single inosine overhang.
[0061] Further described herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the system, cause the system to: receive a first sequence read obtained by ligating a sequencing adaptor to a 3' overhang of a first strand of a cell-free DNA duplex; extend the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and ligate the 3' inosine-extended end of the sequencing adaptor to the 3' overhang of the first strand of the cell-free DNA duplex. The instructions instruct the system to: bind to the 5' end of the second strand, wherein the bound 3' inosine extension provides a 5' inosine extension of the second strand of the cell-free DNA duplex; determine one or more bases in the first sequence read to be soft-clipped; soft-clip bases from the first sequence reads corresponding to the 5' inosine extension of the second strand of the cell-free DNA duplex; and determine a length or sequence of a 3' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine extension of the second strand of the cell-free DNA duplex. In some embodiments, the instructions further instruct the system to detect the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex.
[0062] In some embodiments of the above-described storage medium, attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises extending the 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; and attaching the second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0063] In some embodiments of the above storage medium, one or more bases in the first sequence read are determined to be soft-clipped by a method including aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read.
[0064] In some embodiments of the above storage medium, the instructions, when executed by one or more processors, further direct the system to receive a second sequence read obtained by sequencing a second strand of the cell-free DNA duplex and determine one or more bases in the second sequence read to be soft-clipped. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence and identifying unaligned portions of the second sequence read. In some embodiments, the first sequence read and the second sequence read are related via a unique molecular identifier (UMI). In some embodiments, one or more bases in the first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to the second sequence read and identifying unaligned portions of the first sequence read. In some embodiments, one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. In some embodiments, the second sequence read is further obtained by attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex, extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex, and attaching the 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, the instructions, when executed by the one or more processors, further direct the system to soft-clip bases corresponding to the 3' inosine extension of the first strand of the cell-free DNA duplex from the second sequence read, and determine the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex.In some embodiments, attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises attaching a single inosine to the 3' overhang of the second strand of the cell-free DNA duplex and attaching the second sequencing adaptor to the inosine. In some embodiments, the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the single inosine attached to the 3' overhang of the second strand of the cell-free DNA duplex.
[0065] In some embodiments of the above storage medium, the sequencing adaptor or the second sequencing adaptor comprises a UMI.
[0066] In some embodiments of the above storage medium, the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. In some embodiments, the adaptor can be a Y-full-length adaptor (e.g., a Y-shaped adaptor including an index sequence), a stubby adaptor, or a hairpin adaptor.
[0067] In some embodiments of the above storage medium, the cell-free DNA duplexes are obtained from a subject suspected of having or determined to have cancer.
[0068] In some embodiments of the above storage medium, the cell-free DNA duplex is obtained from a subject.
[0069] In some embodiments of the above storage medium, the cell-free DNA duplexes are obtained from a liquid biopsy sample. For example, in some embodiments, the sample is a liquid biopsy sample, including blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva.
[0070] In some embodiments of the above storage medium, the cell-free DNA duplexes are circulating tumor DNA (ctDNA) duplexes.
[0071] In some embodiments of the above storage medium, the sequencing adaptor or the second sequencing adaptor comprises an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence.
[0072] In some embodiments of the above storage medium, the sequencing comprises using massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. In some embodiments, the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technology comprises next-generation sequencing (NGS). In some embodiments, the sequencing is performed using a next-generation sequencer.
[0073] In some embodiments of the above storage medium, the instructions, when executed by the one or more processors, further direct the system to generate a report indicating the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of the 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of the 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the instructions, when executed by the one or more processors, further direct the system to transmit the report to a healthcare provider. In some embodiments, the report is transmitted over a computer network or a peer-to-peer connection.
[0074] In some embodiments of the above-described storage medium, the instructions, when executed by the one or more processors, further instruct the system to generate, by the one or more processors, a genomic profile of the subject comprising the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of a 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of a 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the subject's genomic profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the subject's genomic profile further comprises results from a nucleic acid sequencing-based test. [Brief explanation of the drawings]
[0075] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of exemplary embodiments and the accompanying drawings.
[0076] [Figure 1A] 1 shows an exemplary process for generating a sequencing construct and determining the length of a 5' overhang from a DNA duplex molecule (e.g., a cfDNA molecule), according to some embodiments. [Figure 1B] 1A shows an exemplary embodiment of the process for creating a sequencing construct and determining the length of 5' overhang from a DNA double-stranded molecule (e.g., cfDNA molecule).In this exemplary method, the 5' overhang of cfDNA is filled with deoxyinosine, and then an adapter is attached to the end of cfDNA.Optionally, cfDNA is PCR amplified before sequencing.Then, the sequences are aligned and analyzed for overhang length and sequence. [Figure 2A]1 shows an exemplary method, according to some embodiments, for generating a sequencing construct useful for determining the length of a 3' overhang from a DNA duplex, such as a cfDNA molecule. [Figure 2B] 1 shows an exemplary method for determining the length of a 3' overhang from a DNA duplex, such as a cfDNA molecule, according to some embodiments. [Figure 3A]
[0023] Figure 1 shows an exemplary method for generating a sequencing construct useful for determining the length of 5' and / or 3' overhangs from DNA duplex molecules (such as cfDNA molecules), as well as a process for determining the length of 5' and / or 3' overhangs of DNA duplex molecules, according to some embodiments. [Figure 3B] 3A shows an exemplary embodiment of a method for creating a sequencing construct and determining the length of a 5' and / or 3' overhang from a DNA double-stranded molecule (such as a cfDNA molecule), according to some embodiments. In this exemplary method, the 5' overhang in the cfDNA is filled with deoxyinosine. One adapter is attached to the 3' overhang of the bottom strand, and another adapter is attached to the opposite end of the molecule. The 3' overhang of the cfDNA is then filled with deoxyinosine. Optionally, the cfDNA is PCR-amplified before sequencing. The sequences are then aligned and analyzed for overhang length and sequence. [Figure 3C]
[0023] Figure 3A illustrates another exemplary embodiment of a method for generating a sequencing construct and determining the length of 5' and / or 3' overhangs from DNA duplex molecules (such as cfDNA molecules), according to some embodiments. Aspects of the method illustrated in Figure 3C can be readily applied to other embodiments of the methods described herein, including the processes illustrated in Figures 1A, 1B, 2A, and 2B. In the exemplary method, 5' overhangs in cfDNA are filled with deoxyinosine, and then a first adaptor is ligated to the deoxyinosine-filled end of the cfDNA. The 3' overhang end in the cfDNA is extended with a single type of nucleotide, represented by "Y," and a second adaptor is ligated to the cfDNA, filling the 3' overhang of the cfDNA with deoxyinosine. Optionally, the cfDNA is PCR-amplified before sequencing. The sequences are then aligned and analyzed for overhang length and sequence. [Figure 4] 1 illustrates a process used to determine cfDNA topology, according to some embodiments. [Figure 5] 1 shows an exemplary read processing matrix for extracting data from paired-end sequencing, according to some embodiments. [Figure 6] 1 illustrates an exemplary computing device or system according to one embodiment of the present disclosure. [Figure 7] 1 illustrates an exemplary computer system or computer network in accordance with some examples of the systems described herein. DETAILED DESCRIPTION OF THE INVENTION
[0077] Described herein are methods and systems for determining the topology of DNA duplex (e.g., cell-free DNA) molecules, specifically the length and sequence of 5' and / or 3' overhangs. Standard methods for assessing cfDNA fail to adequately capture information about the 5' and 3' overhangs of cfDNA to provide a complete description of the topological information or characteristics of cfDNA molecules. Such information may provide information about the onset, progression, diagnosis, treatment, etc. of various diseases, specifically cancer. This loss of information poses an obstacle not only to understanding disease processes but also to utilizing this information to improve human health (e.g., early detection of diseases such as cancer, correlation with treatment efficacy, etc.).
[0078] The methods and systems described herein provide for the generation and analysis of sequencing libraries that capture 5' and / or 3' overhangs that contain the length and / or sequence of the natural ends of DNA duplex molecules. Specifically, the methods and systems described herein can further generate and analyze sequence libraries from cfDNA collected from patients with early-stage cancer or suspected cancer. Thus, the methods and systems described herein may provide a solution to the limitations of other methods for analyzing cfDNA, allowing for better capture and utilization of the information present in cfDNA molecules.
[0079] Thus, in one aspect, a method for determining cell-free DNA topology includes extending the 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill in a 5' overhang of a second strand of the cell-free DNA duplex; attaching a sequencing adaptor to the cell-free DNA duplex; sequencing the first strand of the cell-free DNA duplex to generate first sequence reads; determining, by one or more processors, one or more bases in the first sequence reads to be soft-clipped; soft-clipping, by the one or more processors, bases corresponding to the 3' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence reads; and determining, by the one or more processors, the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine stretch of the first strand of the cell-free DNA duplex. The method includes attaching a second sequencing adaptor to a 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to a 5' end of the first strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. The method may further include determining, by the one or more processors, one or more bases attached to the 5' end of the first sequence read that are soft-clipped, soft-clipping, by the one or more processors, bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence read, and determining, by the one or more processors, the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. In some embodiments, the method further includes detecting, by the one or more processors, the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 5' overhang of the first strand and / or the second strand of the cell-free DNA duplex.In some embodiments, the method further includes detecting, by the one or more processors, the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 3' overhang of the first strand and / or the second strand of the cell-free DNA duplex.
[0080] In another aspect, there is a method of making a sequencing construct, the method comprising: attaching a sequencing adaptor to a 3' overhang of a first strand of a cell-free DNA duplex; extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the sequencing adaptor to a 5' end of a second strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the second strand of the cell-free DNA duplex. The method further includes analyzing the sequencing construct to determine the topology of the cell-free DNA duplex, e.g., by sequencing the second strand of the cell-free DNA duplex to generate first sequence reads, determining, by one or more processors, one or more bases in the first sequence reads to be soft-clipped, soft-clipping, by one or more processors, bases corresponding to the 5' inosine stretch of the second strand of the cell-free DNA duplex from the first sequence reads, and determining, by one or more processors, the length or sequence of a 3' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the second strand of the cell-free DNA duplex. In some embodiments, the method further includes detecting, by one or more processors, the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 5' overhang of the first strand and / or the second strand of the cell-free DNA duplex. In some embodiments, the method further includes detecting, by the one or more processors, the presence or absence of a disease (e.g., cancer) based on the length or sequence of the 3' overhang of the first strand and / or the second strand of the cell-free DNA duplex.
[0081] Also described are systems and computer-readable storage media for performing the methods described herein.
[0082] definition Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0083] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Any reference to "or" herein is intended to include "and / or" unless specifically stated otherwise.
[0084] "About" and "approximately" are intended to mean an acceptable degree of error for the quantity measured, generally given the nature or precision of the measurements. Exemplary degrees of error are within 20 percent (%), typically within 10%, and more typically within 5% of a given value or range of values.
[0085] As used herein, the terms "comprising" (and any form or variation of "comprising", such as "comprise" and "comprises"), "having" (and any form or variation of "having", such as "have" and "has"), "including" (and any form or variation of "includes" and "include"), or "containing" (and any form or variation of "containing", such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unlisted additives, components, integers, elements, or method steps.
[0086] As used herein, the terms "individual," "patient," or "subject" are used interchangeably and refer to any single animal, e.g., mammal (including such non-human animals as dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non-human primates), for which treatment is desired. In certain embodiments, an individual, patient, or subject herein is a human.
[0087] The terms "cancer" and "tumor" are used interchangeably herein. These terms refer to the presence of cells that have characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of tumors, but such cells can exist alone in an animal or can be non-tumorigenic cancer cells, such as leukemia cells. These terms include solid tumors, soft tissue tumors, or metastatic lesions. As used herein, the term "cancer" includes precancerous as well as malignant cancers.
[0088] As used herein, "treatment" (and grammatical variations such as "treat" or "treating") refers to a clinical intervention (e.g., administration of an anti-cancer agent or anti-cancer therapy) that seeks to alter the natural course of the individual being treated. Desirable effects of treatment include, but are not limited to, preventing recurrence of the disease, alleviating symptoms, reducing any direct or indirect pathological consequences of the disease, preventing metastasis, reducing the rate of disease progression, ameliorating or palliation of the disease state, and remission or improved prognosis.
[0089] When a range of values is provided, it is understood that each intervening value between the upper and lower limits of that range, and any other stated or intervening value within that stated range, is included within the scope of the disclosure. Where the stated range includes an upper or lower limit, ranges excluding either of those included limits are also included in the disclosure.
[0090] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be readily apparent to those skilled in the art, and the generic principles herein may be applied to other embodiments. Thus, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.
[0091] 1-7 illustrate processes and systems according to various embodiments. In the exemplary processes, some blocks may be arbitrarily combined, the order of some blocks may be arbitrarily changed, and some blocks may be arbitrarily omitted. In some instances, additional steps may be performed in combination with the exemplary processes. Accordingly, the operations illustrated (and described in more detail below) are exemplary in nature and, therefore, should not be considered limiting.
[0092] The disclosures of all publications, patents, and patent applications referenced herein are each incorporated herein by reference in their entirety. To the extent that a reference incorporated by reference conflicts with the present disclosure, the present disclosure shall control. Method for determining the topology of a protrusion
[0093] Provided herein are methods for generating sequencing constructs for DNA duplexes (e.g., cfDNA molecules) and analyzing the sequencing constructs to determine the topology of the DNA duplex molecules. For example, the ends of cfDNA can include 5' and / or 3' overhangs, and the length and / or sequence of such overhangs can be determined according to the methods described herein. For example, the method can include filling gaps in the overhangs using nucleotides such as inosine nucleotides (e.g., deoxyinosine) or analogs thereof. Adapters can then be attached to the ends of the cfDNA duplex molecules, and the first and / or second strands can be sequenced, processed (e.g., soft-clipped inosine bases), and analyzed for cfDNA topology, including the length and sequence of the 5' and / or 3' overhangs.
[0094] As further described herein, the method can include generating a sequencing construct for a DNA duplex molecule (e.g., cfDNA). The sequencing construct can be generated so that the 5' overhang can be analyzed (e.g., length and / or sequence determined), the 3' overhang can be analyzed (e.g., length and / or sequence determined), or both the 5' overhang and the 3' overhang can be analyzed. Although certain embodiments are described with reference to a single DNA duplex molecule (e.g., a single cfDNA duplex), it is understood that the method can be applied to the construction of a sequencing library in which multiple DNA duplex molecules, which may have different topologies, can be generated and / or analyzed in parallel.
[0095] In some instances, one or both strands of the cfDNA duplex may contain 5' and / or 3' overhangs. Sequencing construct generation and analysis methods utilize nucleic acid analogs, such as deoxyinosine (dI), to fill gaps in the 5' and / or 3' overhangs, thereby preserving the sequence and length data of the overhangs for further downstream applications. Deoxyinosine can be incorporated or added by polymerases during extension, including, but not limited to, standard or high-fidelity preparations of Taq polymerase and its variants (e.g., LongAmp Taq, Epimark® Taq, Hemo Klen Taq, OneTaq), Bst DNA polymerase, Bsu DNA polymerase, phi29 DNA polymerase, and T7 DNA polymerase. Other polymerases, such as Pfu, KOD, or Tth DNA polymerase, can also be used.
[0096] Additional enzymes may be used to incorporate sequencing adapters at each end of the cfDNA duplex. For example, in some embodiments, a ligase may be used to join the sequencing adapters to the cfDNA ends by ligation. In some embodiments, a terminal transferase, e.g., terminal deoxynucleotidyl transferase, may be used to extend the end of the 3'-overhang of the cfDNA molecule to generate a nucleotide tail with a known sequence that pairs with the adapter. In some embodiments, the adapter comprises a sequence complementary to the extended sequence generated by the terminal transferase at the 3'-overhang end of the cfDNA molecule. The terminal transferase may attach one or more (e.g., multiple) nucleotide bases to the 3'-overhang end of the cfDNA duplex molecule. The attached bases may be, for example, standard bases (e.g., A, C, T, or G). In some embodiments, the attached bases are of the same nucleotide base type. The adapter may comprise a sequencing primer binding site and, in some embodiments, may further comprise a sample index sequence and / or a UMI sequence. Once the adapters are attached to the 5' and 3' ends of the cfDNA duplex, the first strand and / or the second strand of the cfDNA duplex can be sequenced to generate first and / or second sequencing reads.
[0097] Once a sequencing construct or sequencing library is generated, one or both strands of the DNA duplex molecule can be sequenced to generate one or more sequencing reads. Sequencing can be performed by any suitable process. In some implementations of the method, sequencing is performed by next-generation sequencing. As used herein, "next-generation sequencing" (or "NGS") can also be referred to as "massively parallel sequencing" (or "MPS") and refers to any sequencing method that determines the nucleotide sequence of individual nucleic acid molecules (e.g., in single-molecule sequencing) or clonally expanded proxies of individual nucleic acid molecules in a high-throughput manner. Next-generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference. Other examples of sequencing methods suitable for use in implementing the methods and systems disclosed herein are described, for example, in International Patent Application Publication No. WO 2012 / 092426. In some embodiments, sequencing may include, for example, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, or direct sequencing. In some embodiments, sequencing may be performed using, for example, Sanger sequencing. In some instances, sequencing may include paired-end sequencing techniques that allow both ends of a fragment to be sequenced, generating high-quality, mappable sequence data, for example, for detecting 5' and / or 3' overhang sequences of cfDNA.
[0098] The disclosed methods and systems may be implemented using a sequencing platform, such as a Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, and / or Polonator platform. In some embodiments, the sequencing may include Illumina MiSeq sequencing. In some embodiments, the sequencing may include Illumina HiSeq sequencing. In some embodiments, the sequencing may include Illumina NovaSeq sequencing. Optimized methods for sequencing nucleic acids extracted from samples are described in more detail, for example, in International Patent Application Publication No. WO 2020 / 236941, the entire contents of which are incorporated herein by reference.
[0099] Incorporating dI to fill the gaps in the overhang provides a clear boundary for each sequence read, where the cfDNA sequence ends and the dI-filled sequence in the gaps begins. During amplification (e.g., PCR amplification), dI is converted to one or more nucleotides (e.g., A, T, G, C). During alignment, the nucleotides representing the ends of the irregularities (filled with dI) are not aligned (mapped) to the reference genome sequence, thus identifying where the original cfDNA sequence (e.g., fragment) ends and where the dI-filled sequence in the gaps begins. The inclusion of dI causes incorrect nucleotide pairings that can be identified by sequence alignment. Computational removal of dI sequences by soft clipping provides information about the length of the 5' and / or 3' overhang sequences of each cfDNA molecule. In combination with sequence alignment, the sequences of the 5' and / or 3' overhangs can be confirmed, as described in more detail below.
[0100] The sequence reads can be analyzed to identify one or more bases in the sequence reads that are soft-clipped. As further described herein, the method can include incorporating an inosine base during the creation of a sequencing construct. The soft clipping can be based on the presence of an inosine stretch that represents an overhang on the complementary strand. Thus, the length of the overhang in one strand can be based on the soft clipping of one or more inosine stretches in the complementary strand. Given the length of the overhang and the sequence near the overhang, the sequence of the overhang itself can be determined.
[0101] Determining one or more bases in a soft-clipped sequence read includes aligning the sequence read with a reference sequence or the sequence of a complementary strand. The complementary strand may be determined using standard molecular barcode (i.e., unique molecular identifier (UMI)) technologies. These UMI technologies can be either exogenous unique identifiers or non-unique identifiers. For example, an exogenous barcode attached to a molecule is not unique, but after alignment, other features of the molecule determined from the read (e.g., coordinates, portions of endogenous sequence, etc.) are combined with the non-unique barcode to create a UMI. Alignment is the process of matching a query sequence read with one or more additional sequence reads or reference sequences. Alignment may further include mapping the sequence read to a location, such as a genomic position or locus, within the reference sequence. In some embodiments, the sequence read may be aligned to a known reference sequence (e.g., a wild-type sequence). In some embodiments, the reference sequence may be obtained from a database of human genomes (e.g., the HG19 human reference genome) or cancer mutations (e.g., COSMIC). Methods for sequence alignment of sequence reads are described, for example, in Trapnell, C. and Salzberg, SL Nature Biotech., 2009, 27:455-457. Optimization of sequence alignment has been described in the art, for example, in International Patent Application Publication No. 2012 / 092426. Additional descriptions of sequence alignment methods are described in more detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire contents of which are incorporated herein by reference.
[0102] In some embodiments, the methods and systems disclosed herein may integrate the use of multiple individually tailored alignment methods or algorithms to optimize base-calling performance in sequencing methods, particularly methods that rely on massively parallel sequencing (MPS). In some embodiments, the disclosed methods and systems may include the use of one or more global alignment algorithms. In some embodiments, the disclosed methods and systems may include the use of one or more local alignment algorithms.Examples of alignment algorithms that may be used include, but are not limited to, the Burrows-Wheeler Alignment (BWA) software bundle (see, e.g., Li, et al. (2009), "Fast and Accurate Short-Read Alignment with Burrows-Wheeler Transform", Bioinformatics 25:1754-60; Li, et al. (2010), "Fast and Accurate Long-Read Alignment with Burrows-Wheeler Transform", Bioinformatics epub. PMID:20080505), the Smith-Waterman algorithm (see, e.g., Smith, et al. (1981), "Identification of Common Molecular Subsequences", J. Molecular Biology 147(1):195-197), the Stripped Smith-Waterman algorithm (see, e.g., Farrar (2007), "Striped Smith-Waterman Speeds Database Searches Six Times Over Other SIMD Implementations”, Bioinformatics 23(2):156-161), the Needleman-Wunsch algorithm (Needleman, et al. (1970) “A General Method Applicable to the Search for Similarities in the Amino Acid Sequence of Two Proteins”, J. Molecular Biology 48(3):443-53), or any combination thereof.
[0103] The disclosed methods and systems can be implemented using soft clipping analysis. A first sequence and / or a second sequence can be aligned to a reference sequence, or a first sequence can be aligned to a second sequence read, or vice versa. The 5' and / or 3' overhang sequences, in which the overhang gaps are filled with inosine, do not align with the reference sequence, and these inosine-filled gaps are therefore computationally trimmed from the first and / or second reads of the first and / or second sequences of the cfDNA molecule. The length of the 5' and / or 3' overhang can be calculated based on the number of misaligned inosine bases that are soft-clipped. This information can then be further processed to identify additional DNA topological characteristics, including the sequence of the 5' and / or 3' overhang, and their association with diseases such as cancer can be analyzed.
[0104] FIG. 1A shows an exemplary method for determining cell-free DNA topology, such as determining the length of a 5' overhang of a cell-free DNA duplex, according to some embodiments. As shown at 102 in FIG. 1A, the method includes extending the 3' end of a first strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of a second strand of the cell-free DNA duplex. In some embodiments, the inosine is deoxyinosine (dI). At 104, the method includes attaching a sequencing adapter to the cell-free DNA duplex. In some embodiments, attaching the sequencing adapter includes ligating the sequencing adapter to the cell-free DNA duplex. The sequencing adapter may be, for example, a Y-shaped sequencing adapter. In some embodiments, the adapter may be a Y-full-length adapter (e.g., a Y-shaped adapter including an index sequence), a stubby adapter, or a hairpin adapter. In some embodiments, generating a sequencing construct may include, for example, amplifying the first and second strands of the cell-free DNA duplex prior to sequencing. The amplification process may allow for binding of the sample index to the cell-free DNA duplex. However, in some embodiments, the sample index is included in the sequence adapter, eliminating the need to incorporate the sample index into the downstream amplification process. In some embodiments, amplifying includes performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0105] At 106, the method includes sequencing the first strand of the cell-free DNA duplex to generate first sequence reads. In some embodiments, the sequencing includes next-generation sequencing ("NGS"). In some embodiments, the sequencing includes paired-end sequencing. In some embodiments, the sequencing adaptor or the second sequencing adaptor includes an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence. In some embodiments, the sequencing includes use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. In some embodiments, the sequencing includes massively parallel sequencing, and the massively parallel sequencing technology includes next-generation sequencing (NGS). In some embodiments, the sequencing is performed using a next-generation sequencer.
[0106] At 108 in FIG. 1A, the method further includes determining, by one or more processors, one or more bases in the first sequence read to be soft-clipped. Determining one or more bases in the first sequence read to be soft-clipped may include, for example, aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read. The unaligned portions of the first sequence read may be associated with inosine bases in the extended first strand. Because these inosine bases are artificially generated, they do not align with the reference sequence or the complementary strand, and thus represent overhangs in the complementary strand. Thus, these unaligned bases in the 3' portion of the sequence read may be computationally identified for soft-clipping.
[0107] In some embodiments, the method includes sequencing a second strand of the cell-free DNA duplex to generate a second sequence read. One or more bases in the first sequence read to be soft-clipped can be determined by aligning the first sequence (i.e., the first strand) to the second sequence read (i.e., the second strand) to identify unaligned portions of the first sequence read. Matching the first sequence read to the second sequence read can include using a UMI common between the first sequence read and the second sequence read. The unaligned 3' portion of the sequence read can be identified for soft-clipping.
[0108] At 110 in Figure 1A, the method includes soft-clipping, by one or more processors, bases corresponding to the 3' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence reads. At 112, the method includes determining, by one or more processors, the length or sequence of a 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine stretch of the first strand of the cell-free DNA duplex.
[0109] FIG. 1B illustrates an exemplary embodiment of the process described in FIG. 1A. While the double-stranded DNA duplex in the example shown in FIG. 1B includes a 5' overhang on each of the top and bottom strands, the method can be performed even when only one strand has a 5' overhang or when both strands have a 5' overhang. In the example shown at 114, the 3' ends of the top and bottom strands are extended using an inosine base (e.g., deoxyinosine) to fill the 5' overhang. Optionally, a polymerase capable of attaching a single nucleotide (e.g., an inosine base) to the 3' end (e.g., a blunt end) of a DNA duplex molecule can be used to generate a single 3' nucleotide overhang (e.g., a single 3' inosine overhang), as shown at 116 in FIG. 1B. The polymerase may be, for example, Taq polymerase. In another embodiment, a terminal transferase can be used to attach one or more (e.g., multiple) nucleotide bases to the 3' overhang end of a cfDNA duplex molecule. The attached base can be, for example, a standard base (e.g., A, C, T, or G). A 3' overhang can increase the efficiency of binding of the sequencing adapter, and the sequencing adapter can include a 3' overhang that complements the 3' overhang of the cfDNA molecule. In some embodiments, the attached base is of the same nucleotide base type. Then, at 118, a sequencing adapter can be attached to the DNA duplex molecule. Optionally, if the 3' end includes a tail, the adapter can include a 3' overhang that complements the tail. For example, if the 3' end includes a single inosine tail, the adapter can include a 3' overhang that includes a cytosine base, as shown in FIG. 1B. This exemplary embodiment further illustrates an optional step involving amplification of the first strand of the cell-free DNA duplex, for example, by polymerase chain reaction (PCR). In this example, the amplification process adds a sample index to the first and second strands of the cell-free DNA duplex. In some embodiments, the sequencing adaptor can include a unique molecular identifier (UMI) to enable pairing of the first strand (top) and the second strand (bottom).
[0110] As shown at the top of FIG. 1B , once the sequencing construct is prepared, the nucleic acid molecule can be sequenced at 122 to provide a sequencing read. Optionally, the sequencing construct can be amplified, e.g., by PCR, prior to sequencing, as shown at 120, which can allow for the incorporation of an index sequence, such as a sample index. In the exemplary embodiment shown, the sequencing read is aligned, e.g., to a reference sequence, at 124. In another example, the top and bottom strands of the same DNA duplex molecule are aligned to each other. The alignment may include the use of unique molecular identifiers (e.g., contained in the sequencing adapter), which allow the top and bottom strands to match each other. Portions of the sequence read that do not align with the reference sequence or complementary strand are identified for soft clipping, as shown at 126. That is, bases that do not align with the reference sequence or complementary strand can be identified as inosine bases used to artificially extend the 3' end. The soft-clipped bases can be used to determine the length of the 5' overhang of the complementary strand if inosine bases are used to fill in the 5' overhang.
[0111] In some embodiments, a method for determining cell-free DNA topology is provided, the method comprising: extending the 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill a 5' overhang of a second strand of the cell-free DNA duplex; attaching a sequencing adaptor to the cell-free DNA duplex; sequencing the first strand of the cell-free DNA duplex to generate first sequence reads; determining, by one or more processors, one or more bases in the first sequence reads to be soft-clipped; soft-clipping, by the one or more processors, bases corresponding to the 3' inosine extension of the first strand of the cell-free DNA duplex from the first sequence reads; and determining, by the one or more processors, the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a single 3' inosine overhang. In some embodiments, the sequencing adaptor comprises a 3' nucleotide overhang (eg, a cytosine overhang) that is complementary to the single 3' inosine overhang.
[0112] In some embodiments, the method further includes detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the disease is cancer.
[0113] The methods described herein may additionally or alternatively include generating a sequencing construct for analyzing the 3' overhang of a DNA molecule, for example, as shown in Figure 2A. As shown at 202, the method includes attaching (e.g., by ligation) a sequencing adapter to the 3' overhang of the first strand of the cell-free DNA duplex. In some implementations, the sequencing adapter is ligated directly to the 3' overhang of the first strand. In other implementations, the 3' overhang of the first strand can be extended by one or more bases (i.e., providing a 3' extension or tail). The sequencing adapter can include a 3' overhang that is complementary to the 3' extension of the first strand of the cell-free DNA duplex. Including an overhang in the adapter that is complementary to the extension or tail attached to the 3' overhang improves the efficiency of the ligation reaction. The 3' extension or tail can be, for example, a single inosine. Inosine is an analog of guanine and therefore pairs preferentially with cytosine bases (although it may also pair with other canonical nucleotide bases, i.e., adenine, thymine, or guanine). If a single inosine tail is included at the 3' end, the adapter may include, for example, a 3' overhang containing a cytosine base complementary to the inosine base.
[0114] In some embodiments, sequencing adaptors are attached to the 3' overhang of the first strand of a cell-free DNA duplex using one or more enzymes with terminal transferase and ligase activity. These enzymes extend the 3' overhang and attach a sequencing adaptor to the 3' overhang extension. This can be done, for example, using the Adaptase™ module from xGen™, which contains enzymes with terminal transferase and ligase activity. Terminal transferase can generate a 3' tail, also referred to as an "adapter tail" or "AdT." In some implementations, the 3' tail provided by terminal transferase activity contains multiple bases of the same base type. The sequencing adaptor includes a 3' overhang that is complementary to the 3' base attached to the 3' overhang of the DNA duplex molecule. This process is further illustrated in Figure 3C.
[0115] At 204 in Figure 2A, the method further includes extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex. In some embodiments, the inosine is deoxyinosine (dI). At 206, the 3' inosine-extended end of the sequencing adaptor is attached (e.g., by ligation) to the 5' end of the second strand of the cell-free DNA duplex. The attached 3' inosine-extended end of the sequencing adaptor thereby provides a 5' inosine extension of the second strand of the cell-free DNA duplex. That is, the inosine extension is between the sequencing adaptor and the original second strand of the DNA duplex molecule.
[0116] 2B shows an exemplary method for determining cell-free DNA topology, such as determining the length of a 3' overhang of a cell-free DNA duplex, based on the method for generating a sequencing construct shown in FIG. 2A. At 208, the second strand of the cell-free DNA duplex is sequenced to generate a first sequence read. In some embodiments, the sequencing comprises next-generation sequencing ("NGS"). In some embodiments, the sequencing comprises paired-end sequencing. In some embodiments, the sequencing adaptor or the second sequencing adaptor comprises an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence. In some embodiments, the sequencing comprises the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. In some embodiments, the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technology comprises next-generation sequencing (NGS). In some embodiments, the sequencing is performed using a next-generation sequencer.
[0117] At 210 in FIG. 2B , the method further includes determining, by one or more processors, one or more bases in the first sequence read to be soft-clipped. Determining one or more bases in the first sequence read to be soft-clipped may include, for example, aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read. The unaligned portions of the first sequence read may be associated with inosine bases in the extended first strand. Because these inosine bases are artificially generated, they do not align with the reference sequence or the complementary strand, and thus represent overhangs in the complementary strand. Thus, these unaligned bases in the 5′ portion of the sequence read may be identified for soft-clipping.
[0118] In some embodiments, the method includes sequencing a second strand of the cell-free DNA duplex to generate a second sequence read. One or more bases in the first sequence read to be soft-clipped can be determined by aligning the first sequence (i.e., the first strand) to the second sequence read (i.e., the second strand) to identify unaligned portions of the first sequence read. Matching the first sequence read to the second sequence read can include using a UMI common between the first sequence read and the second sequence read. The unaligned 5' portion of the sequence read can be identified for soft-clipping.
[0119] At 212 in Figure 2B, the method includes soft-clipping, by one or more processors, bases corresponding to the 5' inosine stretch of the second strand of the cell-free DNA duplex from the first sequence reads. At 214, the method includes determining, by one or more processors, a length of a 3' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the second strand of the cell-free DNA duplex. The sequence of the 3' overhang can be determined based on the length of the 3' overhang of the first strand of the cell-free DNA duplex and the sequence of the first strand of the cell-free DNA.
[0120] In some embodiments, a method of generating a sequencing construct is provided, the method comprising: attaching a sequencing adaptor to a 3' overhang of a first strand of a cell-free DNA duplex; extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the sequencing adaptor to the 5' end of the second strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the second strand of the cell-free DNA duplex. In some embodiments, attaching a sequencing adaptor to the 3' overhang of the first strand of the cell-free DNA duplex comprises extending the 3' overhang of the first strand of the cell-free DNA duplex to provide a 3' extension, wherein the sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the first strand of the cell-free DNA duplex, and attaching the sequencing adaptor to the 3' extension of the first strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the first strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0121] In some embodiments, a method for determining cell-free DNA topology is provided, the method comprising: generating a sequencing construct described herein; sequencing a second strand of a cell-free DNA duplex to generate first sequence reads; determining, by one or more processors, one or more bases in the first sequence reads to be soft-clipped; soft-clipping, by one or more processors, bases corresponding to the 5' inosine stretch of the second strand of the cell-free DNA duplex from the first sequence reads; and determining, by one or more processors, the length or sequence of a 3' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the second strand of the cell-free DNA duplex. In some embodiments, the method comprises sequencing the first strand of the cell-free DNA duplex to generate second sequence reads; and determining, by one or more processors, one or more bases in the second sequence reads to be soft-clipped.
[0122] In some embodiments, the method comprises attaching a second sequencing adaptor to a 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with inosine bases to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to a 5' end of the first strand of the cell-free DNA duplex, wherein the 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. In some embodiments, the method further includes sequencing a first strand of the cell-free DNA duplex to generate second sequence reads, determining, by one or more processors, one or more bases attached to the 5' end of the second sequence reads to be soft-clipped, soft-clipping, by one or more processors, bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the second sequence reads, and determining, by one or more processors, the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. In some embodiments, attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises attaching a single inosine to the 3' overhang of the second strand of the cell-free DNA duplex and attaching a second sequencing adaptor to the inosine. In some embodiments, the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the single inosine attached to the 3' overhang of the second strand of the cell-free DNA duplex.
[0123] In some embodiments, the method comprises amplifying the first strand and the second strand of the cell-free DNA duplex, as described above. In some embodiments, the method comprises sequencing, as described above. In some embodiments, determining one or more bases in the first sequence read to be soft-clipped comprises aligning the first sequence read and / or sequence reads, as described above.
[0124] In some embodiments, the method further comprises detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of a 3' overhang of the first strand of the cell-free DNA duplex. In some embodiments, the method further comprises detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the disease is cancer.
[0125] In some instances, a DNA duplex molecule includes both a 5' overhang and a 3' overhang. The methods provided herein can be used to generate sequencing constructs and analyze such constructs to determine the topology of both the 5' overhang and the 3' overhang of a DNA duplex. Figure 3A illustrates a method for determining the topology of a DNA duplex (e.g., cfDNA), such as determining the length of the 3' overhang and the length of the 5' overhang of a DNA duplex molecule (e.g., a cell-free DNA duplex molecule). As shown at 302, the method includes extending the 3' end of a first strand of the DNA duplex with an inosine base to fill in the 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the inosine is deoxyinosine (dI). The 3' end of the first strand of the DNA duplex molecule can be extended using a polymerase. Optionally, a single 3' nucleotide overhang (e.g., a single 3' inosine overhang) may be generated using a polymerase capable of attaching a single nucleotide (e.g., an inosine base) to the 3' end (e.g., a blunt end) of a DNA duplex molecule. The polymerase may be, for example, Taq polymerase. In another embodiment, a terminal transferase may be used to attach one or more (e.g., multiple) nucleotide bases to the 3' overhang end of a cfDNA duplex molecule. The attached bases may be, for example, standard bases (e.g., A, C, T, or G). The 3' overhang may increase the efficiency of binding of a sequencing adaptor, and the sequencing adaptor may include a 3' overhang that is complementary to the 3' overhang of the cfDNA molecule. In some embodiments, the attached bases are of the same nucleotide base type.
[0126] At 304, the method includes attaching a first sequencing adaptor and a second sequencing adaptor to the DNA duplex molecule. In some embodiments, attaching the sequencing adaptor includes ligating the sequencing adaptor to the cell-free DNA duplex. The sequencing adaptor may be, for example, a Y-shaped sequencing adaptor. In some embodiments, the adaptor may be a Y-full-length adaptor (e.g., a Y-shaped adaptor including an index sequence), a stubby adaptor, or a hairpin adaptor. As shown in Figures 3B and 3C, the first sequencing adaptor may be attached to the end of a DNA duplex molecule having a 5' overhang. The first and second sequencing adaptors may be attached to the DNA duplex molecule in the same reaction (e.g., as shown in Figure 3B) or may be attached to the DNA duplex molecule in sequential reactions (e.g., as shown in Figure 3B). In the exemplary process shown in Figure 3B, the first sequencing adaptor includes a 3' overhang that is complementary to the 3' tail of the first strand of the DNA duplex molecule. Similarly, the second sequencing adaptor includes a 3' overhang that is complementary to the 3' tail of the second strand of the DNA duplex molecule. As described above, a polymerase capable of attaching a single nucleotide (e.g., an inosine base) to the 3' end (e.g., a blunt end) of a DNA duplex molecule can be used to generate a single 3' nucleotide overhang (e.g., a single 3' inosine overhang). Inosine is an analog of guanine and therefore preferentially pairs with cytosine bases (although it can also pair with other canonical nucleotide bases (i.e., adenine, thymine, or guanine)). For example, if the 3' end includes an inosine tail, the adaptor can include a 3' overhang that includes a single nucleotide (e.g., a cytosine base) that preferentially complements the inosine base. The polymerase can be, for example, Taq polymerase. In another embodiment, a terminal transferase can be used to attach one or more (e.g., multiple) nucleotide bases to the 3' overhang end of a cfDNA duplex molecule. The bases to be attached can be, for example, standard bases (eg, A, C, T, or G).The 3' overhang may increase the binding efficiency of the sequencing adaptor, and the sequencing adaptor may include a 3' overhang that is complementary to the 3' overhang of the cfDNA molecule. In some embodiments, the bound bases are of the same nucleotide base type.
[0127] In some implementations, the first and second sequencing adaptors are sequentially attached to the DNA duplex molecule, e.g., attaching the first sequencing adaptor to the DNA duplex molecule followed by attaching the second sequencing adaptor to the DNA duplex molecule. The first sequencing adaptor can be attached as described above. To attach the second sequencing adaptor, an enzyme with terminal transferase can be used to extend the 3' overhang of the second strand of the DNA molecule, as shown in FIG. 3C. This can be done, for example, using the Adaptase™ module from xGen™, which includes an enzyme with terminal transferase and ligase activity. The terminal transferase can generate a 3' tail, also referred to as an "adapter tail" or "AdT." In some implementations, the 3' tail provided by the terminal transferase activity includes multiple bases of the same base type. The sequencing adaptor includes a 3' overhang that is complementary to the 3' base attached to the 3' overhang of the DNA duplex molecule.
[0128] At 306 in Figure 3A, the method further includes extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex. In some embodiments, the inosine is deoxyinosine (dI). At 308, the 3' inosine-extended end of the sequencing adaptor is attached (e.g., by ligation) to the 5' end of the second strand of the cell-free DNA duplex. The attached 3' inosine-extended end of the sequencing adaptor thereby provides a 5' inosine extension of the second strand of the cell-free DNA duplex. That is, the inosine extension is between the sequencing adaptor and the original second strand of the DNA duplex molecule.
[0129] In some embodiments, the sequencing construct is amplified, for example, prior to sequencing. The amplification process may allow for binding of the sample index to the cell-free DNA duplex. However, in some embodiments, the sample index is included in the sequence adapter, eliminating the need to incorporate the sample index into the downstream amplification process. In some embodiments, amplifying includes performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0130] The first strand of the DNA duplex molecule is sequenced at 310 to generate a first sequence read. In some implementations, both the first strand and the second strand of the DNA duplex molecule are sequenced to generate a first sequence read and a second sequence read, respectively. In some embodiments, the sequencing comprises next-generation sequencing ("NGS"). In some embodiments, the sequencing comprises paired-end sequencing. In some embodiments, the sequencing adaptor or the second sequencing adaptor comprises an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence. In some embodiments, the sequencing comprises use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. In some embodiments, the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technology comprises next-generation sequencing (NGS). In some embodiments, the sequencing is performed using a next-generation sequencer.
[0131] At 312, the method includes determining, by one or more processors, one or more bases attached to the 5' end and / or 3' end of the first sequence read to be soft-clipped. As described above, during amplification (e.g., PCR amplification), inosine bases are converted to one or more nucleotides (e.g., A, T, G, C). During alignment, nucleotides representing inosine bases are not aligned (mapped) to the reference genome sequence, thus identifying where the original cfDNA sequence (e.g., fragment) ends and where the inosine base sequence of the gap in the overhang begins. Therefore, determining one or more bases in the first sequence read to be soft-clipped may include, for example, aligning the first sequence read to the reference sequence to identify unaligned portions of the first sequence read. The unaligned portions of the first sequence read (i.e., unaligned nucleotides) may be associated with inosine bases in the extended first strand. These nucleotides representing inosine bases are artificially generated and therefore do not align with the reference sequence or the complementary strand, thus representing overhangs in the complementary strand. Therefore, these non-aligned nucleotide bases in the 5' and / or 3' portions of the sequence read can be identified for soft clipping.
[0132] In some embodiments, the method includes sequencing a second strand of the cell-free DNA duplex to generate a second sequence read. One or more bases in the first sequence read to be soft-clipped can be determined by aligning the first sequence (i.e., the first strand) to the second sequence read (i.e., the second strand) to identify unaligned portions of the first sequence read. Matching the first sequence read to the second sequence read can include using a UMI common between the first sequence read and the second sequence read. Unaligned 5' and / or 3' portions of the sequence read can be identified for soft-clipping.
[0133] At 314, the method includes soft-clipping, by the one or more processors, bases corresponding to the 3' inosine stretch and / or the 5' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence reads. At 316, the method includes determining, by the one or more processors, the length or sequence of a 5' overhang of a second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine stretch of the first strand of the cell-free DNA duplex and / or determining the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex.
[0134] In some embodiments, methods for determining cell-free DNA topology are provided, the method comprising: extending a 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill in a 5' overhang of a second strand of the cell-free DNA duplex; attaching a sequencing adaptor to the cell-free DNA duplex; sequencing the first strand of the cell-free DNA duplex to generate first sequence reads; determining, by one or more processors, one or more bases in the first sequence reads to be soft-clipped; soft-clipping, by the one or more processors, bases corresponding to the 3' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence reads; and determining, by the one or more processors, the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine stretch of the first strand of the cell-free DNA duplex. In some embodiments, the method includes attaching a second sequencing adaptor to a 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to a 5' end of the first strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the first strand cell-free DNA duplex. determining, by the one or more processors, one or more bases attached to the 5' end of the first sequence read to be soft-clipped; soft-clipping, by the one or more processors, bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the first sequence read; and determining, by the one or more processors, the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex.In some embodiments, the method comprises extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex; attaching a second sequencing adaptor to the cell-free DNA duplex; soft-clipping, by one or more processors, bases corresponding to the 3' inosine extension of the second strand of the cell-free DNA duplex from the second sequence reads; and determining, by the one or more processors, the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the second strand of the cell-free DNA duplex.
[0135] In some embodiments, attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises extending the 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex, and attaching the second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. In some embodiments, the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type.
[0136] In some embodiments, the method further comprises detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of a 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the method further comprises detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the method further comprises detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of a 5' overhang of the first strand of the cell-free DNA duplex. In some embodiments, the method further comprises detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of a 3' overhang of the first strand of the cell-free DNA duplex. In some embodiments, the method further comprises detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the disease is cancer.
[0137] In some embodiments, the method comprises amplifying the first strand and the second strand of the cell-free DNA duplex, as described above. In some embodiments, the method comprises sequencing, as described above. In some embodiments, determining one or more bases in the first sequence read that are soft-clipped comprises aligning the first sequence read and / or the second sequence read, as described above.
[0138] Figure 4 illustrates a process used to determine cfDNA topology, according to some embodiments. Sequence reads obtained by the methods described herein are received by one or more processors. In the process illustrated in Figure 4, Read 1 and Read 2 represent reads sequenced from both ends of a single strand of a duplex. The sequence reads are subjected to preprocessing before further analysis. This may include, for example, trimming poly-G regions (e.g., dark cycles), adapter sequences, and / or portions of the sequence read that are deemed low quality (e.g., based on a sequencing quality score below a predetermined threshold, such as a Q score of Q10 or less, Q20 or less, or Q30 or less). If the sequencing read includes an adapter tail (a 3'-end tail added by terminal transferase), this adapter tail may be trimmed. The resulting sequencing reads (i.e., after preprocessing) may be further analyzed to determine one or more bases in the sequencing read that are soft-clipped, for example, by aligning the sequencing read to a reference sequence. Based on the alignment, one or more bases for soft clipping are identified, and the length to be soft clipped (i.e., the "uneven" length) is measured. This length indicates the length of the overhang on the complementary strand. The genomic content of the complementary strand can be extracted to provide the sequence of the overhang.
[0139] Figure 5 shows an exemplary read processing matrix for extracting data from paired-end sequencing, according to some embodiments, including the sequence construction and analysis methods shown in Figures 3A and 3B.
[0140] In some implementations of any of the methods described herein, DNA duplex molecules are obtained from an individual. The method may further include generating a DNA duplex molecule overhang profile for the individual, the profile including information about a plurality of DNA duplex molecule overhangs. For example, the information may include, for the plurality of DNA duplex molecules, the length of a 3' overhang of a first strand of the DNA duplex molecule, the length of a 5' overhang of a first strand of the DNA duplex molecule, the length of a 3' overhang of a second strand of the DNA duplex molecule, and / or the length of a 5' overhang of a second strand of the DNA duplex molecule. In some implementations, the information includes the sequence of the 5' overhang or the 3' overhang. In some implementations, the information includes: (1) the ratio of the length of a 5' overhang of a first strand of a DNA duplex molecule to the length of a 3' overhang of a first strand of a DNA duplex molecule; (2) the ratio of the length of a 5' overhang of a first strand of a DNA duplex molecule to the length of a 3' overhang of a second strand of a DNA duplex molecule; (3) the ratio of the length of a 5' overhang of a first strand of a DNA duplex molecule to the length of a 5' overhang of a second strand of a DNA duplex molecule; (4) the ratio of the length of a 3' overhang of a first strand of a DNA duplex molecule to the length of a 5' overhang of a second strand of a DNA duplex molecule; or (5) the ratio of the length of a 3' overhang of a first strand of a DNA duplex molecule to the length of a 3' overhang of a second strand of a DNA duplex molecule. In some implementations, the information includes (1) the length of the 5' overhang of the first strand of the DNA duplex molecule, the length of the 3' overhang of the first strand of the DNA duplex molecule, the length of the 5' overhang of the second strand of the DNA duplex molecule, the length of the 3' overhang of the second strand of the DNA duplex molecule, and (2) the ratio of the duplex lengths.
[0141] The method may further include comparing the DNA duplex molecule overhang profile to a reference DNA duplex molecule overhang profile, for example, using one or more processors. In some implementations, the reference DNA duplex molecule overhang profile is based on DNA duplex molecules from a normal sample, multiple normal samples, or a synthetically generated normal sample. In some implementations, the reference DNA duplex molecule overhang profile is based on DNA duplex molecules from a sample obtained from an individual with cancer or multiple individuals with cancer. In some implementations, the reference DNA duplex molecule overhang profile is based on DNA duplex molecules from a sample obtained from an individual with an abnormal fetus or multiple individuals with abnormal fetuses. In some implementations, the reference DNA duplex molecule overhang profile is based on DNA duplex molecules from a sample obtained from an individual who has undergone stable transplantation or multiple individuals who have undergone stable transplantation. In some implementations, the reference DNA duplex molecule overhang profile is based on DNA duplex molecules from a matched normal sample obtained from the individual. In some implementations, the reference DNA duplex molecule overhang profile is based on DNA duplex molecules from a previous sample obtained from the individual. This prior sample can be a normal baseline or a disease state baseline (ie, a profile of DNA duplex molecular overhangs representing a pre-disease state).
[0142] In some embodiments, the methods described herein further include generating, by the one or more processors, a report indicating the length of the 3' overhang of the first strand of the cell-free DNA duplex, the length of the 3' overhang of the second strand of the cell-free DNA duplex, the length of the 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of the 5' overhang of the second strand of the cell-free DNA duplex, or other DNA duplex molecule overhang profile information. In some embodiments, the method further includes transmitting the report to a healthcare provider. In some embodiments, the report is transmitted via a computer network, a peer-to-peer connection, or an application programming interface (API).
[0143] In some embodiments, the method further comprises generating a genomic profile of the subject, the genomic profile comprising the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of a 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of a 5' overhang of the second strand of the cell-free DNA duplex. In some embodiments, the subject's genomic profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the subject's genomic profile further comprises results from a nucleic acid sequencing-based test.
[0144] DNA topological information, such as the length of the 3'-overhang and / or 5'-overhang and / or sequence information determined according to the methods described herein, can be used to detect the presence, possible presence, or absence of disease (e.g., cancer). This is particularly useful, for example, in determining the presence or absence of disease in early-stage cancer patients or cancer patients with low cancer levels, or in detecting disease recurrence. For example, in some implementations, the circulating tumor DNA (ctDNA) content of an individual is less than 1%, less than 0.8%, less than 0.5%, less than 0.3%, less than 0.1%, or less than 0.05%.
[0145] sample The disclosed methods and systems can be used with any of a variety of samples (also referred to herein as specimens) containing nucleic acid (e.g., DNA) collected from a subject (e.g., a patient). Examples of samples include, but are not limited to, a liquid biopsy sample, a blood sample (e.g., a peripheral whole blood sample), a plasma sample, a serum sample, a lymph sample, a saliva sample, a sputum sample, a urine sample, a gynecological fluid sample, a circulating tumor cell (CTC) sample, a cerebrospinal fluid (CSF) sample, a pericardial fluid sample, a pleural effusion sample, an ascites (peritoneal fluid) sample, a fecal (or stool) sample, or other bodily fluid, secretion, and / or excretory sample (or a sample derived therefrom).
[0146] In some embodiments, the sample may be collected by needle biopsy, fine needle aspiration, collection cup or tube, oral swab, nasal swab, vaginal swab, or cytology smear, or the like.
[0147] In some embodiments, the sample is a liquid biopsy sample and may include, for example, whole blood, plasma, serum, urine, stool, sputum, saliva, or cerebrospinal fluid. In some embodiments, the sample may be a liquid biopsy sample and may include circulating tumor cells (CTCs). In some embodiments, the sample may be a liquid biopsy sample and may include cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.
[0148] In some embodiments, the disclosed methods may further include analyzing a primary control (e.g., a normal tissue sample). In some embodiments, the disclosed methods may further include determining whether a primary control is available and, if so, isolating control nucleic acid (e.g., DNA) from the primary control. In some embodiments, if a primary control is not available, the sample may include an optional normal control. In some embodiments, the method includes evaluating the sample (e.g., a normal sample) using a method described herein. In some embodiments, the disclosed methods may further include determining that a primary control is not available and marking the sample for analysis without a matching control.
[0149] In some embodiments, nucleic acids extracted from a sample may include deoxyribonucleic acid (DNA) molecules. Examples of DNA that may be suitable for analysis by the disclosed methods include, but are not limited to, genomic DNA or fragments thereof, mitochondrial DNA or fragments thereof, cell-free DNA (cfDNA), and circulating tumor DNA (ctDNA). Cell-free DNA (cfDNA) is composed of fragments of DNA that are released from normal and / or cancer cells during apoptosis and necrosis and circulate in the bloodstream and / or accumulate in other body fluids. Circulating tumor DNA (ctDNA) is composed of fragments of DNA released from cancer cells and tumors that circulate in the bloodstream and / or accumulate in other body fluids.
[0150] In some embodiments, cell-free DNA or circulating tumor DNA is extracted from a liquid sample. In some embodiments, samples with low nucleated cellularity may require more, e.g., larger volumes, for DNA extraction.
[0151] In some embodiments, the method for determining cell-free DNA topology further comprises obtaining cell-free DNA duplexes from a subject. In some embodiments, the cell-free DNA duplexes are obtained from a liquid biopsy sample. In some embodiments, the sample is a liquid biopsy sample, including blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the cell-free DNA duplexes are circulating tumor DNA (ctDNA) duplexes.
[0152] subject In some embodiments, the sample is obtained (e.g., taken) from a subject (e.g., a human patient) having or suspected of having a condition or disease (e.g., a hyperproliferative disease or a non-cancer indication). In some embodiments, the hyperproliferative disease is cancer. In some embodiments, the cancer is a solid tumor or a metastatic form thereof. In some embodiments, the cancer is a blood cancer, such as leukemia or lymphoma.
[0153] In some embodiments, the subject has cancer or is at risk of having cancer. For example, in some embodiments, the subject has a genetic predisposition to cancer (e.g., having a genetic mutation that increases their baseline risk of developing cancer). In some embodiments, the subject has been exposed to an environmental perturbation (e.g., radiation or a chemical) that increases their risk of developing cancer. In some embodiments, the subject needs to be monitored for the development of cancer. In some embodiments, the subject needs to be monitored for cancer progression or regression, for example, after being treated with an anti-cancer therapy (or anti-cancer treatment). In some embodiments, the subject needs to be monitored for cancer recurrence. In some embodiments, the subject needs to be monitored for minimal residual disease (MRD). In some embodiments, the subject has been or is being treated for cancer. In some embodiments, the subject has not been treated with an anti-cancer therapy (or anti-cancer treatment).
[0154] In some embodiments, a subject (e.g., a patient) is being treated or has previously been treated with one or more anti-cancer therapies. In some embodiments, for example, for a patient who has previously been treated with a targeted anti-cancer therapy, a post-targeted therapy sample (e.g., specimen) is obtained (e.g., collected). In some embodiments, the post-targeted therapy sample is a sample obtained after completion of the targeted therapy. In some embodiments, the one or more anti-cancer therapies (or anti-cancer treatments) include, but are not limited to, surgery (e.g., surgical resection), radiation therapy, or chemotherapy, and combinations thereof.
[0155] In some embodiments, the patient has not been previously treated with an anti-cancer therapy. In some embodiments, for example, in the case of a patient who has not been previously treated with a targeted anti-cancer therapy, the sample comprises a liquid biopsy (e.g., an initial liquid biopsy or a liquid biopsy after recurrence).
[0156] In some embodiments, the method for determining cell-free DNA topology further comprises obtaining cell-free DNA duplexes from a subject. In some embodiments, the cell-free DNA duplexes are obtained from a patient suspected of having cancer or a subject determined to have cancer. In some embodiments, the cell-free DNA duplexes are obtained from a liquid biopsy sample. In some embodiments, the sample is a liquid biopsy sample, including blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the cell-free DNA duplexes are circulating tumor DNA (ctDNA) duplexes. In some embodiments, the method further comprises treating the subject with an anti-cancer therapy.
[0157] Nucleic Acid Extraction and Processing DNA can be extracted from liquid biopsy samples, including, but not limited to, blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva, or other bodily fluid samples, using any of a variety of techniques known to those of skill in the art (see, e.g., Example 1 of International Patent Application Publication No. WO 2012 / 092426; Tan, et al. (2009), "DNA, RNA, and Protein Extraction: The Past and The Present," J. Biomed. Biotech. 2009:574398; the technical literature for the Maxwell® 16 LEV Blood DNA Kit (Promega Corporation, Madison, WI); and the Maxwell 16 Buccal Swab LEV DNA Purification Kit Technical Manual (Promega Literature #TM333, January 1, 2011, Promega Corporation, Madison, WI)).
[0158] A typical DNA extraction procedure includes, for example, (i) collecting a liquid sample from which DNA is to be extracted, (ii) treating the liquid sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA, and (iii) purifying the DNA from the supernatant to remove detergents, proteins, salts, or other reagents used in previous steps. The DNA sample can optionally be further treated with RNase to digest RNA in the sample.
[0159] Examples of suitable techniques for DNA purification include, but are not limited to, (i) precipitation in ice-cold ethanol or isopropanol followed by centrifugation (precipitating DNA can be enhanced by increasing the ionic strength, e.g., by adding sodium acetate); (ii) phenol-chloroform extraction followed by centrifugation to separate the aqueous phase containing the nucleic acids from the organic phase containing the denatured proteins; and (iii) solid-phase chromatography, in which nucleic acids adsorb to a solid phase (e.g., silica or other) depending on the pH and salt concentration of the buffer.
[0160] In some instances, cellular and histone proteins bound to DNA can be removed by adding proteases, or by precipitating the proteins with sodium acetate or ammonium acetate, or through extraction with a phenol-chloroform mixture prior to the DNA precipitation step.
[0161] In some instances, DNA may be extracted using any of a variety of suitable commercially available DNA extraction and purification kits, including, but not limited to, the QIAamp (for isolating genomic DNA from human samples) and DNAeasy (for isolating genomic DNA from animal or plant samples) kits from Qiagen (Germantown, MD), or the Maxwell® and ReliaPrep™ series from Promega (Madison, WI).
[0162] In some examples, the disclosed methods may further include determining or obtaining a yield value for nucleic acid extracted from the sample and comparing the determined value to a reference value. For example, if the determined or obtained value is less than the reference value, the nucleic acid may be amplified before proceeding with library construction. In some examples, the disclosed methods may further include determining or obtaining a value for the size (or average size) of nucleic acid fragments in the sample and comparing the determined or obtained value to a reference value, e.g., a size (or average size) of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs (bps). In some examples, one or more parameters described herein may be adjusted or selected in response to this determination.
[0163] After isolation, nucleic acids are typically dissolved in a slightly alkaline buffer, such as Tris-EDTA (TE) buffer, or in ultrapure water.
[0164] System and storage medium Further described herein are systems and computer readable storage media containing instructions for causing the systems to perform the methods described herein. An exemplary system includes, for example, one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, direct the system to: receive a first sequence read obtained by extending the 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill in a 5' overhang of a second strand of the cell-free DNA duplex; attach a sequencing adapter to the cell-free DNA duplex; sequence the first strand of the cell-free DNA duplex to generate a first sequence read; determine one or more bases in the first sequence read to be soft-clipped; soft-clip a base corresponding to the 3' inosine extension of the first strand of the cell-free DNA duplex from the first sequence read; and determine the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft-clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex.
[0165] In another example, a system includes one or more processors and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive a first sequence read obtained by ligating a sequencing adaptor to a 3' overhang of a first strand of a cell-free DNA duplex; extend the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and ligate the 3' inosine-extended end of the sequencing adaptor to a second strand of the cell-free DNA duplex. wherein the bound 3' inosine extended end provides a 5' inosine extension of a second strand of the cell-free DNA duplex; determining one or more bases in the first sequence read to be soft-clipped; soft-clipping bases from the first sequence read that correspond to the 5' inosine extension of the second strand of the cell-free DNA duplex; and determining the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex based on the soft-clipping of the 5' inosine extension of the second strand of the cell-free DNA duplex.
[0166] FIG. 6 illustrates an example of a computing device or system according to one embodiment. Device 600 may be a host computer connected to a network. Device 900 may be a client computer or a server. As illustrated in FIG. 6, device 600 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device). The device may include, for example, one or more processors 610, input devices 620, output devices 630, memory or storage devices 640, communication devices 660, and a nucleic acid sequencer 670. Software 650 resident in memory or storage device 640 may include, for example, an operating system and software for implementing methods described herein. Input devices 620 and output devices 630 may generally correspond to those described herein and may be connectable to or integrated with a computer.
[0167] Input device 620 may be any suitable device that provides input, such as a touchscreen, a keyboard or keypad, a mouse, or a voice recognition device. Output device 630 may be any suitable device that provides output, such as a touchscreen, a tactile device, or a speaker.
[0168] Storage 640 may be any suitable device that provides storage (e.g., electrical, magnetic, or optical memory, including RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communications device 660 may include any suitable device that can send and receive signals over a network, such as a network interface chip or device. The components of the computer may be connected in any suitable manner, for example, via a wired medium (e.g., a physical system bus 680, an Ethernet connection, or any other wired transmission technology) or wirelessly (e.g., Bluetooth, Wi-Fi, or any other wireless technology).
[0169] The software modules 650 can be stored as executable instructions in storage 640 and executed by processor 610 and can include, for example, an operating system and / or processes that embody the functionality of the methods of the present disclosure (e.g., embodied in the devices described above).
[0170] The detection module 650 may also be stored in and / or transferred to any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (e.g., those described herein) and may fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of the present disclosure, a computer-readable storage medium may be any medium, such as storage 640, that may contain or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer-readable storage media may include memory units, such as hard drives, flash drives, and distribution modules, that operate as a single functional unit. Additionally, the various processes described herein may be embodied as modules configured to operate in accordance with the above embodiments and techniques. Furthermore, while processes may be shown and / or described separately, those skilled in the art will understand that the above processes may be routines or modules within other processes.
[0171] The software 650 may also be propagated within any transmission medium for use by or in connection with an instruction execution system, apparatus, or device such as those described above, which may fetch instructions associated with the software from and execute the instructions. In the context of this disclosure, a transmission medium may be any medium that may communicate, propagate, or transmit programming for use by or in connection with an instruction execution system, apparatus, or device. Transmission-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.
[0172] Device 600 can be connected to a network (e.g., network 704 shown in FIG. 7 and / or described below), which can be any suitable type of interconnected communications system. The network can implement any suitable communications protocol and can be protected by any suitable security protocol. The network can include any suitable arrangement of network links capable of implementing the transmission and reception of network signals, such as a wireless network connection (T1 or T3 line), a cable network, DSL, or telephone lines.
[0173] Device 600 may be implemented using any operating system, e.g., an operating system suitable for operating on a network. Software module 650 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure may be deployed in different configurations (e.g., in a client / server arrangement, or via a web browser as a web-based application or web service). In some embodiments, the operating system is executed by one or more processors, e.g., processor 610.
[0174] The device 600 can further include a sequencer 670, which can be any suitable nucleic acid sequencing instrument.
[0175] 7 illustrates an example of a computing system, according to one embodiment. In system 700, device 600 (e.g., as described above and illustrated in FIG. 6) is connected to network 704, which is also connected to device 706. In some embodiments, device 706 is a sequencer. Exemplary sequencing instruments may include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX System, Illumina / Solexa's Genome Analyzer (GA), Illumina's HiSeq 2500, HiSeq 3000, HiSeq 4000, and NovaSeq 6000 sequencing systems, Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope Gene sequencing system, or Pacific Biosciences' PacBio RS system.
[0176] Devices 600 and 706 can communicate using a suitable communication interface over a network 704, such as, for example, a local area network (LAN), a virtual private network (VPN), or the Internet. In some embodiments, network 704 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 600 and 706 can communicate partially or entirely via wireless or wired communications, such as Ethernet, IEEE 802.11b wireless, etc. Additionally, devices 600 and 706 can communicate over a second network, such as, for example, a mobile / cellular network, using a suitable communication interface. Communications between devices 600 and 706 can further include or communicate with various servers, such as a mail server, a mobile server, a media server, a telephone server, etc. In some embodiments, devices 600 and 706 can communicate directly (instead of or in addition to communicating over network 704) via wireless or wired communications, such as, for example, Ethernet, IEEE 802.11b wireless, etc. In some embodiments, devices 900 and 1006 communicate via communication 708, which may be a direct connection or may occur over a network (eg, network 704).
[0177] One or all of devices 600 and 706 generally include logic (e.g., http web server logic) or are programmed to format data accessed from local or remote databases or other sources of data and content to provide and / or receive information over network 704 in accordance with various examples described herein.
[0178] Illustrative Embodiments The following embodiments are illustrative and are not intended to limit the scope of the inventions described herein. Exemplary embodiments include, but are not limited to: Embodiment 1. A method for determining cell-free DNA topology, comprising: extending the 3' end of the first strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the second strand of the cell-free DNA duplex; ligating a sequencing adaptor to the cell-free DNA duplex; sequencing a first strand of the cell-free DNA duplex to generate a first sequence read; determining, by the one or more processors, one or more bases in the first sequence read to be soft-clipped; soft clipping, by the one or more processors, bases corresponding to 3' inosine stretches of the first strand of the cell-free DNA duplex from the first sequence reads; and determining, by the one or more processors, the length or sequence of a 5' overhang of a second strand of the cell-free DNA duplex based on soft clipping of the 3' inosine stretch of the first strand of the cell-free DNA duplex. Embodiment 2. The method of embodiment 1, further comprising detecting, by the one or more processors, the presence or absence of the disease based on the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 3. The method of embodiment 2, wherein the disease is cancer. Embodiment 4. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with inosine bases to fill in the 3' overhang of the second strand of the cell-free DNA duplex; ligating a 3' inosine-extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the attached 3' inosine-extended end provides a 5' inosine-extended portion of the first strand cell-free DNA duplex; determining, by the one or more processors, one or more bases attached to the 5′ end of the first sequence read to be soft-clipped; soft clipping, by the one or more processors, bases from the first sequence read that correspond to the 5' inosine stretch of the first strand of the cell-free DNA duplex; 4. The method of any one of embodiments 1-3, further comprising determining, by the one or more processors, the length or sequence of a 3' overhang of a second strand of the cell-free DNA duplex based on soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. Embodiment 5. The method of embodiment 4, comprising detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex. Embodiment 6. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises: extending a 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; 6. The method of embodiment 4 or 5, comprising attaching a second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. Embodiment 7. The method of embodiment 6, wherein the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type. Embodiment 8. The method of any one of embodiments 1 to 7, wherein determining one or more bases in the first sequence read to be soft-clipped comprises aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read. Embodiment 9. Sequencing a second strand of the cell-free DNA duplex to generate second sequence reads; and determining, by one or more processors, one or more bases in the second sequence reads to be soft-clipped. Embodiment 10. The method of embodiment 9, wherein determining one or more bases in the second sequence read to be soft-clipped comprises aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. Embodiment 11. The method of embodiment 9 or 10, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI). Embodiment 12. The method of embodiment 11, wherein determining one or more bases in the first sequence read to be soft-clipped comprises aligning the first sequence to a second sequence read to identify unaligned portions of the first sequence read. Embodiment 13. The method of embodiment 11 or 12, wherein determining one or more bases in the second sequence read to be soft-clipped comprises aligning the second sequence to the first sequence read to identify unaligned portions of the second sequence read. Embodiment 14. Extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex; ligating a second sequencing adaptor to the cell-free DNA duplex; soft-clipping, by the one or more processors, bases corresponding to the 3' inosine stretch of the second strand of the cell-free DNA duplex from the second sequence reads; and determining, by the one or more processors, the length or sequence of a 5' overhang of a first strand of the cell-free DNA duplex based on soft-clipping of a 3' inosine stretch of a second strand of the cell-free DNA duplex. Embodiment 15. The method of embodiment 14, comprising detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex. Embodiment 16 The method of any one of embodiments 1-15, wherein extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a single 3' inosine overhang. Embodiment 17. The method of embodiment 16, wherein the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the 3' inosine overhang. Embodiment 18 The method of any one of embodiments 1 to 17, comprising amplifying the cell-free DNA duplex and binding the sample index to the cell-free DNA duplex. Embodiment 19. The method of any one of embodiments 1 to 18, wherein the sequencing adaptor or the second sequencing adaptor comprises a sample index. Embodiment 20. The method of any one of embodiments 11 to 19, wherein the sequencing adaptor or the second sequencing adaptor comprises a UMI. Embodiment 21 The method of any one of embodiments 1 to 20, wherein the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. Embodiment 22. A method for making a sequencing construct, comprising: Attaching a sequencing adaptor to the 3' overhang of the first strand of the cell-free DNA duplex; extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and attaching a 3' inosine extended end of the sequencing adaptor to the 5' end of the second strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the second strand of the cell-free DNA duplex. Embodiment 23. Attaching a sequencing adaptor to the 3' overhang of the first strand of the cell-free DNA duplex comprises: extending a 3' overhang of a first strand of the cell-free DNA duplex to provide a 3' extension, wherein the sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the first strand of the cell-free DNA duplex; Attaching a sequencing adaptor to the 3' extension of the first strand of the cell-free DNA duplex. Embodiment 24 The method of embodiment 23, wherein the 3' overhang of the first strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type. Embodiment 25. A method for determining cell-free DNA topology, comprising: Creating a sequencing construct according to the method of any one of embodiments 22 to 24; sequencing a second strand of the cell-free DNA duplex to generate a first sequence read; determining, by the one or more processors, one or more bases in the first sequence read to be soft-clipped; soft clipping, by one or more processors, bases corresponding to 5' inosine stretches of the second strand of the cell-free DNA duplex from the first sequence reads; determining, by the one or more processors, the length or sequence of a 3' overhang of a first strand of the cell-free DNA duplex based on soft clipping of a 5' inosine stretch of a second strand of the cell-free DNA duplex. Embodiment 26. The method of embodiment 25, further comprising detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex. Embodiment 27. The method of embodiment 26, wherein the disease is cancer. Embodiment 28. The method of any one of embodiments 1 to 27, wherein determining one or more bases in the 25th sequence read to be soft-clipped comprises aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read. Embodiment 29. Sequencing a first strand of a cell-free DNA duplex to generate second sequence reads; 29. The method of any one of embodiments 25 to 28, further comprising determining, by one or more processors, one or more bases in the second sequence reads to be soft-clipped. Embodiment 30. The method of embodiment 29, wherein determining one or more bases in the second sequence read to be soft-clipped comprises aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. Embodiment 31. The method of embodiment 29 or 30, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI). Embodiment 32. The method of embodiment 21, wherein determining one or more bases in the first sequence read to be soft-clipped comprises aligning the first sequence to a second sequence read to identify unaligned portions of the first sequence read. Embodiment 33. The method of embodiment 11 or 12, wherein determining one or more bases in the second sequence read to be soft-clipped comprises aligning the second sequence to the first sequence read to identify unaligned portions of the second sequence read. Embodiment 34. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with inosine bases to fill in the 3' overhang of the second strand of the cell-free DNA duplex; 34. The method of any one of embodiments 22-33, comprising attaching a 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. Embodiment 35. Sequencing a first strand of a cell-free DNA duplex to generate second sequence reads; determining, by the one or more processors, one or more bases attached to the 5' end of the second sequence read to be soft-clipped; soft-clipping, by the one or more processors, bases corresponding to the 5' inosine stretch of the first strand of the cell-free DNA duplex from the second sequence reads; 35. The method of embodiment 34, further comprising determining, by the one or more processors, the length or sequence of a 3' overhang of the second strand of the cell-free DNA duplex based on soft-clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex. Embodiment 36. The method of embodiment 35, further comprising detecting, by one or more processors, the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex. Embodiment 37. The method of any one of embodiments 34-36, wherein attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises attaching a single inosine to the 3' overhang of the second strand of the cell-free DNA duplex, and attaching the second sequencing adaptor to the single inosine. Embodiment 38. The method of embodiment 37, wherein the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the single inosine attached to the 3' overhang of the second strand of the cell-free DNA duplex. Embodiment 39. The method of any one of embodiments 35 to 38, wherein determining one or more bases attached to the 5' end of the second sequence read to be soft-clipped comprises aligning the second sequence read to a reference sequence to identify unaligned portions of the second sequence read. Embodiment 40. The method of any one of embodiments 35 to 39, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI). Embodiment 41. The method of embodiment 40, wherein determining one or more bases attached to the 5' end of the second sequence read to be soft-clipped comprises aligning the second sequence read to the first sequence read to identify an unaligned portion of the second sequence read. Embodiment 42. The method of embodiment 40 or 41, wherein determining one or more bases attached to the 5' end of the first sequence read to be soft-clipped comprises aligning the first sequence read to a second sequence read to identify an unaligned portion of the first sequence read. Embodiment 43 The method of any one of embodiments 22 to 42, comprising amplifying the cell-free DNA duplex and binding the sample index to the cell-free DNA duplex. Embodiment 44. The method of any one of embodiments 22 to 42, wherein the sequencing adaptor or the second sequencing adaptor comprises a sample index. Embodiment 45. The method of any one of embodiments 40 to 44, wherein the sequencing adaptor or the second sequencing adaptor comprises a UMI. Embodiment 46 The method of any one of embodiments 22 to 45, wherein the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. Embodiment 47 The method of any one of embodiments 1 to 46, wherein the cell-free DNA duplex is obtained from a subject suspected of having cancer or determined to have cancer. Embodiment 48 The method of embodiment 47, further comprising treating the subject with an anti-cancer therapy. Embodiment 49 The method of any one of embodiments 1 to 48, further comprising obtaining cell-free DNA duplexes from the subject. Embodiment 50 The method of any one of embodiments 1 to 49, wherein the cell-free DNA duplexes are obtained from a liquid biopsy sample. Embodiment 51. The method of embodiment 50, wherein the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. Embodiment 52 The method of any one of embodiments 1 to 51, wherein the cell-free DNA duplexes are circulating tumor DNA (ctDNA) duplexes. Embodiment 53. The method of any one of embodiments 1 to 52, wherein the sequencing adaptor or second sequencing adaptor comprises an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence. Embodiment 54 The method of any one of embodiments 1 to 53, further comprising amplifying the first strand and the second strand of the cell-free DNA duplex. Embodiment 55. The method of embodiment 54, wherein amplifying comprises performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique. Embodiment 56. The method of any one of embodiments 1 to 55, wherein the sequencing comprises the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. Embodiment 57. The method of embodiment 56, wherein the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technology comprises next-generation sequencing (NGS). Embodiment 58. The method of any one of embodiments 1 to 57, wherein sequencing is carried out using a next-generation sequencer. Embodiment 59. The method of any one of embodiments 1 to 58, further comprising generating, by the one or more processors, a report indicating the length of the 3' overhang of the first strand of the cell-free DNA duplex, the length of the 3' overhang of the second strand of the cell-free DNA duplex, the length of the 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of the 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 60. The method of embodiment 59, further comprising sending the report to a healthcare provider. Embodiment 61. The method of embodiment 60, wherein the report is transmitted via a computer network or a peer-to-peer connection. Embodiment 62. The method of any one of embodiments 1-61, further comprising generating a genomic profile of the subject, comprising the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of a 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of a 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 63. The method of embodiment 62, wherein the subject's genomic profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. Embodiment 64. The method of embodiment 62 or 63, wherein the subject's genomic profile further comprises results from a nucleic acid sequencing-based test. Embodiment 65. A system comprising: one or more processors; a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receiving a first sequence read obtained by extending the 3' end of a first strand of the cell-free DNA duplex with an inosine base to fill in a 5' overhang of a second strand of the cell-free DNA duplex; attaching a sequencing adaptor to the cell-free DNA duplex; and sequencing the first strand of the cell-free DNA duplex to generate a first sequence read; determining one or more bases in the first sequence read to be soft-clipped; soft clipping bases corresponding to 3' inosine stretches of the first strand of the cell-free DNA duplex from the first sequence read; determining the length or sequence of a 5' overhang of a second strand of the cell-free DNA duplex based on soft clipping of the 3' inosine extension of the first strand of the cell-free DNA duplex. Embodiment 66. The system of embodiment 65, wherein the instructions further instruct the system to detect the presence or absence of a disease based on the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 67. The method of embodiment 66, wherein the disease is cancer. Embodiment 68. The system of any one of embodiments 65-67, wherein the first sequence read is further obtained by: attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extension of the first strand cell-free DNA duplex. Embodiment 69. The instructions, when executed by one or more processors, provide a system with: soft clipping bases corresponding to 5' inosine stretches of the first strand of the cell-free DNA duplex from the first sequence read; 69. The system of embodiment 68, further instructed to: determine the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex based on soft-clipping the 5' inosine stretch of the first strand of the cell-free DNA duplex. Embodiment 70. The system of embodiment 69, wherein the instructions further instruct the system to detect the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex. Embodiment 71. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises: extending a 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; Attaching a second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. Embodiment 72. The system of claim 71, wherein the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type. Embodiment 73. The system of any one of embodiments 65 to 72, wherein one or more bases in the first sequence read are determined to be soft-clipped by a method comprising aligning the first sequence read to a reference sequence and identifying unaligned portions of the first sequence read. Embodiment 74. The instructions, when executed by one or more processors, cause the system to: receiving second sequence reads obtained by sequencing a second strand of the cell-free DNA duplex; 74. The system of any one of embodiments 65 to 73, further instructed to: determine one or more bases in the second sequence read to be soft-clipped. Embodiment 75. The system of embodiment 74, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. Embodiment 76. The system of embodiment 74 or 75, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI). Embodiment 77. The system of embodiment 76, wherein one or more bases in a first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to a second sequence read to identify unaligned portions of the first sequence read. Embodiment 78. The system of embodiment 76 or 77, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. Embodiment 79. The system of any one of embodiments 74 to 78, wherein the second sequence read is further obtained by extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex, and binding a second sequencing adaptor to the cell-free DNA duplex. Embodiment 80. The instructions, when executed by one or more processors, cause the system to: soft clipping bases corresponding to the 3' inosine stretch of the second strand of the cell-free DNA duplex from the second sequence read; 80. The system of embodiment 79, further instructed to: determine the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex based on soft-clipping of the 3' inosine stretch of the second strand of the cell-free DNA duplex. Embodiment 81 The system of any one of embodiments 65-80, wherein extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a 3' single inosine overhang. Embodiment 82. The system of embodiment 81, wherein the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the 3' inosine overhang. Embodiment 83. A system comprising: one or more processors; a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receiving a first sequence read obtained by binding a sequencing adaptor to a 3' overhang of a first strand of the cell-free DNA duplex, extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex, and binding the 3' inosine extended end of the sequencing adaptor to a 5' end of a second strand of the cell-free DNA duplex, wherein the bound 3' inosine extended end provides a 5' inosine extended portion of the second strand of the cell-free DNA duplex; determining one or more bases in the first sequence read to be soft-clipped; soft clipping bases corresponding to 5' inosine stretches of the second strand of the cell-free DNA duplex from the first sequence read; determining the length or sequence of a 3' overhang of a first strand of the cell-free DNA duplex based on soft clipping of a 5' inosine extension of a second strand of the cell-free DNA duplex. Embodiment 84. The system of embodiment 83, wherein the instructions further instruct the system to detect the presence or absence of a disease based on the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex. Embodiment 85. The system described in embodiment 84, wherein the disease is cancer. Embodiment 86. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises: extending a 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; Attaching a second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. Embodiment 87. The system of embodiment 86, wherein the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type. Embodiment 88. The system of any one of embodiments 83 to 87, wherein one or more bases in the first sequence read are determined to be soft-clipped by a method comprising aligning the first sequence read to a reference sequence and identifying unaligned portions of the first sequence read. Embodiment 89. The instructions, when executed by one or more processors, provide a system with: receiving second sequence reads obtained by sequencing a second strand of the cell-free DNA duplex; 89. The system of any one of embodiments 83 to 88, further instructed to: determine one or more bases in the second sequence read to be soft-clipped. Embodiment 90. The system of embodiment 89, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. Embodiment 91. The system of embodiment 89 or 90, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI). Embodiment 92. The system of embodiment 91, wherein one or more bases in a first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to a second sequence read to identify unaligned portions of the first sequence read. Embodiment 93. The system of embodiment 91 or 92, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. Embodiment 94. The second sequence read comprises: Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with inosine bases to fill in the 3' overhang of the second strand of the cell-free DNA duplex; 94. The system of any one of embodiments 89-93, further obtained by binding a 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the 3' inosine extended end provides a 5' inosine extension of the first strand of the cell-free DNA duplex. Embodiment 95. The instructions, when executed by one or more processors, provide a system with: soft clipping bases corresponding to the 3' inosine stretch of the first strand of the cell-free DNA duplex from the second sequence read; 95. The system of embodiment 94, further instructed to: determine the length or sequence of a 5' overhang of a second strand of the cell-free DNA duplex based on soft-clipping of the 3' inosine stretch of the first strand of the cell-free DNA duplex. Embodiment 96. The system of embodiment 94 or 95, wherein attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises attaching a single inosine to the 3' overhang of the second strand of the cell-free DNA duplex, and attaching the second sequencing adaptor to the single inosine. Embodiment 97. The system of embodiment 96, wherein the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the single inosine attached to the 3' terminal overhang of the second strand of the cell-free DNA duplex. Embodiment 98. A system described in any one of embodiments 65 to 97, wherein the sequencing adaptor or the second sequencing adaptor comprises a UMI. Embodiment 99. The system of any one of embodiments 65 to 98, wherein the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. Embodiment 100. The system of any one of embodiments 65 to 99, further comprising a nucleic acid amplifier configured to amplify the cell-free DNA duplex and bind the sample index to the cell-free DNA. Embodiment 101. The system of embodiment 100, wherein the nucleic acid amplifier is a thermal cycler. Embodiment 102. The system of any one of embodiments 65 to 101, wherein the cell-free DNA duplexes are obtained from a subject suspected of having cancer or determined to have cancer. Embodiment 103. The system of any one of embodiments 65 to 102, wherein the cell-free DNA duplex is obtained from a subject. Embodiment 104. The system of any one of embodiments 65 to 103, wherein the cell-free DNA duplexes are obtained from a liquid biopsy sample. Embodiment 105. The system of embodiment 104, wherein the sample is a liquid biopsy sample and includes blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. Embodiment 106. The system of any one of embodiments 65 to 105, wherein the cell-free DNA duplex is a circulating tumor DNA (ctDNA) duplex. Embodiment 107. The system of any one of embodiments 65 to 106, wherein the sequencing adapter or the second sequencing adapter comprises an amplification primer binding site, a flow cell adapter sequence, or a substrate adapter sequence. Embodiment 108. The system of any one of embodiments 65 to 107, further comprising a nucleic acid amplifier configured to amplify the first and second strands of the cell-free DNA duplex. Embodiment 109. The system of embodiment 108, wherein amplifying comprises performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique. Embodiment 110. The system of any one of embodiments 65 to 109, further comprising a sequencer configured to sequence the first strand of the cell-free DNA duplex and / or the second strand of the cell-free DNA duplex. Embodiment 111. A system described in any one of embodiments 65 to 110, wherein the sequencing comprises the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. Embodiment 112. The system described in embodiment 111, wherein the sequencing comprises massively parallel sequencing and the massively parallel sequencing technology comprises next-generation sequencing (NGS). Embodiment 113. A system described in any one of embodiments 65 to 112, wherein sequencing is performed using a next-generation sequencer. Embodiment 114. The system of any one of embodiments 65 to 113, wherein the instructions, when executed by the one or more processors, further instruct the system to generate a report indicating the length of the 3' overhang of the first strand of the cell-free DNA duplex, the length of the 3' overhang of the second strand of the cell-free DNA duplex, the length of the 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of the 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 115. The system described in embodiment 114, wherein the instructions, when executed by one or more processors, further instruct the system to send a report to a healthcare provider. Embodiment 116. A system described in embodiment 114 or 115, wherein the report is transmitted via a computer network or a peer-to-peer connection. Embodiment 117. The system of any one of embodiments 65 to 116, wherein the instructions, when executed by the one or more processors, further instruct the system to generate, by the one or more processors, a genomic profile of the subject, comprising the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of a 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of a 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 118. The method of embodiment 117, wherein the subject's genomic profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. Embodiment 119. The method of embodiment 117 or 118, wherein the subject's genomic profile further comprises results from a nucleic acid sequencing-based test. Embodiment 120. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to: receiving a first sequence read obtained by extending the 3' end of a first strand of the cell-free DNA duplex with an inosine base to fill in a 5' overhang of a second strand of the cell-free DNA duplex; attaching a sequencing adaptor to the cell-free DNA duplex; and sequencing the first strand of the cell-free DNA duplex to generate a first sequence read; determining one or more bases in the first sequence read to be soft-clipped; soft clipping bases corresponding to 3' inosine stretches of the first strand of the cell-free DNA duplex from the first sequence read; determining the length or sequence of a 5' overhang of a second strand of the cell-free DNA duplex based on soft-clipping of a 3' inosine stretch of a first strand of the cell-free DNA duplex. Embodiment 121. The storage medium of embodiment 120, wherein the instructions further instruct the system to detect the presence or absence of a disease based on the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 122. A storage medium described in embodiment 121, wherein the disease is cancer. Embodiment 123. The storage medium of any one of embodiments 120 to 122, wherein the first sequence read is further obtained by: attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; and attaching the 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the attached 3' inosine extended end provides a 5' inosine extended end of the first strand of the cell-free DNA duplex. Embodiment 124. The instructions, when executed by one or more processors, provide a system with: soft clipping bases corresponding to 5' inosine stretches of the first strand of the cell-free DNA duplex from the first sequence read; 124. The storage medium of embodiment 123, further instructing the storage medium to: determine the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex based on soft-clipping the 5' inosine stretch of the first strand of the cell-free DNA duplex. Embodiment 125. The storage medium of embodiment 124, wherein the instructions further instruct the system to detect the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex. Embodiment 126. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises: extending a 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; Attaching a second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. Embodiment 127. The storage medium of embodiment 126, wherein the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type. Embodiment 128. A storage medium described in any one of embodiments 120 to 127, wherein one or more bases in the first sequence read are determined to be soft-clipped by a method comprising aligning the first sequence read to a reference sequence and identifying unaligned portions of the first sequence read. Embodiment 129. The instructions, when executed by one or more processors, cause the system to: receiving second sequence reads obtained by sequencing a second strand of the cell-free DNA duplex; 129. A storage medium according to any one of embodiments 120 to 128, further instructing the storage medium to: determine one or more bases in the second sequence read to be soft-clipped. Embodiment 130. The storage medium of embodiment 129, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. Embodiment 131. A storage medium described in embodiment 129 or 130, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI). Embodiment 132. The storage medium of embodiment 131, wherein one or more bases in a first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to a second sequence read to identify unaligned portions of the first sequence read. Embodiment 133. A storage medium described in embodiment 131 or 132, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. Embodiment 134. The storage medium of any one of embodiments 129 to 133, wherein the second sequence read is further obtained by extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex, and binding a second sequencing adaptor to the cell-free DNA duplex. Embodiment 135. The instructions, when executed by one or more processors, cause the system to: soft clipping bases corresponding to the 3' inosine stretch of the second strand of the cell-free DNA duplex from the second sequence read; 135. The storage medium of embodiment 134, further instructing the storage medium to: determine the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex based on soft-clipping of the 3' inosine stretch of the second strand of the cell-free DNA duplex. Embodiment 136. The storage medium of any one of embodiments 120 to 135, wherein extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a 3' single inosine overhang. Embodiment 137. The storage medium of embodiment 136, wherein the sequencing adapter comprises a 3' cytosine overhang that is complementary to the 3' single inosine overhang. Embodiment 138. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the system, cause the system to: receiving a first sequence read obtained by binding a sequencing adaptor to a 3' overhang of a first strand of the cell-free DNA duplex, extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex, and binding the 3' inosine extended end of the sequencing adaptor to a 5' end of a second strand of the cell-free DNA duplex, wherein the bound 3' inosine extended end provides a 5' inosine extended portion of the second strand of the cell-free DNA duplex; determining one or more bases in the first sequence read to be soft-clipped; soft clipping bases corresponding to 5' inosine stretches of the second strand of the cell-free DNA duplex from the first sequence read; determining the length or sequence of a 3' overhang of a first strand of the cell-free DNA duplex based on soft-clipping of a 5' inosine stretch of a second strand of the cell-free DNA duplex. Embodiment 139. The storage medium of embodiment 138, wherein the instructions further instruct the system to detect the presence or absence of a disease based on the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex. Embodiment 140. The storage medium of embodiment 139, wherein the disease is cancer. Embodiment 141. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises: extending a 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; Attaching a second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex. Embodiment 142. The storage medium of embodiment 141, wherein the 3' overhang of the second strand of the cell-free DNA duplex is extended using nucleotide bases of the same base type. Embodiment 143. A storage medium described in any one of embodiments 138 to 142, wherein one or more bases in the first sequence read are determined to be soft-clipped by a method comprising aligning the first sequence read to a reference sequence and identifying unaligned portions of the first sequence read. Embodiment 144. The instructions, when executed by one or more processors, provide a system: receiving second sequence reads obtained by sequencing a second strand of the cell-free DNA duplex; 144. A storage medium described in any one of embodiments 138 to 143, further instructing the storage medium to: determine one or more bases in the second sequence read to be soft-clipped. Embodiment 145. The storage medium of embodiment 144, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read. Embodiment 146. A storage medium described in embodiment 144 or 145, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI). Embodiment 147. The storage medium of embodiment 146, wherein one or more bases in a first sequence read are determined to be soft-clipped according to a method comprising aligning the first sequence to a second sequence read to identify unaligned portions of the first sequence read. Embodiment 148. A storage medium described in embodiment 146 or 147, wherein one or more bases in the second sequence read are determined to be soft-clipped according to a method comprising aligning the second sequence to the first sequence read and identifying unaligned portions of the second sequence read. Embodiment 149. The second sequence read comprises: Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with inosine bases to fill in the 3' overhang of the second strand of the cell-free DNA duplex; 149. The storage medium of any one of embodiments 144 to 148, further obtained by binding a 3' inosine extended end of the second sequencing adaptor to the 5' end of the first strand of the cell-free DNA duplex, wherein the 3' inosine extended end provides a 5' inosine extended portion of the first strand of the cell-free DNA duplex. Embodiment 150. The instructions, when executed by one or more processors, cause the system to: soft clipping bases corresponding to the 3' inosine stretch of the first strand of the cell-free DNA duplex from the second sequence read; 150. The storage medium of embodiment 149, further instructing the storage medium to: determine the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on soft-clipping the 3' inosine stretch of the first strand of the cell-free DNA duplex. Embodiment 151. The storage medium of embodiment 149 or 150, wherein attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex comprises attaching a single inosine to the 3' overhang of the second strand of the cell-free DNA duplex, and attaching the second sequencing adaptor to the single inosine. Embodiment 152. The storage medium of embodiment 151, wherein the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the single inosine attached to the 3' end overhang of the second strand of the cell-free DNA duplex. Embodiment 153. A storage medium described in any one of embodiments 120 to 152, wherein the sequencing adaptor or the second sequencing adaptor comprises a UMI. Embodiment 154. A storage medium according to any one of embodiments 120 to 153, wherein the sequencing adaptor or the second sequencing adaptor is a Y-shaped sequencing adaptor. Embodiment 155. A storage medium according to any one of embodiments 120 to 154, wherein the cell-free DNA duplex is obtained from a subject suspected of having cancer or determined to have cancer. Embodiment 156. A storage medium according to any one of embodiments 120 to 155, wherein the cell-free DNA duplex is obtained from a subject. Embodiment 157. A storage medium according to any one of embodiments 120 to 156, wherein the cell-free DNA duplex is obtained from a liquid biopsy sample. Embodiment 158. A storage medium according to embodiment 157, wherein the sample is a liquid biopsy sample, including blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. Embodiment 159. A storage medium according to any one of embodiments 120 to 158, wherein the cell-free DNA duplex is a circulating tumor DNA (ctDNA) duplex. Embodiment 160. A storage medium according to any one of embodiments 120 to 159, wherein the sequencing adaptor or the second sequencing adaptor comprises an amplification primer binding site, a flow cell adaptor sequence, or a substrate adaptor sequence. Embodiment 161. A storage medium described in any one of embodiments 120 to 160, wherein the sequencing comprises the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. Embodiment 162. A storage medium described in embodiment 161, wherein the sequencing includes massively parallel sequencing and the massively parallel sequencing technology includes next-generation sequencing (NGS). Embodiment 163. A storage medium described in any one of embodiments 120 to 162, wherein sequencing is performed using a next-generation sequencer. Embodiment 164. The storage medium of any one of embodiments 120 to 163, wherein the instructions, when executed by one or more processors, further instruct the system to generate a report indicating the length of the 3' overhang of the first strand of the cell-free DNA duplex, the length of the 3' overhang of the second strand of the cell-free DNA duplex, the length of the 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of the 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 165. The storage medium of embodiment 164, wherein the instructions, when executed by one or more processors, further instruct the system to send a report to a healthcare provider. Embodiment 166. A storage medium described in embodiment 164 or 165, in which the report is transmitted via a computer network or a peer-to-peer connection. Embodiment 167. The storage medium of any one of embodiments 120 to 166, wherein the instructions, when executed by the one or more processors, further instruct the system to generate, by the one or more processors, a genomic profile of the subject, comprising the length of a 3' overhang of the first strand of the cell-free DNA duplex, the length of a 3' overhang of the second strand of the cell-free DNA duplex, the length of a 5' overhang of the first strand of the cell-free DNA duplex, and / or the length of a 5' overhang of the second strand of the cell-free DNA duplex. Embodiment 168. The storage medium of embodiment 167, wherein the subject's genomic profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. Embodiment 169. The storage medium of embodiment 167 or 168, wherein the subject's genomic profile further comprises results from a nucleic acid sequencing-based test.
[0179] From the foregoing, it should be understood that, while particular implementations of the disclosed method and system have been illustrated and described, various modifications can be made thereto and are contemplated herein. It is also not intended that the present invention be limited by the specific examples provided herein. While the present invention has been described with reference to the above specification, the descriptions and illustrations of preferred embodiments herein are not meant to be construed in a limiting sense. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon various conditions and variables. Various modifications in form and details of the embodiments of the present invention will be apparent to those skilled in the art. Accordingly, the present invention is also intended to encompass any such modifications, variations, and equivalents.
Claims
1. 1. A method for determining cell-free DNA topology, comprising: extending the 3' end of a first strand of a cell-free DNA duplex with an inosine base to fill in a 5' overhang of a second strand of said cell-free DNA duplex; ligating a sequencing adaptor to the cell-free DNA duplex; sequencing the first strand of the cell-free DNA duplex to generate first sequence reads; determining, by one or more processors, one or more bases in the first sequence read to be soft-clipped; and soft-clipping, by the one or more processors, bases corresponding to 3' inosine stretches of the first strand of the cell-free DNA duplex from the first sequence reads; determining, by the one or more processors, the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex based on the soft clipping of the 3' inosine stretch of the first strand of the cell-free DNA duplex.
2. 10. The method of claim 1, further comprising detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 5' overhang of the second strand of the cell-free DNA duplex.
3. The method of claim 2 , wherein the disease is cancer.
4. Attaching a second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' end of the second sequencing adaptor with an inosine base to fill in the 3' overhang of the second strand of the cell-free DNA duplex; ligating a 3' inosine-extended end of the second sequencing adaptor to a 5' end of the first strand of a cell-free DNA duplex, wherein the attached 3' inosine-extended end provides a 5' inosine-extended portion of the first strand cell-free DNA duplex; determining, by one or more processors, one or more bases attached to the 5′ end of the first sequence read to be soft-clipped; soft-clipping, by the one or more processors, bases from the first sequence reads that correspond to the 5' inosine stretch of the first strand of the cell-free DNA duplex; 10. The method of claim 1, further comprising determining, by the one or more processors, a length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex based on the soft clipping of the 5' inosine stretch of the first strand of the cell-free DNA duplex.
5. 5. The method of claim 4, comprising detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 3' overhang of the second strand of the cell-free DNA duplex.
6. Attaching the second sequencing adaptor to the 3' overhang of the second strand of the cell-free DNA duplex; extending the 3' overhang of the second strand of the cell-free DNA duplex to provide a 3' extension, wherein the second sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the second strand of the cell-free DNA duplex; and binding the second sequencing adaptor to the 3' extension of the second strand of the cell-free DNA duplex.
7. 2. The method of claim 1, wherein determining the one or more bases in the first sequence read to be soft-clipped comprises aligning the first sequence read to a reference sequence to identify unaligned portions of the first sequence read.
8. sequencing the second strand of the cell-free DNA duplex to generate second sequence reads; 10. The method of claim 1, further comprising determining, by one or more processors, one or more bases in the second sequence reads to be soft-clipped.
9. 9. The method of claim 8, wherein determining the one or more bases in the second sequence read to be soft-clipped comprises aligning the second sequence to a reference sequence to identify unaligned portions of the second sequence read.
10. 9. The method of claim 8, wherein the first sequence read and the second sequence read are associated via a unique molecular identifier (UMI).
11. determining the one or more bases in the first sequence read to be soft-clipped comprises aligning the first sequence to the second sequence read to identify unaligned portions of the first sequence read; or 11. The method of claim 10, wherein determining the one or more bases in the second sequence read to be soft-clipped or soft-clipped comprises aligning the second sequence to the first sequence read to identify unaligned portions of the second sequence read.
12. extending the 3' end of the second strand of the cell-free DNA duplex with an inosine base to fill in the 5' overhang of the first strand of the cell-free DNA duplex; ligating a second sequencing adaptor to the cell-free DNA duplex; soft-clipping, by the one or more processors, bases corresponding to 3' inosine stretches of the second strand of the cell-free DNA duplex from the second sequence reads; and determining, by the one or more processors, the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex based on the soft clipping of the 3' inosine stretch of the second strand of the cell-free DNA duplex.
13. 13. The method of claim 12, comprising detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 5' overhang of the first strand of the cell-free DNA duplex.
14. 10. The method of claim 1, wherein extending the 3' end of the first strand of the cell-free DNA duplex comprises forming a single 3' inosine overhang.
15. 15. The method of claim 14, wherein the sequencing adaptor comprises a 3' cytosine overhang that is complementary to the 3' inosine overhang.
16. 1. A method of making a sequencing construct, comprising: Attaching a sequencing adaptor to the 3' overhang of the first strand of the cell-free DNA duplex; extending the 3' end of the sequencing adaptor with an inosine base to fill in the 3' overhang of the first strand of the cell-free DNA duplex; and attaching a 3' inosine-extended end of the sequencing adaptor to a 5' end of a second strand of the cell-free DNA duplex, wherein the attached 3' inosine-extended end provides a 5' inosine-extended portion of the second strand of the cell-free DNA duplex.
17. Attaching the sequencing adaptor to the 3' overhang of the first strand of the cell-free DNA duplex; extending the 3' overhang of the first strand of the cell-free DNA duplex to provide a 3' extension, wherein the sequencing adaptor comprises a 3' overhang that is complementary to the 3' extension of the first strand of the cell-free DNA duplex; and binding the sequencing adaptor to the 3' extension of the first strand of the cell-free DNA duplex.
18. 1. A method for determining cell-free DNA topology, comprising: Producing the sequencing construct according to the method of any one of claims 22 to 24; sequencing the second strand of the cell-free DNA duplex to generate first sequence reads; determining, by one or more processors, one or more bases in the first sequence read to be soft-clipped; and soft-clipping, by the one or more processors, bases corresponding to the 5' inosine stretch of the second strand of the cell-free DNA duplex from the first sequence reads; determining, by the one or more processors, the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex based on the soft clipping of the 5' inosine stretch of the second strand of the cell-free DNA duplex.
19. 20. The method of claim 18, further comprising detecting, by the one or more processors, the presence or absence of a disease based on the length or sequence of the 3' overhang of the first strand of the cell-free DNA duplex.
20. 20. The method of claim 19, wherein the disease is cancer.
Citation Information
Patent Citations
Methods for generating nucleic acid libraries
JP2019532014A
Methods and systems for sequence alignment and variant calling
JP2020505052A
Cell-free DNA damage assay and its clinical application
JP2021531016A
Systems and methods for determining consensus base calls in nucleic acid sequencing
US20210065847A1
Methods for duplex sequencing of cell-free DNA and applications thereof
WO2020264565A1