Suppression of nucleic acid errors
Deep-depth WGS with in silico error correction and UMIs in the Ultima Genomics platform addresses high costs and error rates in ctDNA detection, enabling accurate and cost-effective cancer monitoring in low-burden diseases.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2026-04-03
AI Technical Summary
Current sequencing technologies face high costs and error rates in detecting low-burden diseases, particularly in cancer detection, due to the dilution of circulating tumor DNA (ctDNA) in plasma samples, limiting the depth and accuracy of whole-genome sequencing (WGS) for reliable ctDNA analysis.
Employing deep-depth WGS with in silico error correction using the Ultima Genomics' mnSBS platform, incorporating unique molecular identifiers (UMIs) and double-strand sequencing to reduce sequencing errors and enhance detection sensitivity.
Achieves accurate ctDNA detection at a low error rate of approximately 10⁻⁷, enabling reliable cancer monitoring and detection of low-burden diseases without requiring corresponding tumor sequencing, with cost-effective sequencing at $1/Gb.
Smart Images

Figure 2026510423000001_ABST
Abstract
Description
[Technical Field]
[0001] Related applications This application claims priority based on U.S. Provisional Application No. 63 / 380915, filed on 25 October 2022, which is incorporated herein by reference in its entirety.
[0002] The present invention is generally related to the field of medical diagnostics. In particular, embodiments of this disclosure relate to methods and systems for reducing sequencing error rates in cancer detection and other fields requiring low-error sequencing. [Background technology]
[0003] Monitoring of circulating cell-free DNA (cfDNA) has shown to be a promising clinical tool for non-invasive cancer detection. While analysis of cancer-specific epigenetic markers such as DNA methylation and histone modifications has been applied to cfDNA for the detection of various cancers, a mutation-based approach using direct genomic sequencing of somatic variants found in circulating tumor DNA (ctDNA) provides more specific and clinically relevant information. Therefore, ctDNA genomic sequencing is preferred for clinical applications, particularly in low-disease-burden situations, such as early cancer screening, detection of minimal residual disease (MRD) after treatment or surgery, and monitoring for recurrence of novel resistance mutations for treatment guidance. In these situations, where the tumor ratio is low, a method with high sensitivity is required for reliable detection.
[0004] Existing methods for ctDNA detection utilize targeted sequencing protocols that increase the number of genomes sequenced at targeted sites. However, high-throughput targeted sequencing rapidly depletes the genomes available for sequencing (1,000–10,000 genome equivalents (GE) per mL of plasma), thus setting a design-based upper limit for ctDNA detection. Once a limited number of GEs have been sequenced, further increasing the sequencing depth at target sites offers no further benefit. Alternatively, to overcome these limitations, whole-genome sequencing (WGS) approaches leverage broad coverage instead of depth, eliminating reliance on single-site detection to enhance ctDNA characterization in low tumor-to-tumor environments. For example, a recent technique, MRDetect, uses mutation profiles from primary tumors to provide information for detecting tumor-derived single nucleotide variants (SNVs) across the entire genome in cfDNA, where the number of available GEs is no longer a limiting factor for ctDNA detection success.
[0005] The detection challenges posed by dilution are significant, requiring extensive, accurate, and deep cfDNA sequencing. Therefore, whole-genome, low-error, and high-coverage methods are necessary for reliable ctDNA analysis. However, the costs associated with these approaches are often exorbitant, especially for clinical applications. While the cost of genome sequencing has decreased rapidly since the introduction of high-throughput next-generation sequencing, this decline has stalled in recent years. Consequently, sequencing costs remain a major barrier to performing deep WGS in liquid biopsies with low tumor ratios (approximately 10⁻⁵) where shallow WGS is insufficient for ctDNA detection, particularly in clinically important applications. To estimate the cost of performing WGS to successfully detect ctDNA at such ratios, it is possible to model the probability of detecting a single mutation in a cfDNA sample, considering the number of gene-enhancing mutants (GEs), tumor ratio, and sequencing depth. 17To detect single mutations in tumor-derived DNA at this low ratio level, it is estimated that sequencing depths of more than 100-fold are required, which makes the per-sample cost of WGS extremely high for large-scale applications of the established technology (approximately US$2,000 per sample using an Illumina Novaseq S4 flow cell with v1.5 reagent costs).
[0006] Recently, Ultima Genomics has developed a new low-cost, high-throughput sequencing method utilizing mostly natural sequencing-by-synthesis (mnSBS). This Ultima sequencing platform generates approximately 10 billion single-ended reads per run at $1 per GB, resulting in a significant reduction in sequencing costs compared to current platforms. This cost-effectiveness makes it promising for many genomics applications, and the approach is currently being applied to reference samples for Genome-In-A-Bottle and 1000 Genomes, as well as single-cell RNA-seq studies. However, mnSBS / Ultima sequencing has not been used for clinical cfDNA samples for ctDNA sequencing. Furthermore, it should be noted that the error rate profile of this new sequencing method has not been fully characterized, nor has it been rigorously compared with competing technologies. The important point is that, in order to apply ctDNA to the monitoring of clinical diseases in the future, it is especially important to accurately estimate the error rate because ctDNA has a low negative charge and high detection sensitivity.
[0007] Therefore, there is a need for methods to reduce sequencing error rates in disease detection. The technical challenges posed by the dilution of ctDNA in low-burden disease environments can be overcome by increasing sequencing depth with lower error rates. Error suppression techniques using specific molecular identifiers (UMIs) or double-strand sequencing can improve the accuracy of distinguishing true somatic variant calls from sequencing-induced errors, and when combined with deep genome sequencing, can optimize the detection of low-burden diseases in clinical settings. [Overview of the Initiative]
[0008] This specification investigates the usefulness of deep-depth WGS for ctDNA detection by sequencing circulating cell-free DNA reads obtained from plasma samples of healthy controls, cancer patients, and patient-derived xenografted mouse models using Ultima Genomics' mnSBS sequencing platform. The initial proof-of-principle study demonstrates that deep-depth WGS (approximately 100x) with in silico error correction enables ctDNA detection within a range of 1 in 1,000,000. Leveraging the high cost-efficiency and high throughput of Ultima Genomics sequencing (10⁹ reads per run, $1 / Gb), a double-strand sequencing library with high cell-free DNA coverage of 2.7 × 10⁶ reads is constructed. -7 This enables the achievement of a low error rate. This allows for accurate assessment of the disease burden in melanoma patients after treatment, even without tumor information. Thus, the usefulness of deep WGS in clinical samples for ctDNA detection is demonstrated.
[0009] This disclosure discloses a system and method for detecting cancer while reducing sequencing error rates, in a particular aspect of this disclosure.
[0010] In one embodiment, the method includes extracting DNA from a group of plasma samples, preparing a whole-genome library using a double-strand adapter, the whole-genome library being prepared by ligating a double-strand adapter having a unique molecular identifier (UMI) to each end of multiple strands of the extracted DNA, and amplifying the extracted DNA by a first PCR, preparing the whole-genome library, selecting a subset of the whole-genome library, amplifying the subset by a second PCR to increase the PCR duplicate, sequencing multiple double-strand reads from the amplified subset, aligning the multiple double-strand reads to the host genome and denoising these multiple double-strand reads based on this alignment, detecting the presence of a variant in at least one of the multiple double-strand reads, determining the signature of the variant, comparing the signature of the variant with a group of signatures of disease-specific variants, and determining the type of disease based on the comparison.
[0011] The accompanying drawings incorporated herein and forming part thereof illustrate various exemplary embodiments and, together with the description herein, are intended to illustrate the principles of the embodiments of this disclosure. [Brief explanation of the drawing]
[0012] [Figure 1] This flowchart shows a method for detecting mutation signatures to determine the state of cancer using the technology disclosed herein. [Figure 2A] This is a series of graphs showing the error rate and sequencing coverage when the tumor ratio is 10⁻⁵ or less, using the technology disclosed herein. [Figure 2B] This flowchart shows the pre-analysis process for preparing a cfDNA library using the technology disclosed herein. [Figure 2C] This graph shows the sequencing depth of the corresponding Illumina and Ultima datasets using the techniques disclosed herein. [Figure 2D] This is a comparison of normalized read coverage of sequenced matching cfDNA samples using the techniques disclosed herein. [Figure 2E] This is a comparison of the tumor ratios of copy number-based variants (CNVs) and single-nucleotide variants (SNVs) using the techniques disclosed herein. [Figure 2F] This is a graph of an in silico mixing test using the technology disclosed herein. [Figure 3A] This is a series of graphs showing double-strand whole-genome sequencing (WGS) of samples from mice (left) and patients (right) using the techniques disclosed herein. [Figure 3B] This graph shows a comparison of allele frequencies of variants calculated using unfiltered sequencing reads with the technology disclosed herein. [Figure 3C] This graph shows a comparison between the model allele ratios of patients with progressive disease at the double-strand corrected position and the copy number-based tumor ratio estimates, using the technology disclosed herein. [Figure 3D] This graph shows an example of trinucleotide frequencies from UV signatures associated with melanoma. [Figure 3E] This is a conditional comparison of cosine similarity with SBS7 or SBS1B mutations using the techniques disclosed herein. [Figure 3F] This graph shows the signature score and ctDNA detection in an in silico mixed test of metastatic melanoma samples using the technology disclosed herein. [Figure 3G] This graph shows a series of signature scores for melanoma signature 7 in a plasma cfDNA sample using double-stranded WGS, as disclosed herein. [Figure 3H] This graph shows the estimated tumor ratio in samples with high signature scores, as determined by the technology disclosed herein. [Figure 4]A series of graphs showing the frequency of cfDNA fragment lengths in a single-end Ultima sequencing dataset paired with paired-end Illumina sequencing, according to the technology disclosed herein. [Figure 5A] A graph showing the UG-specific blacklist according to the technology disclosed herein. [Figure 5B] A diagram showing the overlap between the UG blacklist and other low-confidence regions according to the technology disclosed herein. [Figure 5C] A graph showing the overlap between SNVs and low-confidence regions in melanoma tumor tissue according to the technology disclosed herein. [Figure 6A] A heatmap of cosine similarity in cancer-free samples and high-load ctDNA samples according to the technology disclosed herein. [Figure 6B] A box plot of cosine similarity for three correction methods in the same cancer-free samples and high-load ctDNA samples according to the technology disclosed herein. [Figure 7A] A graph of deconvolving double-strand corrected mutations into representative mutation signatures. [Figure 7B] A correlation plot between the age at cancer diagnosis and the number of clock-like mutations due to SBS1A and SBS1B according to the technology disclosed herein. [Figure 8] A graph of tumor ratio estimates based on tumor-nonspecific copy numbers in cancer-free control samples and preoperative melanoma plasma according to the technology disclosed herein. [Figure 9A] A graph showing the homopolymer size between two PCR duplicates according to the technology disclosed herein. [Figure 9B] A graph showing the homopolymer size between the read and the aligned reference according to the technology disclosed herein. [Figure 9C]This graph shows the frequency of homopolymer sizes across the entire human genome, as demonstrated by the technologies disclosed herein. [Figure 9D] This graph shows the indel call accuracy for each PCR overlap family size using the technology disclosed herein. [Figure 10] This graph shows single-nucleotide variant analysis in corresponding Ultima and Illumina sequencing datasets using the techniques disclosed herein. [Figure 11A] This is a flowchart of a sequencing process that provides predictable and error-tolerant motifs using the technology disclosed herein. [Figure 11B] This graph shows the error rates for each sequencing platform using the technologies disclosed herein. [Figure 12A] This is a graph of double-stranded WGS libraries from three start inputs sequenced with 1 to 13x coverage using the technology disclosed herein. [Figure 12B] This graph shows the overlap rate of the samples in Figure 12A. [Figure 12C] This graph shows the effect of downsampling experiments using the techniques disclosed herein. [Figure 12D] This graph shows that double-stranded coverage is significantly higher under fixed coverage conditions using the technology disclosed herein. [Figure 12E] This graph compares the number of double-stranded mutants detected using fgbio and decision trees, as disclosed herein. [Figure 13] This graph shows the mutation error rate in mouse PDX samples using the technology disclosed herein. [Figure 14] This is a bar graph showing the number of preoperative samples demonstrated in validation experiments using the technology disclosed herein. [Figure 15] This graph shows the detection of chemotherapy mutation signatures in free plasma DNA using the technology disclosed herein. [Figure 16A] This is a bar graph showing the apobec signature and its measured values using the technology disclosed herein. [Figure 16B] This is a bar graph showing the apobec signature and its measured values using the technology disclosed herein. [Figure 17] This is a computing node according to an embodiment of the disclosure. [Modes for carrying out the invention]
[0013] Examples of embodiments of the present disclosure are shown here in detail in the accompanying drawings. Wherever possible, identical or similar parts are shown with the same reference numeral throughout the drawings.
[0014] The systems, devices, and methods disclosed herein are described in detail with reference to the drawings and with illustrative examples. The examples considered herein are for illustrative purposes only and are provided to aid in the description of the apparatus, devices, systems, and methods described herein. No feature or component shown in the drawings or described below should be construed as essential to any particular implementation of these devices, systems, or methods unless specifically designated as essential.
[0015] Furthermore, regarding any method described herein, whether or not it is described in conjunction with a flow diagram, unless otherwise specified or required by context, the order of each step explicitly or implicitly indicated in the execution of the method does not implicitly mean that these steps must be performed in the order presented, but rather that they may be performed in a different order or in parallel.
[0016] As used herein, the term “exemplary” means “an example,” not “ideal.” Furthermore, as used herein, the terms “one (a)” and “one (an)” are not limited in quantity, but indicate that there is one or more items being referred to.
[0017] Cell-free DNA (cfDNA) sequencing for low-burden cancer monitoring is constrained by the dilution of circulating tumor DNA (ctDNA), the abundance of genomic material in plasma samples, and the pre-processing error rate resulting from library preparation and sequencing errors. Historically, sequencing costs have driven the development of deep-target sequencing approaches to overcome dilution in ctDNA detection; however, these techniques are limited by the abundance of cfDNA in samples, impinging an upper limit on the maximum coverage depth in targeted panels.
[0018] Whole-genome sequencing (WGS) is an orthogonal approach to ctDNA-based cancer detection, overcoming the low abundance of cfDNA and integrating signals across the entire tumor mutation landscape, thereby replacing depth with breadth. However, the high cost of WGS limits its practical coverage depth and widespread adoption.
[0019] The reduction in sequencing costs enables advanced ctDNA cancer monitoring via WGS. A newly introduced low-cost WGS (Ultima Genomics, $1 / Gb) was applied to plasma samples with approximately 120x coverage. While copy number and single nucleotide variation profiles were comparable between the corresponding Ultima and Illumina datasets, the deeper WGS coverage allows for ctDNA detection in the range of one part per million. Furthermore, these low-cost sequencing methods were also used for double-strand error-corrected sequencing on a whole-genome scale, showing approximately 3,000x error reduction compared to raw sequencing reads in plasma from patient-derived xenograft mouse models, and a low error rate of approximately 10⁻⁷ in plasma samples from metastatic melanoma patients. Highly denoised plasma WGS can be used for cancer monitoring in more challenging low-load melanoma situations without corresponding tumor sequencing. In this context, double-strand corrected WGS enables disease monitoring using known mutation signature patterns without the need for corresponding tumors, opening the way for novel cancer monitoring.
[0020] Deep whole-genome sequencing (DEEP WGS) can be used for ctDNA detection using low-pass sequencing. Low-cost WGS (Ultima Genomics, $1 / Gb) can be used on plasma samples with 120x coverage. Copy number and single nucleotide variation profiles were comparable between the corresponding Ultima and Illumina datasets, but the deeper coverage of WGS enabled ctDNA detection in the range of 1 in 1,000,000. Furthermore, when these low-cost sequencing methods were used to perform double-strand error correction on a whole-genome scale, an error reduction of approximately 1 / 3000 was achieved in plasma from patient-derived xenografted mouse models compared to raw sequencing reads, and approximately 10 in plasma samples from patients with metastatic melanoma. -7The error rate was shown. Highly denoised plasma WGS was utilized for cancer monitoring in the more challenging situation of low-load melanoma where corresponding tumor sequencing was unavailable. In this situation, double-strand corrected WGS made it possible to utilize known mutation signature patterns for disease monitoring without requiring a corresponding tumor, opening the way for novel cancer monitoring. Sequencing was performed using Ultima Genomics' sequencing platform from 31 plasma samples (n=8 healthy control samples, n=19 cancer patient samples, n=4 patient-derived xenograft mouse samples) at a rate of 2.6 × 10⁶ 9 ±1.4 × 10 8 It is possible to sequence circulating cfDNA reads. Deep WGS with in silico error correction (approximately 100x) enables ctDNA detection at the 1 in 1,000,000 level on the sequenced genome. Ultima Genomics offers cost-effective and ultra-high throughput sequencing (10 per run). 9 Using reads ($1 / Gb), a double-stranded sequencing library of cfDNA was prepared with 4.15 ± 3.23 coverage (range 1.02 to 11.55), resulting in 2.7 × 10⁶ -7This enables the achievement of a low error rate. Taken together, the results of sequencing and in silico error correction demonstrate the feasibility and usefulness of deep WGS for ctDNA detection in clinical samples. High-throughput short-read sequencing has revolutionized the field of liquid biopsy, and sequencing costs have historically decreased at a faster rate than Moore's Law. However, cost reductions have stagnated over the past decade. The emergence of high-throughput and low-cost sequencing platforms will enable deeper and more extensive sequencing of samples and patient populations. To this end, a comparative analysis of Ultima and Illumina short-read sequencing platforms was conducted, demonstrating that these two approaches have comparable tumor information-driven analytical capabilities in circulating tumor DNA detection (Figures 2A-2F). Importantly, these tumor information-driven approaches enable ctDNA detection at a range of one in a million, thereby making such analyses suitable for minimal residual disease detection. To further extend and demonstrate the usefulness of this technology, we tackled the more challenging task of ctDNA detection in tumor-nonspecific environments where the disease state or tumor origin is unknown. In this critical clinical setting, whole-genome double-strand correction was employed to achieve a low error rate, enabling ctDNA detection in a preoperative setting by deconvolving the accumulation of cell-free DNA mutations into a representative mutation signature without corresponding tumor sequencing (Figures 3A-3H). Compared with commonly used off-the-shelf panels, whole-genome analysis offers the advantage of sequencing breadth and enables the detection of rare tumor-derived mutations that may not be present in targeted panels. Those skilled in the art can envision that these methods can be utilized for novel cancer monitoring in low-burden disease scenarios, providing a powerful tool for diagnosing cancer and detecting recurrence at the earliest possible stage, ultimately leading to improved overall patient outcomes.In addition to detecting new cancers, this method can also be used for cancer screening (for example, screening for bladder cancer, melanoma, lung cancer, Lynch syndrome-related cancers, and BRCA syndrome-related cancers based on APOBEC, UV, tobacco, MSI, and BRCA signatures).
[0021] This method is not only useful for novel detection (e.g., signatures) for tumor monitoring, but also for tumor information-based approaches. In this case, double-strand sequencing reduces errors and improves signal resolution for detecting the accumulation of mutations already identified in tumors. Figure 2C shows the use of a signature-independent tumor information-based approach.
[0022] In one embodiment, this method is useful not only for monitoring but also for non-invasive whole-genome characterization of mutations in cancer (e.g., to identify actionable driver mutations or mutations that stratify patients for specific therapies). This is achieved by reducing various types of errors. Figure 2B shows this characterization, which is independent of signature analysis.
[0023] In one embodiment, this method enables non-invasive detection and discovery of driver mutations in somatic mosaicism. Figure 2E shows the detection of non-malignant mutations, specifically clonal hematopoietic mutations.
[0024] Figure 1 is a flowchart illustrating an exemplary method for detecting variants in DNA using double-strand sequencing and noise reduction, according to an exemplary embodiment of the present disclosure. For example, exemplary method 100 (e.g., steps 102-118) may be performed automatically by a processor or partially on user request.
[0025] According to one embodiment, an exemplary method 100 for detecting a variant may include one or more of the following steps. In step 102, the method includes extracting DNA from a group of plasma samples. Plasma can be obtained from any human bodily fluid, such as urine, saliva, peritoneal fluid, or cerebrospinal fluid. The extracted DNA may be cfDNA or genomic DNA.
[0026] Genomic DNA can be extracted from tissue and blood samples using the QiAamp DNA Mini Kit (Qiagen, catalog number: 563034) and QiAamp DNA Blood Kit (Qiagen, catalog number: 51104), respectively, and sheared to 450 bp (Covaris). In an exemplary experiment, a sequencing library was prepared on 1 μg of DNA using the TruSeq DNA PCR-Free Library Preparation Kit (Illumina), followed by one additional bead cleanup after end repair and adapter ligation. The extracted DNA was quantified using a Qubit 3.0 fluorometer, and length analysis was performed using an Agilent Bioanalyzer or High Sensitivity Fragment Analyzer. 2 × 150 bp paired-end sequencing was performed using Illumina HiSeq X or NovaSeq v1.0 instruments.
[0027] Cell-free DNA can be extracted from plasma using the Magbind cfDNA Extraction Kit (Omega Biotek, M3298). The manufacturer's recommended procedure was followed, but the elution volume was increased to 35 μL, and the elution time was increased to 20 minutes at 1,600 rpm (room temperature) on a thermomixer. The extracted cfDNA was quantified using a Qubit 3.0 fluorometer, and length analysis was performed using an Agilent Bioanalyzer or High Sensitivity Fragment Analyzer.
[0028] Step 104 of the method involves preparing a whole-genome library using a double-strand adapter. In some embodiments, the library is prepared by ligating a double-strand adapter having a trinucleotide pair Unique Molecule Identifier (UMI) and amplifying the amount of DNA to a sequencing-ready level using a first PCR.
[0029] The double-strand adapter contains a 3-base pair UMI, which allows the upper strand of DNA to be traced back to the lower strand of the original DNA molecule. In rare mutation environments such as early detection or minimal residual disease, the ctDNA content can drop to less than 1 in 10,000. Therefore, if 1 mL of plasma contains 1,000 to 10,000 GEs, there is at most about one circulating tumor read overlapping each somatic locus, which is lower than the error rate per base of high-throughput sequencers (approximately 1 error per 1,000 bases). To overcome this problem, deep-target sequencing approaches can use UMIs incorporated during library preparation for sequencing error correction. While strand-independent UMIs can help reduce sequencing errors, using UMIs that ligate forward and reverse DNA strands (i.e., double-strand sequencing) can reduce errors that occur on one strand (such as G>T conversion due to oxidative DNA damage) or errors that occur during library preparation.
[0030] The extracted cfDNA library can be prepared using a method similar to that used in Illumina whole-genome sequencing, but the full-length adapter has been replaced with a stubby Y adapter containing three UMI bases (IDT Duplex Seq adapter (1080799)).
[0031] The first PCR amplification generates a set of PCR duplicates that increase the amount of DNA, thereby enabling sequencing. These duplicates can be used to eliminate sequencing errors when two or more molecules with the same UMI are remapped.
[0032] In one exemplary experiment, six cycles of PCR were performed using indexed primers for input masses greater than 5 ng, and eight cycles for masses less than 5 ng. The library was quantified as described above. To improve the recovery rate of duplicates in human samples, an additional six cycles of PCR were performed on 4 ng of the prepared library before Ultima library conversion. No additional PCR cycles were performed on mouse PDX samples before Ultima library conversion.
[0033] Step 106 involves selecting a subset of the whole-genome library and amplifying this subset using a second PCR to increase the amount of PCR duplicates. Similar to the first PCR amplification, these duplicates can be used to eliminate sequencing errors when two or more molecules with the same UMI are remapped.
[0034] Step 108 may involve sequencing a series of reads from the amplified subset.
[0035] Illumina sequencing: The Illumina sequencing library was sequenced using 2x150 pair-end sequencing on HiSeq X or NovaSeq 1.0.
[0036] Ultima sequencing: The Illumina sequencing library was sent to Ultima Genomics (Newark, California) for library conversion and sequencing.
[0037] Step 110 involves aligning the reads to a host genome (e.g., the human genome) and denoising the double-stranded reads using the original sequencing reads and the reduced read information. FastQ reads were trimmed using cutadapt (version X) for adapter trimming and UMI trimming. The trimmed reads were then aligned to the human genome (version hg38) using bwa mem (parameters: -K 100000000 -p -v3 -t 16 -Y). The trimmed UMI was added to the alignment file as an additional RX tag. Single-strand and double-strand correction was performed using the fgbio toolset (version 2.0). For single-strand error correction, reads were grouped by UMI (fgbio GroupReadsByUmi -s edit) and a consensus call was performed (fgbio CallMolecularConsensusReads --min-reads 2). The resulting error-reduced fastQ was realigned to the human genome using bwa mem. For double-strand error correction, single-strand consensus sequencing was performed independently on the upper-strand and lower-strand mapping reads. Subsequently, the lower-strand mapping reads were converted to their inverse complementary sequences and merged with the upper-strand mapping reads. The reads were then grouped again by UMI, and error correction was performed to obtain double-strand corrected reads. Reads belonging to the double-strand family that were not corrected were processed to measure the following read-specific features. 1. Base pairing at variant position 2. Edit distance between the reference and the lead. 3. Total number of single nucleotide variants present on the reads 4. Read mapping quality 5. Position along the reads from both ends of the DNA fragment
[0038] Step 112 involves extracting variants from double-stranded reads. Variant types are defined by mutations such as a change of the C base to T (denoted as C>T) and the base pairs adjacent to the mutated base. For example, A[A>T]A is one type of variant. A[A>C]A is another type of variant, as are A[A>G]A and A[C>A]A. There are 96 variant types.
[0039] Reads are filtered based on the following criteria: 1) whether all reads in the double-stranded family have the same variant; 2) whether the edit distance is less than 2; 3) whether the total number of single-nucleotide variants on the read is less than 10; 4) whether the mapping quality is the highest possible value (60 for BWA MEM); and 5) whether the position along the read is 10 or greater.
[0040] Step 114 of this method involves comparing the extracted variants with a set of cancer-specific variant signatures. Certain cancers have very well-defined variant types, i.e., cancer-specific mutation signatures. For example, melanoma is a cancer that has a clear signature associated with UV exposure of the skin. Some lung cancers exhibit signatures associated with tobacco exposure. This signature matching process is shown in Figure 3E. The ratio measurement estimates shown along the y-axis are obtained using copy number-based tumor ratio estimates.
[0041] Copy number analysis was performed using ichorCNA (version). Tumor proportions were estimated after correcting for library and sequencing artifacts, using a normal group panel obtained from cancer-free controls (CTRL-01~CTRL-05) sequenced on the same instrument as the sample.
[0042] Under low tumor burden conditions, not all tumor-derived somatic mutations are expected to appear in the cell-free DNA sequencing pool; any genomic locus is expected to be covered by at most one circulating tumor DNA read. Therefore, to accurately quantify ctDNA content, a read-based (TF) estimation framework is needed, rather than a locus-based (TF) tumor ratio estimation. Genomically-wide mutations obtained from sequencing reads can be integrated and summarized as a weighted sum of single-nucleotide substitution (SBS) reference mutation signatures. As expected, a reanalysis of the publicly available PCAWG melanoma dataset revealed that UV-related SBS7 mutation signatures were the most abundant in the whole-genome sequencing (WGS) dataset of melanoma, and the clock-like signature 1A / 1B showed a weak correlation with patient age at tissue sampling (Spearman's ρ=0.26, p=0.0071, Figures 6A-6B). The trinucleotide sequence background of cfDNA variants was investigated through variant signature analysis at each level of noise reduction (UMI-independent, single-stranded, and double-stranded) to examine the possible origins of the mutations in these samples. The cosine similarity between UV-related SBS7 signatures and high-loading samples was highest after double-strand correction (mean cosine similarity to SBS7 between double-stranded and uncorrected samples was 0.967±0.02 and 0.365±0.03, p-value = 1.6×10⁻⁶). -4 (Wilcoxon rank-sum test) (Figures 7A-7B). Similar improvements were observed when measuring cosine similarity between clock-like signatures and cancer-free controls, highlighting the importance of double-strand correction in accurate signature analysis (Figures 7A-7B).
[0043] Based on the ability of novel mutation identification in error-corrected cfDNA WGS, which provides profiles corresponding to SBS7 signatures and clock-like signatures for identifying melanoma- and aging-related circulating DNA fragments respectively, a tumor-independent approach for ctDNA detection was developed based on mutation patterns. As a first step, the SBS mutation signatures were deconvolved from plasma-derived ctDNA mixtures using a non-negative maximum likelihood model, and the tumor ratio was estimated by normalizing the weights of the tumor-related SBS signatures by the total number of mutations and sequencing depth. Next, signature scores were calculated to determine whether the cancer-related SBS signatures could better explain the observed mutation profiles compared to random shuffling of cancer-related motifs. To analytically validate this approach, an in silico mixing experiment was performed. In this experiment, two high-loading ctDNA samples (MEL-12.A and MEL-12.B) of reads after double-strand noise removal and a cancer-free control (CTRL-06) were combined at various ratios (expected tumor ratio 0-1%) at 10-fold sequence depth (after double-strand consensus). As a result, when the expected tumor ratio was 10 -4 %, the estimated tumor ratio could be easily detected, and a signature score with high specificity for melanoma was obtained under the dilution condition of 10 -3 % (Figure 3F).
[0044] In one embodiment, a signature-based ctDNA detection platform was applied for preoperative ctDNA detection (i.e., tumor-independent ctDNA detection). Plasma samples obtained from 4 stage III melanoma patients, 3 cancer-free controls, and 1 stage IV melanoma patient who did not respond to treatment (collected at 5 different time points) were sequenced. Notably, the signature scores for ctDNA detection showed complete separation between samples from cancer-free controls and melanoma patients (Figures 3G-3H). In contrast, these two groups could not be distinguished by copy number-based analysis (Figure 8).
[0045] If the tumor burden in plasma is high (over 10%, source), tumor genotyping can be performed via sequencing of cfDNA and normal tissue. Mutect2 can be used for normal tissue. A quality threshold is set to retain only SNVs. Then, four blacklists (encode, gnomad, local blacklist, centromere) are applied to create the final tumor panel.
[0046] Read counts were performed using hmmcopy in 1 Mbp bins (duplicate reads and reads with mapping quality less than 60 were excluded). Read counts were adjusted using hmmcopy according to mapping potential and GC content. For both the Illumina and Ultima datasets, separate normal group panels were created using cancer-free controls (CTRL-01 to CTRL-05, n=5 per sequencing instrument). Tumor ratios were estimated using ichorCNA (version indicated). For illustrative purposes (Figure 1X, Supplementary Figure Y), corrected log2 read counts output by ichorCNA were used. Bins marked as copy number increase, amplification, or high-level amplification by ichorCNA were marked and colored as chromosome increase (pink). Bins marked as homozygous deletion and hemizygous deletion were marked and colored as chromosome loss (blue). Regions with neutral copy numbers were marked as neutral (black). Bins with corrected log2 read counts in the range of -0.05 to 0.05 were also marked as neutral (black).
[0047] Step 116 involves determining the cancer status based on the agreement between variant signatures.
[0048] The variants detected using the noise reduction method described above were used. Variants with allele frequencies exceeding 30% were presumed to be germline mutations and were excluded. The remaining reads were aggregated, and the frequency of occurrence of the variants in the trinucleotide sequence context was calculated. These trinucleotide variant frequencies were compared to the trinucleotide variant frequencies of publicly available references corresponding to different biological processes. In this context, given that the processed samples were obtained from cancer-free controls and melanoma patients, it was assumed that the sample trinucleotide variant frequencies would be a combination of age-related trinucleotide variant frequencies and UV damage-related trinucleotide frequencies. The sample frequencies were fitted to the reference using the non-negative maximum likelihood method. To eliminate false positives, a permutation test was performed. In this test, the trinucleotide frequencies in the UV-related (melanoma) reference signature were randomly varied, and non-negative maximum likelihood fitting was performed. If a sample showed a stronger fit to the randomly varied frequency than the original, it was considered a false positive. This process was repeated 10,000 times to obtain signature scores. Samples with a signature score of less than 0.001 were considered acceptable. Samples exceeding this threshold were considered cancer-negative.
[0049] In some embodiments, the method may include comprehensive WGS. In addition to WGS, embodiments of the disclosure may also use less comprehensive sequencing methods, such as whole exome sequencing (WES) or SNP genotyping. Various enrichment techniques may be applied, including, but not limited to, exome enrichment, targeted gene enrichment, and / or specific mutation enrichment. For example, exome enrichment may include whole exome sequencing. Targeted gene enrichment may include sequencing of a whole gene, including one or more introns, exons, and / or coding sequences. Specific mutation enrichment may target specific locations in the genome, including, for example, exons, introns, and / or other user-defined locations. Techniques for achieving enrichment of at least one of the aforementioned regions are generally hybridization-based or primer-based. Examples of these include targeted variant sequencing, targeted gene sequencing, whole exome sequencing, targeted PCR, nested PCR, and / or linear PCR.
[0050] When input is limited, such as within cell-free DNA, it is possible to comprehensively sequence the sample (i.e., sequence all available molecules). Therefore, some embodiments of this disclosure may include comprehensive whole-exome sequencing, comprehensive cfDNA sequencing, and / or comprehensive target sequencing.
[0051] Figures 2A to 2F illustrate ultra-trace ctDNA detection requiring high-depth sequencing coverage and low error rates.
[0052] Figure 2A shows a group of graphs illustrating the simulated sequencing coverage. Simulation analysis shows that when the tumor ratio is 10 -5 The following situations indicate that lower error rates and higher sequencing coverage are necessary for accurate ctDNA detection.
[0053] Simulations for Figure 2A were performed assuming a tumor mutation compendium consisting of 10,000 SNVs at different error rates (10⁻³, 10⁻⁴, and 10⁻⁵), coverage (1, 10, and 100), and tumor ratios (0, 10⁶, and 10⁵). For each of the 50,000 SNV mutations, the coverage was simulated using a Poisson distribution. Each simulated base pair was classified as either ctDNA or cfDNA according to the tumor ratio, and errors misclassified as ctDNA were determined based on the error rate. The estimated tumor ratio was calculated by summing the number of ctDNA molecules and the number of errors and dividing by the total number of simulated base pairs.
[0054] Figure 2B shows the pretreatment workflow for cfDNA library preparation. This workflow is similar to the embodiment shown in Figure 1 and includes obtaining plasma and extracting cfDNA therefrom. The double-stranded DNA library is prepared for sequencing, which is done using Illumina sequencing or Ultima library conversion and subsequent Ultima sequencing.
[0055] Figure 2C is a graph comparing the sequencing depth (genomic equivalent) of the corresponding Illumina and Ultima datasets across 15 matching cfDNA samples.
[0056] Figure 2D shows a comparison of normalized read coverage for matching cfDNA samples (chromosome level) from Illumina (top panel) and Ultima (bottom panel).
[0057] Figure 2E compares tumor ratios based on copy number variation (CNV) and single nucleotide variant (SNV) using the Illumina and Ultima datasets. The graph on the left shows the estimated tumor ratios based on CNV, measured using Illumina or Ultima sequencing, for the corresponding samples using ichorCNA. Prior to tumor ratio estimation, a corresponding cancer-free control was used to create a normal group panel.
[0058] The graph on the right side of Figure 2E shows tumor ratio estimates based on single nucleotide variants, measured by Illumina or Ultima sequencing. Somatic SNVs were identified through matched tumor-normal sequencing. Two samples with low ctDNA ratios (e.g., measured via CNV analysis and less than 5%) were excluded without tumor sequencing.
[0059] Figure 2F shows the predicted tumor ratio scores with and without error suppression. It illustrates the effect of applying Ultima-specific quality filtering to perform analytical noise reduction based on tumor information (red) versus not applying it (blue) in an in silico mixed study (50 replicas per tumor ratio, 80x coverage per replica) of metastatic melanoma sample MEL-01 and cancer-free control CTRL-05.
[0060] Figures 3A to 3H show double-strand correction that enables the detection of ctDNA without tumor sequencing. Figure 3A shows the error rates in mouse and human DNA in double-strand sequencing, single-strand sequencing, and the uncorrected group. The graph on the left shows the error rate for double-strand WGS sequencing in mouse PDX samples (n=3). The open circles in the graph on the left indicate samples in which no sequencing errors were detected. The graph on the right shows the cross-referencing results of double-strand WGS sequencing in the patient sample MEL-12.D with tumor mutation profiles from 107 melanoma patients obtained from the PCAWG (Patient-Cross-Cancer-Wide Genome Analysis) consortium. Base substitutions matching somatic mutations in tumors were considered errors (after removing germline and somatic mutations from matching patient data).
[0061] To initially test the accuracy of double-strand error correction, double-strand libraries were prepared using cfDNA obtained from the plasma of mice carrying patient-derived xenografts (n=4, NOD / ShiLtJ strain, n=1 lung cancer, n=3 diffuse large B-cell lymphoma). Tumor ratios were defined as the percentage of reads that uniquely map to the human genome, and were 0.4%, 40%, 73%, and 96%. To estimate the error rate of the double-strand libraries, the level of mutations at well-characterized homozygous variant sites in NOD / ShiLtJ mice was investigated for three samples with an increased number of reads mapped to mice. Overall, 4.2 × 10⁶ 6 Of the total number of base pairs, only two base pairs were sequenced that did not match the known genotype of the mouse, resulting in an error rate of 4.75 × 10⁻⁶. -7 This was the case (Figure 3A). These results are in line with previous reports using whole-genome double-strand sequencing (Abascal et al., 2021, used a similar protocol with an error rate of 2 × 10⁻⁶). -7 This is consistent with the report on the matter.
[0062] Figure 3B is a series of graphs comparing variant allele frequencies calculated using unfiltered sequencing reads. Variant allele frequencies are shown at the locations where variants were detected using uncorrected reads (left column) and at the locations where variants were detected using double-stranded corrected reads (right column). The top and bottom rows represent representative examples from cancer-free patient samples and high-burden patient samples, respectively.
[0063] Figure 3C is a graph comparing the model allele frequencies at the double-strand corrected position (only allele frequencies less than 30%) with the estimated tumor ratio based on copy number in patients with progressive disease (samples MEL-12.A~E).
[0064] Figures 3D and 3E illustrate exemplary methods of signature matching between sequencing reads. The Signature 7 reference in Figure 3D is a publicly available signature associated with UV exposure (i.e., a melanoma-specific signature). The signatures in Figure 4E show uncorrected, single-strand corrected, and double-strand corrected signatures for melanoma patients with a control signature and a MEL-12.D signature. Each bar on the signature represents a specific trinucleotide mutation (i.e., 96 bars), and the y-axis shows the relative proportion of the trinucleotide mutation. The double-strand corrected MEL-12.D signature is compared to the reference signature, and the cancer-like signature is only evident in cancer patients after double-strand correction.
[0065] Figure 3F is a graph showing signature scores and ctDNA detection in an in silico mixed study of metastatic melanoma samples MEL-12.A / B and cancer-free control CTRL-06 (10 replicas per tumor ratio, 10x coverage per replica).
[0066] In the upper panel, when decomposing the trinucleotide frequencies of the sample into a reference signature, signature scores are used to estimate the contribution to signature SBS& (UV-related in melanoma).
[0067] The lower panel shows ctDNA detection based on predicted tumor ratios. Using Z-score estimation, detection of the variant signature SBS7 was calculated compared to detection in replicates with TF=0. True variants originating from the high-load sample MEL-12.A / B or the cancer-free sample CTRL-06 are shown in blue (filled circles: MEL-12.A / B; open circles: CTRL-06). Error bars represent the standard deviation of the number of variants per replicate at a given predicted tumor ratio.
[0068] Figure 3G is a graph showing a series of signature scores for melanoma signature 7 in plasma cfDNA samples using double-stranded WGS (n=9 melanoma samples, n=3 controls). Red samples are taken at different time points in the clinical course of patient MEL-12 with stage IV melanoma. Blue samples are from separate patients (MEL-08 to MEL-11). Pink samples represent control samples.
[0069] Figure 3H is a graph showing the estimated tumor ratio in samples with elevated signature scores. The X-axis represents the clinical time point for each patient sample. The tumor ratio was estimated by multiplying the number of single nucleotide variants found in double-strand corrected reads by the weight of Signature 7 after reference signature decomposition, and then normalizing by coverage depth. A tumor mutation profile consisting of 10,000 SNVs was assumed for the estimation of the tumor ratio.
[0070] Figure 4 shows cfDNA fragment lengths in single-ended sequencing datasets correlated with paired-ended sequencing. For cfDNA molecules shorter than 200 base pairs, single-ended Ultima reads accurately reproduce fragment lengths compared to paired-ended Illumina sequencing.
[0071] Figures 5A to 5C illustrate the effect of artifact blacklisting in single-nucleotide variant detection.
[0072] Figure 5A is a graph showing the UG-specific blacklist. The UG-specific blacklist includes low GC content regions, tandem repeats, regions with poor mapping capabilities, regions with large coverage variations, and homopolymer regions exceeding 10 base pairs.
[0073] Figure 5B shows the overlap between the UG blacklist and other low-confidence regions. These other low-confidence regions include centromeres, simple repeats, regions encoding blacklists, and gnomad regions with AF values greater than 0.001.
[0074] Figure 5C is a graph showing the overlap between SNVs and low confidence regions in melanoma tumor tissue. The effect of blacklisting on somatic single nucleotide variant (SNV) retrieval is shown in 107 melanoma tissue samples obtained from the PCAWG consortium.
[0075] Figures 6A and 6B show the cosine similarity between the clock-like signature and the UV-related signature, SBS1B and SBS7, in the high-load and cancer-free samples, respectively.
[0076] Figure 6A is a heatmap of cosine similarity in cancer-free samples and high-load ctDNA samples. It shows the cosine similarity of double-strand corrected reads, single-strand corrected reads, and uncorrected reads from cancer-free samples (n=3) or high-load ctDNA samples (n=5, both derived from stage IVB melanoma patient MEL-12). The X-axis is arranged alphabetically without hierarchical clustering.
[0077] Figure 6B is a box plot showing the cosine similarity of three correction methods for the same cancer-free and high-load ctDNA samples using the techniques disclosed herein.
[0078] Figures 7A and 7B show the reanalysis of 107 melanoma mutation signatures by the PCAWG (Plant-Comprehensive Cancer Analysis Working Group) consortium.
[0079] Figure 7A is a graph showing the signature ratios of several variant signatures. The decomposition of double-strand corrected mutations into representative variant signatures was analyzed using a non-negative maximum likelihood model. The box plots are sorted in ascending order of median for each signature.
[0080] Figure 7B shows a correlation plot between age at cancer diagnosis and the number of clock-like mutations attributable to SBS1A and SBS1B. The number of mutations was obtained by multiplying the total number of mutations found after double-strand correction by the weights of SBS1A and SBS1B.
[0081] Figure 8 shows tumor ratio estimates based on tumor-nonspecific copy number in a cancer-free control sample (n=3) and preoperative melanoma plasma (n=4).
[0082] In another embodiment, whole-genome sequencing can be performed without using double-stranded sequencing, and tumor ratio estimation based on SNVs can be achieved.
[0083] Tumor proportion estimation based on SNVs was performed by counting reads of cell-free DNA containing corresponding tumor-specific somatic mutations (mutation calling pipeline is described later). Platform-specific blacklists were created to limit the impact of genomic problem regions. In Illumina sequencing, regions corresponding to the ENCODE blacklist (source), centromeres (source), simple repeat regions (source), and locations with high mutation rates (GNOMAD, AF>0.001, source) were not considered. In Ultima sequencing, Ultima-specific low-confidence regions consisting of homopolymers, AT-rich regions, tandem repeats, and regions with low mapping and large coverage variability were also excluded.
[0084] To limit the impact of sequencing errors, a custom script was used to perform platform-specific noise reduction. Illumina alignment files were filtered to include read pairs overlapping somatic mutation locations. Paired-end reads were filtered using X, Y, and Z criteria and retained only if both R1 and R2 had somatic mutations or reference base pairs. Tumor ratios were estimated by dividing the number of filtered reads containing somatic mutations by the total number of filtered reads.
[0085] The Ultima alignment file was subsetted to include reads that overlapped with somatic mutation locations. Reads were filtered using the X, Y, and Z criteria. Tumor proportions were estimated by dividing the number of filtered reads containing somatic mutations by the total number of filtered reads.
[0086] Training set and feature space of the SNV model The training set was obtained from plasma samples enriched with ctDNASNV fragments (true labels) derived from specific melanoma tumors and cfDNASNV reads (false labels) obtained from healthy controls without known cancers as listed in Appendix xx. Candidate reads were extracted from custom denoised alignment files. For the true-label set, patients with high-burden metastatic disease were used, and only reads representing matched tumor variants were retained.
[0087] Using a custom deep learning model, we improved the signal-to-noise ratio in a manner similar to previous research (Widman et al, 2022) and effectively classified candidate SNV reads. Candidate SNV reads were extracted using pysam (v0.15.2). Furthermore, strong region-specific and sequencing technique-specific features were encoded as input to the deep learning model architecture using a custom Python (v3.6.8) script. Two separate input structures corresponding to each component of the ensemble model are described below.
[0088] For MLPs, a tabular set of feature values is provided as input.
[0089] Feature selection for this purpose was performed on filtered SNV reads under both true and false labeling. Specific features and their corresponding single-variable AUC performance are listed in Table xx. As highlighted in previous studies (Widman et al, 2022), tissue-specific transcriptional features help define the likelihood of somatic mutations being observed in genomic regions. Local tumor mutation density was classified by quantifying WGS SNV mutation calls from the PCAWG database (edge ref 81), and the total number of SNV mutations was counted from available melanoma-derived tumor samples. In addition, local histone CHiP-Seq marks and tissue-specific bulk RNA expression values were reported as standardized RPKM values from primary tissue alignments in ENCODE (edge ref 95). Region-specific DNAse sensitivity peaks (transitioned to GRCh38) were also included, obtained from narrowpeak files reported in ENCODE (edge ref 95, 96). Melanoma-specific ATAC peak calls reported in TCGA (edge ref 82) were also included.
[0090] Since the deep learning model is designed to operate on a read-level compendium, the values for the features defined above were calculated using a sliding window around each candidate read. The optimal length for this sliding window has been defined in previous research (Widman et al, 2022). Furthermore, region-specific chromatin annotation tracks (migrated to ChromHMM-GRCh38) (edge ref 83) were obtained from ENCODE. Hi-C compartment information was extracted using the Hi-C SNIPER (edge ref 97) bed file. Finally, region-specific features for replication timing and mean expression values (migrated to GRCh38) were obtained from existing literature (edge ref 37).
[0091] Furthermore, Ultima-specific read-level features were included. These include the following: • X-FC1 - Number of features (SNPs) on the same read • X-FC2 - Number of features (SNPs) that passed through the filter on the same read (same number of base pairs as the reference). • Propagation from X-FLAGS-bam file flags ·ed_1 - The edit (Levenshtein) distance from the reference to the read before the target SNV. • ed_2 - The edit (Levenshtein) distance from the reference to the read after the target SNV.
[0092] Next, for the CNN, we used a one-hot encoded tensor structure of candidate reads similar to the previous study Widman et al, 2022. Each read was encoded with a variant (sequencing artifacts / noise from a healthy control or somatic mutations from a heavily burdened tumor plasma sample). The encoded tensor has an image-like structure with a 12 × 240 shape. Rows correspond to the one-hot encoded nucleotides (N, A, C, T, G) corresponding to the reference and read. The second-to-last row dimension is used to mark the position along the read highlighting the target SNV. Finally, the presence / absence (0 / 1) of cycle skips (defined by Ultima) is encoded along the last row dimension to add further relevance to the trinucleotide context of the target SNV. Columns correspond to the nucleotides along the length of the read, respectively. The maximum read length is 200, but an extra 40 base pairs are padded with the reference genome to add additional relevant contextual information.
[0093] SNV model design and training Deep learning models have an ensemble structure, consisting of two main components: a region-specific / read-specific multilayer perceptron (MLP) and a sequence-based convolutional neural network (CNN), whose weight matrices are learned collaboratively.
[0094] The MLP, which takes a feature matrix as input, is constructed by stacking four high-density blocks in series. Each block is defined as consisting of a fully connected layer with ReLU activation. Furthermore, for the purpose of regularization, the inputs to each fully connected layer are batch normalized, and the outputs are passed through a dropout layer.
[0095] The CNN consists of four one-dimensional convolutional layers with nonlinear ReLU activation that sequentially extract information at different spatial resolutions. Furthermore, similar to classical deep learning frameworks, each convolutional layer (after nonlinear activation) is followed by a maximum pooling layer. The output is then passed through the three high-density blocks defined above.
[0096] Subsequently, the latent outputs of both the MLP and CNN are concatenated and passed through a single high-density block. Finally, a single sigmoid-activated fully connected layer is used to obtain a probability score between 0 (sequencing noise) and 1 (true somatic mutation). This probability score reflects the model's estimate of whether the candidate SNV mutations present in the encoded reads are more likely to be signal-derived or noise-derived. The ensemble model is built within Keras (v2.3.0) with Tensorflow (1.14.0).
[0097] To train the ensemble model, we minimize the objective function defined as the binomial cross-entropy loss. Performance metrics are reported within a balanced set.
[0098] UMI correction improves the accuracy of insertion-deletion mutation (indel) detection in the Ultima sequencing dataset. UMI assigns a unique barcode to each DNA molecule. During PCR, this barcode tag (and DNA molecule) is replicated multiple times. This allows for the identification of PCR duplications using UMI. Since the likelihood of the same error occurring in two PCR duplications is low, identified duplications can be used to correct sequencing errors. Note that Ultima's flow sequencing is prone to homopolymer size errors, which are interpreted as false indels. Figure 9A is a graph showing the homopolymer size between two PCR duplications, while Figure 9B is a graph showing the homopolymer size between a read and an aligned reference. For reference, Figure 9C is a graph showing the frequency of homopolymer sizes across the entire human genome using the techniques disclosed herein. Furthermore, to illustrate the accuracy improvement through UMI correction, Figure 9D is a graph showing the insertion-deletion mutation calling accuracy by PCR duplication family size using the techniques disclosed herein.
[0099] UMI-ligated reads enable error-resistant detection of trinucleotide motifs in the Ultima sequencing dataset. UMI can also be used to discover error-resistant single-nucleotide variants. These variants generally fall into the category of cycle-shift motifs, which are specific trinucleotides identified by Ultima as error-resistant. Details regarding cycle shifts are shown in Figure 10, a graph illustrating the single-nucleotide variant analysis of the corresponding Ultima and Illumina sequencing datasets. Figures 11A and 11B demonstrate that flow sequencing provides predictable error-resistant motifs. Figure 11A is a flowchart of the sequencing process that provides predictable error-resistant motifs, and Figure 11B is a graph showing the error rates for each sequencing platform using the techniques disclosed herein.
[0100] Development of a novel machine learning classifier for wet lab improvements and double-strand variant detection. In some embodiments of this disclosure, molecular and computational improvements were applied to the methods to improve the yield of double-stranded molecules. Molecular improvements are demonstrated by more efficient bottleneck processing of DNA molecules (Figures 12A and 12B). Figure 12A is a graph of double-stranded WGS libraries sequenced from three different start inputs with coverage from 1 to 13x. As shown in the graph, double-strand correction was applied and the double-stranded recovery yield (depth of double-stranded-only coverage relative to total sequencing coverage) was measured. Figure 12B shows the overlap rate for n=3 for the samples in Figure 12A.
[0101] Figure 12C shows that downsampling experiments demonstrate that the improved bottleneck processing achieves faster and higher double-strand coverage than other embodiments. This is further illustrated by Figure 12D, a graph showing significantly higher double-strand coverage at fixed coverage. Figure 12E is a graph comparing the number of double-strand mutations discovered using the double-strand method (fgbio) and the decision tree. Figure 13 further elaborates on this by showing the mutation error rate (N=3 per condition) in mouse PDX samples. This data demonstrates the extended applicability of the process in the embodiments of this disclosure, including application to baseline stage III melanoma.
[0102] In some embodiments of this disclosure, more samples are generated, including the preoperative samples shown in Figure 14. Plasma-free DNA was obtained from bladder cancer patients who may or may not have received chemotherapy. Figure 15 shows embodiments of this disclosure detecting chemotherapy-related mutation signatures in most samples that may have received chemotherapy, particularly in application to bladder cancer and the detection of the "APOBEC" signature and chemotherapy. No chemotherapy-derived signals are measured in samples that have never received chemotherapy (green) or in cancer-free control samples (blue).
[0103] Bladder cancer typically exhibits an APOBEC mutation signature. This signature can also be detected in free plasma DNA, as shown in Figures 16A and 16B. Dark bars indicate the APOBEC signature in the tumor, while light bars indicate APOBEC measurements in cfDNA.
[0104] Figure 17 is a schematic diagram of an example of a computing node. Computing node 10 is merely an example of a suitable computing node and is not intended to limit the scope or functionality of the embodiments described herein. Nevertheless, computing node 10 is capable of implementing and / or performing any of the functions described above herein.
[0105] Computing node 10 contains a computer system / server 12, which can operate in conjunction with a number of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, portable devices or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable home electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed computing environments that include any of the above systems or devices.
[0106] The computer system / server 12 can be described in the general context of computer system executable instructions, such as program modules, that are executed by the computer system. Generally, a program module can include routines, programs, objects, components, logic, data structures, etc., that perform a specific task or implement a specific abstract data type. The computer system / server 12 can be implemented in a distributed computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed computing environment, program modules can be located on local and remote computer system storage media, including memory storage devices.
[0107] As shown in Figure 17, the computer system / server 12 within the computing node 10 is represented in the form of a general-purpose computing device. The components of the computer system / server 12 include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components, including the system memory 28, to the processor 16.
[0108] Bus 18 represents one or more types of bus structures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor buses or local buses using various bus architectures. Examples of these architectures include, but are not limited to, Industry Standard Architecture (ISA) buses, Microchannel Architecture (MCA) buses, Extended ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, Peripheral Component Interconnect (PCI) buses, Peripheral Component Interconnect Express (PCIe), and High-Function Microcontroller Bus Architecture (AMBA).
[0109] The computer system / server 12 typically includes various computer system-readable media. These media can be any available media accessible by the computer system / server 12 and include both volatile and non-volatile media, as well as removable and non-removable media.
[0110] The system memory 28 may include computer system-readable media in the form of volatile memory such as random access memory (RAM) 30 and / or cache memory 32. The algorithm computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. For example only, a storage system 34 may be provided for reading from and writing to a non-removable non-volatile magnetic medium (not shown, commonly referred to as a “hard drive”). Not shown, a magnetic disk drive may be provided for reading from and writing to a removable non-volatile magnetic disk (e.g., a “floppy disk”), and an optical disk drive may be provided for reading from and writing to a removable non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical medium. In these cases, each may be connected to the bus 18 by one or more data medium interfaces. As further shown and described below, the memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.
[0111] The program / utility 40 has a set (at least one) of program modules 42, which may be stored in memory 28, for example, together with an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a network environment. The program module 42 generally performs the functions and / or methods of the embodiments described herein.
[0112] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, and display device 24; one or more devices that enable a user to interact with the computer system / server 12; and / or any device (e.g., a network card, modem, etc.) that enables the computer system / server 12 to communicate with one or more other computing devices. Such communication can be performed via the input / output (I / O) interface 22. Furthermore, the computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via the network adapter 20. As shown in the figure, the network adapter 20 communicates with other components of the computer system / server 12 via the bus 18. It should be understood that other hardware and / or software components, not shown in the figure, may be used in conjunction with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0113] In various embodiments, a learning system is provided. In some embodiments, a feature vector is provided to the learning system. Based on the input features, the learning system generates one or more outputs. In some embodiments, the output of the learning system is a feature vector. In some embodiments, the learning system includes an SVM. In other embodiments, the learning system includes an artificial neural network. In some embodiments, the learning system is pre-trained using training data. In some embodiments, the training data is retrospective data. In some embodiments, the retrospective data is stored in a data store. In some embodiments, the learning system may be further trained by manually curating previously generated outputs.
[0114] In some embodiments, the learning system is a trained classifier. In some embodiments, the trained classifier is a random decision tree forest. However, it should be understood from this disclosure that a variety of other classifiers are suitable for use, including linear classifiers, support vector machines (SVMs), or neural networks such as recurrent neural networks (RNNs).
[0115] Appropriate artificial neural networks include, but are not limited to, feedforward neural networks, radial basis function networks, self-organizing maps, learning vector quantization, recurrent neural networks, Hopfield networks, Boltzmann machines, echo-state networks, long short-term memory, bidirectional recurrent neural networks, hierarchical recurrent neural networks, probabilistic neural networks, modular neural networks, associative neural networks, deep neural networks, deep belief networks, convolutional neural networks, convolutional deep belief networks, mass memory and retrieval neural networks, deep Boltzmann machines, deep stacking networks, tensor deep stacking networks, restricted Boltzmann machines with spikes and slabs, complex hierarchical deep models, deep coding networks, multilayer kernel machines, or deep Q-networks.
[0116] This disclosure may be embodied as a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium(s) having computer-readable program instructions therein for performing an aspect of this disclosure.
[0117] A computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. A computer-readable storage medium may be, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or a suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved raised structures on which instructions are recorded, and suitable combinations thereof. As used herein, computer-readable storage media should not be interpreted as transition signals in isolation, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., optical pulses through optical fiber cables), or electrical signals transmitted through wires.
[0118] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network, to an external computer or external storage device. The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer-readable program instructions from the network and transfers those computer-readable program instructions to be stored in a computer-readable storage medium within the respective computing / processing device.
[0119] The computer-readable program instructions for performing the processing described herein may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk and C++, and conventional procedural programming languages such as the C programming language or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on both the user's computer and a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be made to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute a computer-readable program instruction to carry out an aspect of the present disclosure by using state information of the computer-readable program instruction to personalize the electronic circuit.
[0120] Aspects of this disclosure are described herein by reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products by embodiments of this disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and / or combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0121] These computer-readable program instructions may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine, which in turn may be executed via the processor of the computer or other programmable data processing device to form means for implementing functions / operations specified within blocks(or blocks) of a flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium that can instruct a computer, a programmable data processing device, and / or other device to function in a particular manner, in which case the computer-readable storage medium storing the instructions may include a product containing instructions that implement modes of functions / operations specified within blocks(or blocks) of a flowchart and / or block diagram.
[0122] Computer-readable program instructions may also be loaded into a computer, other programmable data processing device, or other device, in which case the instructions executed on the computer, other programmable device, or other device implement the function / operation specified in the block(s) of the flowchart and / or block diagram by performing a series of operational steps on the computer, other programmable device, or other device to generate a computer implementation process.
[0123] The flowcharts and block diagrams shown in the figures illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or instruction containing one or more executable instructions for implementing a specified logical function(s). In some alternative implementations, the functions described within a block may be executed in an order different from that shown in the figure. For example, two consecutively shown blocks may actually be executed almost simultaneously, or the blocks may sometimes be executed in reverse order depending on the functions involved. It should also be understood that each block in the block diagram and / or flowchart diagram, and combinations of blocks in the block diagram and / or flowchart diagram, may be implemented by a special-purpose hardware-based system that performs a specific function or operation, or a combination of special-purpose hardware and computer instructions.
Claims
1. It is a method, Extracting DNA from biological samples, The present invention relates to the preparation of a sequence library with double-stranded adapters, wherein the sequence library is prepared by ligating double-stranded adapters having unique molecular identifiers (UMIs) to each end of multiple strands of extracted DNA, and amplifying the extracted DNA by a first PCR. Selecting a subset of the aforementioned array library, In order to increase the number of PCR duplicates, the subset is amplified by a second PCR, Sequence multiple double-stranded reads from the amplified subset, Aligning the plurality of double-stranded reads to the host genome, and denoising the plurality of double-stranded reads based on the alignment, To detect the presence of a variant in at least one of the plurality of double-stranded reads, Determining the signature of the aforementioned variant, The signature of the variant is compared with the disease-specific variant signature group, Based on the above comparison, the type of disease is determined, Methods that include...
2. The method according to claim 1, wherein the UMI has exactly three base pairs.
3. The method according to claim 1, wherein the UMI is less than 5 base pairs.
4. The method according to claim 1, wherein the type of disease is a type of cancer.
5. The method according to claim 1, wherein the type of cancer includes bladder cancer.
6. The method according to claim 1, wherein the sequence library includes one of a whole genome library or a whole exome library.
7. The method according to claim 6, further comprising preparing a sequence library with double-stranded adapters to reduce errors on one or more strands of extracted DNA.
8. The method according to claim 1, wherein the extracted DNA is amplified by the first PCR, which includes eliminating sequencing errors based on the presence of two or more molecules having the same UMI.
9. The method according to claim 1, wherein sequencing of the plurality of double-stranded reads includes sequencing on a paired-end system.
10. The method according to claim 1, Sequencing the aforementioned plurality of double-stranded reads includes sequencing on a single-ended system. Aligning the aforementioned plurality of double-stranded reads to the host genome is Obtaining multiple single-ended DNA sequencing reads, Separating the upper mapping strand of the DNA from the lower mapping strand, Error reduction processing is performed on each of the upper mapping chain and the lower mapping chain. The lower mapping chain is returned to the upper mapping chain by regrouping it based on UMI, Error correction is performed between the upper and lower chains, including, method.
11. The method according to claim 1, Sequencing the aforementioned plurality of double-stranded reads includes sequencing on a single-ended system. Aligning the sequence with the host genome is Obtaining multiple single-ended DNA sequencing reads, To generate synthetic paired-end reads, To perform error correction on all chains, including, method.
12. The method according to claim 1, wherein sequencing the plurality of double-stranded reads includes sequencing a series of uncorrected reads belonging to a double-stranded family.
13. The method according to claim 12, further comprising processing uncorrected reads belonging to a double-stranded family and measuring read-specific characteristics.
14. The method according to claim 13, further comprising filtering out the uncorrected reads based on the measured read-specific characteristics.
15. The method according to claim 1, further comprising trimming the sequence of reads from the amplified subset.
16. The method according to claim 1, wherein the extracted variant is compared with a group of cancer-specific variant signatures, To calculate the tumor ratio estimate for double-strand corrected signatures, The estimated tumor ratios are plotted to generate double-strand corrected signatures for the extracted variants, The process involves comparing the double-stranded corrected signature with the reference signature, including, method.
17. The method according to claim 16, wherein the signature includes the relative ratio of trinucleotide mutations.
18. The method according to claim 4, Correcting library and sequencing artifacts using a cancer-free control panel sequenced on the same system, To estimate the tumor ratio, Methods that further include the above.
19. The method according to claim 1, further comprising integrating genome-wide mutations from the sequencing reads as a weighted sum of single-nucleotide substitution (SBS) reference mutation signatures.
20. A method according to claim 19, wherein the mutations of the entire genome are integrated, Using a non-negative maximum likelihood model, we deconvolved the SBS mutation signature from a plasma DNA mixture, The tumor proportion is estimated by obtaining the weights of tumor-associated SBS signatures and normalizing them by the total number of mutations and the depth of sequencing. To calculate a signature score to determine whether the aforementioned cancer-related SBS explains the observed mutation profile, A method that includes this.
21. The method according to claim 1, further comprising excluding variants with an allele frequency of more than 30%.
22. The method according to claim 21, To aggregate reads that have variants with an allele frequency of less than 30%, To calculate the frequency of variants in the trinucleotide sequence context of the aggregated reads, The calculated trinucleotide variant frequencies are compared with several reference frequencies for different biological processes. Methods that further include the above.
23. The method according to claim 1, wherein the disease state is determined Randomly changing the trinucleotide frequencies of the reference signature from the aforementioned group, Perform a non-negative maximum likelihood fit between the randomly rearranged trinucleotide frequencies and the signature frequencies, Evaluating the aforementioned matching results that are below the threshold for disease negativity, Methods that include...
24. The method according to claim 1, wherein the DNA is genomic DNA.
25. The method according to claim 1, wherein the DNA is cell-free DNA (cfDNA).
26. The method according to claim 1, wherein the presence of the variant is detected The aforementioned multiple double-stranded reads are provided to a pre-trained machine learning model, From there, we can obtain indications of base variants regardless of the comparison sequence length, Methods that include...
27. The method according to claim 26, wherein the pre-trained machine learning model includes an artificial neural network.
28. The method according to claim 27, wherein the artificial neural network is one of the following: a feedforward neural network, a radial basis function network, a self-organizing map, a learning vector quantization, a recurrent neural network, a Hopfield network, a Boltzmann machine, an echo state network, long short-term memory, a bidirectional recurrent neural network, a hierarchical recurrent neural network, a probabilistic neural network, a modular neural network, an associative neural network, a deep neural network, a deep belief network, a convolutional neural network, a convolutional deep belief network, a large-capacity memory and retrieval neural network, a deep Boltzmann machine, a deep stacking network, a tensor deep stacking network, a restricted Boltzmann machine with spikes and slabs, a complex hierarchical deep model, a deep coding network, a multilayer kernel machine, or a deep Q network.
29. The method according to claim 26, wherein the pre-trained machine learning model includes a trained classifier.
30. The method according to claim 29, wherein the trained classifier is a random decision tree forest.
31. The method according to claim 1, wherein the sample group includes a plasma sample group.
32. The method according to claim 1, wherein the sequence library includes a whole genome sequence library.
33. A computer program product for reducing the sequence determination error rate, wherein the computer program product comprises a computer-readable storage medium, the computer-readable storage medium includes program instructions embodied therewith, the program instructions are executable by a processor, and the computer program product causes the processor to perform the method according to any one of claims 1 to 32.