Variant Capture Minimal Residual Lesion Panel

Polynucleotide libraries with variant sequences address the limitations of existing genomic variant detection methods by enabling rapid and accurate MRD detection through personalized NGS assays.

JP2026515788APending Publication Date: 2026-05-19TWIST BIOSCIENCE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TWIST BIOSCIENCE CORP
Filing Date
2024-04-12
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for identifying genomic variants in complex nucleic acid samples face limitations in scalability, automation, speed, sensitivity, and cost, particularly in detecting minimal residual disease (MRD) which requires highly sensitive and personalized assays.

Method used

Development of polynucleotide libraries containing variant sequences associated with MRD, designed to be synthesized with specific distribution and frequency, and used in personalized NGS assays for accurate MRD detection.

Benefits of technology

Enables rapid and accurate detection of MRD with high sensitivity and specificity, allowing for personalized panels to be designed and manufactured in as little as six days.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515788000001_ABST
    Figure 2026515788000001_ABST
Patent Text Reader

Abstract

This specification describes a composition for detecting genovariants associated with minimal residual disease (MRD). The composition comprises a library containing multiple polynucleotides, each containing at least one variant associated with minimal residual disease (MRD). This specification further describes a method for preparing such a polynucleotide library and a method for detecting MRD in a sample using such a polynucleotide library.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 495,938, filed April 13, 2023, which is incorporated herein by reference in its entirety. All publications, patents, and patent applications referenced herein are incorporated herein by reference whenever they are specifically and individually referred to as if each individual publication, patent, or patent application were incorporated by reference. [Background technology]

[0002] High-fidelity and low-cost identification of genomic variants plays a central role in biotechnology, medicine, and basic biological research. While various methods are known for identifying genomic variants in complex nucleic acid samples, these techniques often suffer from limitations in scalability, automation, speed, sensitivity, accuracy, and cost. [Overview of the project]

[0003] This specification provides polynucleotide libraries. In some embodiments, the polynucleotide library comprises a plurality of polynucleotides. In some embodiments, each polynucleotide of the plurality of polynucleotides comprises a nucleic acid sequence having a center. In some embodiments, the nucleic acid sequence of each polynucleotide comprises at least one variant sequence associated with minimal residual disease (MRD).

[0004] In some embodiments, the position of at least one MRD-associated variant sequence is within 20 bases from the center of each nucleic acid sequence of each polynucleotide. In some embodiments, the polynucleotide library further includes a distribution of the positions of at least one MRD-associated variant sequence in the nucleic acid sequences of all polynucleotides of a plurality of polynucleotides, and this distribution includes the average within 20 bases from the center of each nucleic acid sequence of each polynucleotide. In some embodiments, the nucleic acid sequence of each polynucleotide is 150 bases or less in length. In some embodiments, at least one MRD-associated variant sequence is derived from a genomic sequence. In some embodiments, the genomic sequence is derived from cell-free DNA (cfDNA). In some embodiments, at least one MRD-associated variant sequence is present in the plurality of polynucleotides at a frequency of 0.001% to 0.1% relative to the wild-type genomic sequence. In some embodiments, the plurality of polynucleotides include about 500 variant sequences associated with MRD. In some embodiments, the nucleic acid sequence of each polynucleotide includes a variant sequence of at least one MRD-associated variant sequence. In some embodiments, at least one MRD-associated variant is present in the nucleic acid sequences of at least 150 genes. In some embodiments, at least one MRD-associated variant sequence includes a modification to the nucleic acid sequence of a tumor suppressor gene or oncogene. In some embodiments, the polynucleotide library further includes a background set of polynucleotides, where at least one MRD-associated variant sequence is at least one base pair different from the nucleic acid sequences of the polynucleotides in the background set. In some embodiments, the background set includes cell-free DNA (cfDNA).

[0005] Methods for preparing a polynucleotide library containing multiple polynucleotides are also provided herein. In some embodiments, the method includes the steps of providing at least one variant sequence associated with minimal residual disease (MRD) and synthesizing multiple polynucleotides containing at least one MRD-associated variant sequence to produce a polynucleotide library.

[0006] In some embodiments, the method further comprises the steps of providing a background set of polynucleotides and mixing the background set with a plurality of polynucleotides such that at least one MRD-associated variant sequence is present at a frequency of 2% or less relative to the wild-type genome sequence. In some embodiments, the synthesis includes chemical synthesis, surface synthesis, or coupling of nucleoside phosphoramidites. In some embodiments, the method further comprises the step of sequencing the polynucleotide library.

[0007] Methods for detecting minimal residual disease (MRD) in a sample are also provided herein. In some embodiments, the method includes the steps of providing a polynucleotide library described herein, contacting the polynucleotide library with a sample, and detecting the presence or absence of at least one variant associated with MRD in the sample. [Brief explanation of the drawing]

[0008] [Figure 1] Figures 1A and 1B are schematic diagrams illustrating probe designs according to embodiments of this disclosure. Figure 1A shows a probe design in which the probe does not contain a mutation, while Figure 1B shows a probe design in which the probe itself incorporates a variant sequence. [Figure 2] Figure 2 is a graph of the putative allele fractions for different variants according to the embodiments of this disclosure. Variants include single nucleotide substitutions (SBS), small indels, large indels (>5 base pairs), and structural variants (SV), and the putative allele fractions are shown in 2% increments from 0% to 14% for alignment and K-mer searches. [Figure 3] Figure 3 is a schematic diagram showing an experimental design for testing a probe according to an aspect of this disclosure. [Figure 4]Figure 4 is a schematic diagram illustrating a bioinformatics workflow for analyzing variant detection using a probe, according to an aspect of this disclosure. [Figure 5] Figure 5 is a panel of three bar graphs showing the target variant allele frequencies and the measured variant allele frequencies for a variant sequence-containing ("Ref+Alt") probe and a variant sequence-free ("Ref only") probe according to aspects of this disclosure. The measured variant allele frequency (VAF) is measured as a function of the target VAF for single nucleotide substitution (SBS) sites, indel sites, and all sites. [Figure 6] Figure 6 is a panel of three bar graphs showing the frequency and recall rate of target variant alleles with and without variant sequences ("Ref+Alt") according to aspects of this disclosure. Recall is measured as a function of the target VAF for single nucleotide substitution (SBS) sites, indel sites, and all sites. [Figure 7] Figure 7 shows a panel of dot graphs relating probe panel performance to the following Picard metrics: (A) Off-target rate, (B) Fold80 uniformity, (C) Zero-coverage target rate, and (D) Average target coverage. The probe panel includes breast, CRC, kidney, lung, and melanoma. [Figure 8] Figure 8 is a panel of five bar graphs showing mean positive variant allele frequencies (VAFs) according to aspects of this disclosure, where these VAFs were detected at different target VAF levels for each of the probe panels in Figure 7. [Figure 9] Figure 9 is a panel of five dot graphs showing variant recall at different variant allele frequency (VAF) levels for each of the probe panels in Figure 7, according to aspects of this disclosure. [Figure 10]Figure 10 is a panel of five scatter plots showing the mean recall rates of variant allele frequencies (VAFs) for different variants according to aspects of this disclosure. Variants include single nucleotide substitutions (SBSs), single nucleotide indels, small indels, medium indels, and large indels. [Figure 11] Figure 11 is a panel of five scatter plots showing the sample-by-sample reproduction of variant allele frequencies (VAFs) for different variants according to aspects of this disclosure. Variants include small indels (2–4 base pairs), medium indels (5–9 base pairs), and large indels (10+ base pairs). [Figure 12] Figure 12 is a panel of five scatter plots showing the mean recall of variant allele frequencies (VAFs) for different variants according to aspects of this disclosure. Variants include single nucleotide substitutions (SBSs), single nucleotide indels, small indels, medium indels, and large indels. [Figure 13A] Figures 13A–13B are panels of receiver operating characteristic (ROC) curve analyses to illustrate the relationship between sensitivity and specificity of different target site numbers at each variant allele frequency (VAF) level, according to aspects of this disclosure. Figure 13A is a panel of 10 ROC curve analyses for 0.01% VAF and 0.05% VAF at sites 10, 20, 50, 100, and 197. [Figure 13B] Figures 13A–13B are panels of receiver operating characteristic (ROC) curve analyses to illustrate the relationship between sensitivity and specificity of different target site numbers at each variant allele frequency (VAF) level, according to aspects of this disclosure. Figure 13B is a panel of 10 ROC curve analyses for 0.10% and 2.00% VAFs at 10, 20, 50, 100, and 197 sites. [Figure 14] Figure 14 is a panel of four scatter plots showing the results for the number of positive sites and the very low variant allele frequencies (VAFs) measured on linear and logarithmic scales, according to aspects of this disclosure. [Figure 15]Figure 15 is a schematic diagram showing a plate having 256 clusters according to an aspect of the present disclosure, where each cluster has 121 locates from which polynucleotides extend. [Figure 16] Figure 16 is a schematic diagram illustrating a computer system according to an aspect of the present disclosure. [Figure 17] Figure 17 is a block diagram illustrating the architecture of a computer system according to an aspect of this disclosure. [Figure 18] Figure 18 illustrates a network configured to incorporate multiple computer systems, multiple mobile phones and personal digital assistants, and network-attached storage (NAS) according to an aspect of the present disclosure. [Figure 19] Figure 19 is a block diagram illustrating a multiprocessor computer system using a shared virtual address memory space according to an aspect of this disclosure. [Figure 20] Figure 20A is a frequency plot of polynucleotide representations across a plate, from the synthesis of 29,040 unique polynucleotides from 240 clusters, each having 121 polynucleotides, according to an aspect of the present disclosure. Polynucleotide frequencies are shown as a function of abundance measured in absorbance units. Figure 20B is a series of frequency plots of polynucleotide representations across each individual cluster, according to an aspect of the present disclosure (control clusters are identified by boxes). Polynucleotide frequencies are shown as a function of abundance measured in absorbance units. [Figure 21A] Figure 21A is a schematic diagram illustrating a workflow according to an aspect of this disclosure for forming an adapter-ligated polynucleotide by attaching an adapter containing a unique molecular identifier (UMI) to a polynucleotide. [Figure 21B]Figure 21B is a schematic diagram illustrating a workflow for amplifying adapter-ligated polynucleotides to form a library for sequencing, according to an aspect of this disclosure. [Figure 21C] Figure 21C is a schematic diagram illustrating a workflow for the synthesis of a polynucleotide adapter containing a unique molecular identifier (UMI) according to an aspect of this disclosure. [Figure 21D] Figure 21D is a schematic diagram illustrating a workflow for the synthesis of a polynucleotide adapter containing a unique molecular identifier (UMI) according to an aspect of the present disclosure, the method comprising PCR extension of one strand of the adapter. [Figure 21E] Figure 21E is a schematic diagram illustrating a workflow for the synthesis of a polynucleotide adapter containing a unique molecular identifier (UMI) according to an aspect of the present disclosure, the method comprising PCR extension of one strand of the adapter, followed by restriction enzyme cleavage. [Figure 21F] Figure 21F is a schematic diagram illustrating a workflow for the synthesis of a polynucleotide adapter containing a unique molecular identifier (UMI) according to an aspect of the present disclosure, the method comprising restriction enzyme cleavage. [Figure 22] Figure 22 is a schematic diagram illustrating the workflow of duplex sequencing analysis for variant identification according to an aspect of this disclosure. An asterisk ("*") indicates a potential error introduced by PCR or sequencing, and a plus sign ("+") indicates a true variant. [Figure 23A] Figure 23A is a graph illustrating the design of a variant-targeting synthetic circulating tumor DNA (ctDNA) according to an aspect of this disclosure. Multiple overlapping, or "tiled," polynucleotides are configured to contain the variant site (indicated by a star). Labeled oligonucleotides (y-axis) are shown as a function of labeled genomic coordinates (x-axis) from 0 to 300 at 100-unit intervals. [Figure 23B]Figure 23B is a frequency graph showing the distribution of indel sizes for a synthetic circulating tumor (ctDNA) library containing short, medium (5–10 base pairs), and large (~30 base pairs) sized variants according to aspects of this disclosure. Positive numbers indicate insertions, and negative numbers indicate deletions. The number of labeled variants from 0 to 40 at 20-unit intervals (y-axis) is shown as a function of the labeled indel size from -30 to 10 base pairs at 10-unit intervals. [Figure 23C] Figure 23C is a line graph showing the abundance versus size of background cell-free DNA (cfDNA) according to an embodiment of this disclosure. The abundance was measured by a signal in fluorescence units (FU), and background cfDNA was obtained from healthy donor plasma. The abundance (y-axis) from 0 to 400 fluorescence units at 50-unit intervals is shown as a function of labeled base pairs (bp) at 35, 100, 150, 200, 300, 400, 500, 600, 1000, 2000, and 10380 base pairs. Peaks 1 and 2 are labeled. [Modes for carrying out the invention]

[0009] Minimally residual disease (MRD), also known as molecular residual disease, refers to the small number of tumor cells that may remain in a patient after therapeutic intervention. Detecting these residues and monitoring their abundance is a promising prognostic marker for identifying individuals at risk of relapse or those requiring adjuvant therapy. Due to the low abundance of circulating tumor DNA (ctDNA) in samples obtained during remission, MRD assays need to be highly sensitive. Furthermore, each individual will have a different set of somatic variants, requiring personalized solutions for detection. What is needed is a personalized NGS assay with high sensitivity and specificity for MRD diagnosis.

[0010] This specification provides compositions and methods, including panels and kits, that can be used to address this need and enable accurate assessment of MRD. In some examples, such compositions and methods can enable users to design and / or manufacture fully individualized MRD panels. In some examples, panels can contain up to 100, 200, 300, 400, or 500 targets. In some examples, the design and / or manufacture of the panels can be carried out in as little as six days.

[0011] This specification provides polynucleotide libraries containing at least one variant sequence. In some examples, at least one variant sequence is associated with minimal residual disease (MRD). In some examples, at least one variant sequence is present at a frequency of 0.001% to 0.1% relative to the wild-type genome sequence. In some examples, at least one variant sequence is located within 20 bases of the center of each of the polynucleotides. In some examples, at least one variant sequence is located within 10% of the center of each of the polynucleotides. In some examples, the position of each at least one variant sequence in each of the polynucleotides includes a distribution that includes a mean. In some examples, the mean is at the center of each sequence. In some examples, the mean is within 20 bases of the center of each sequence. In some examples, the mean is within 10% of the center of each sequence. In some examples, the polynucleotides are 150 bases or less in length. In some examples, the polynucleotides are double-stranded.

[0012] This specification further provides kits for detecting MRD. In some examples, the kit detects MRD in samples such as biological samples from patients and / or users. In some examples, the kit includes a polynucleotide library containing at least one variant sequence described herein. In some examples, the kit further includes instructions for use of the kit and / or packaging configured to hold and describe the contents of the kit.

[0013] Furthermore, this specification provides a method for preparing a polynucleotide library containing at least one variant sequence described herein. The library can be used to detect MRDs. In some examples, the method includes the step of providing at least one variant sequence associated with an MRD. In some examples, the method includes the step of synthesizing a plurality of polynucleotides containing at least one variant sequence. In some examples, the method further includes the step of providing a background set of background polynucleotides. In some examples, the method further includes the step of mixing the background set with the plurality of polynucleotides containing at least one variant sequence. In some examples, the step of mixing the background set with the plurality of polynucleotides includes mixing the background set with the plurality of polynucleotides such that at least one variant sequence is present at a frequency of 0%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence.

[0014] Methods for detecting MRD are further provided herein. In some examples, MRD may be detected in samples such as biological samples from patients and / or users. In some examples, the method includes the step of providing a polynucleotide library containing at least one variant sequence, as described herein. In some examples, the method includes the step of contacting the polynucleotide library with a sample. In some examples, the method includes the step of detecting the presence or absence of at least one variant sequence associated with MRD in the sample.

[0015] definition

[0016] Throughout this disclosure, numerical features are presented in range form. It should be understood that this range form is merely for convenience and brevity and should not be interpreted as a firm limitation on the scope of any embodiment. Therefore, range descriptions should be considered as specifically disclosing all possible subranges, and similarly, up to one-tenth of the lower limit unit, unless the context explicitly indicates otherwise. For example, a range description such as 1-6 should be considered to have specifically disclosed subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and similarly, individual values ​​within ranges such as 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the width of the range. The upper and lower limits of such intervening ranges may independently be contained within smaller ranges and are also encompassed within the invention, subject to any explicitly excluded limitations in the expressed range. If the stated scope includes one or both of the limits, the scope excluding either or both of such included limits is also included in the present invention unless the context explicitly indicates otherwise.

[0017] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit any embodiment. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural form unless the context otherwise expressly indicates otherwise. It is further understood that the terms “contains” and / or “contains,” as used herein, identify the presence of an expressed feature, integer, process, operation, element, and / or component, but do not preclude the presence or addition of one or more other features, integers, processes, operations, elements, components, and / or groups thereof. As used herein, the terms “and / or” include any and all combinations of one or more of the enumerated items relating to the subject.

[0018] Unless otherwise explicitly stated or evident from the context, the term “about” in relation to a number or range means, as used herein, a number or range of numbers that is explicitly stated and plus or minus 10%, or a number that is 10% lower than the lower limit and 10% higher than the upper limit for any value listed for a range.

[0019] As used herein, the terms “preselected sequence,” “predefined sequence,” or “predetermined sequence” are used interchangeably. The terms mean that the polymer sequence is known and selected prior to the synthesis or assembly of the polymer. In particular, various aspects of the present invention are described herein with respect to the preparation of nucleic acid molecules and the sequences of polynucleotides, which are known and selected prior to the synthesis or assembly of nucleic acid molecules.

[0020] As used herein, the term “nucleic acid” includes single-stranded molecules as well as double-stranded or triple-stranded nucleic acid molecules. In double-stranded or triple-stranded nucleic acid molecules, their nucleic acid strands do not need to be coextensive (i.e., a double-stranded nucleic acid does not need to be double-stranded along the entire length of both strands). Nucleic acid sequences, if provided, are enumerated in the 5' to 3' direction unless otherwise stated. The methods described herein result in the production of isolated nucleic acid molecules. The methods described herein further result in the production of isolated and purified nucleic acids. The length of a nucleic acid molecule (e.g., polynucleotide), if provided, is given as the number of bases and is abbreviated as nucleotide (nt), base or base pair (bp), kilobase (kb), megabase (Mb), or gigabase (Gb), etc.

[0021] When used herein, the terms “polynucleotide,” “oligonucleotide,” “oligonucleotide,” “oligo,” and “nucleic acid molecule” are interchangeable. A library of synthetic (i.e., de novo or chemically synthesized) polynucleotides described herein may contain multiple polynucleotides that collectively encode one or more genes or gene fragments. In some examples, a polynucleotide library contains coding or non-coding nucleic acid sequences. In some examples, a polynucleotide library encodes multiple cDNA sequences. The reference gene sequence on which the cDNA sequence is based may contain introns, but the cDNA sequence excludes introns. The polynucleotides described herein may encode genes or gene fragments of biological origin. Exemplary organisms include, but are not limited to, prokaryotes (e.g., bacteria) and eukaryotes (e.g., mice, rabbits, humans, and non-human primates). In some examples, a polynucleotide library contains one or more polynucleotides, each of which encodes a sequence of multiple exons. Each polynucleotide in the libraries described herein may encode a different nucleic acid sequence, i.e., a non-identical nucleic acid sequence. In some examples, each polynucleotide in the libraries described herein contains at least one portion that is complementary to the nucleic acid sequence of another polynucleotide in the library. Unless otherwise stated, the polynucleotide sequences described herein may include DNA or RNA. The polynucleotide libraries described herein may contain at least 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10000, 20000, 30000, 50000, 100000, 200000, 500000, 1000000, or more than 1,000,000 polynucleotides. The polynucleotide libraries described herein may contain 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10000, 20000, 30000, 50000, 100000, 200000, 500000 or less, or 1000000 or less polynucleotides.The polynucleotide libraries described herein may contain 10-500, 20-1000, 50-2000, 100-5000, 500-10000, 1000-5000, 10000-50000, 100000-50000, or 50000-1000000 polynucleotides. The polynucleotide libraries described herein may contain approximately 370,000, 400,000, 500,000, or more different polynucleotides.

[0022] Variant Library

[0023] This specification provides polynucleotide libraries configured to detect or measure one or more different sequences. In some examples, these libraries are used as a reference or control. Known methods for producing such libraries may involve isolating nucleic acids from a biological source (blood, plasma, cells, or patient) with an established disease or condition. However, in some cases, such methods provide libraries containing contamination from those biological sources. In some examples, libraries are prepared from biological samples to mimic cell-free DNA (cfDNA) by restriction digestion, sonication, or other methods that produce short nucleic acid fragments. These methods may not mimic the natural fragmentation profile of cfDNA. Furthermore, low-abundance variant sequences may not be detectable from libraries derived from living organisms. This specification provides methods for designing and de novo synthesizing polynucleotide libraries (or sample sets) useful for detecting or measuring the frequency of variant sequences. In some examples, such libraries provide improved accuracy for diagnosing a disease or illness and are substantially free from biological contamination. In some cases, synthetic polynucleotide libraries offer further control over library contents, reliability / reproducibility, lack of dependence on fragmentation methods, and / or other advantages over conventional cell-derived libraries. In some cases, such libraries are mixed with control nucleic acid molecules (e.g., cfDNA) to generate a reference standard for specific variant allele frequencies (VAFs).

[0024] In some embodiments, the polynucleotide library comprises multiple polynucleotides (e.g., sample sets) derived from a genomic sequence. In some examples, the multiple polynucleotides may comprise at least one variant sequence associated with a disease or illness. In some examples, the at least one variant sequence comprises one or more changes compared to the wild-type genomic sequence or background polynucleotides. In some embodiments, the polynucleotide library comprises a background set comprising background polynucleotides, the background set comprising cell-free DNA (cfDNA). In some examples, at least some of the polynucleotides are tiled across each of the at least one variant sequence. In some examples, the polynucleotides are not tiled across each of the at least one variant. In some examples, the background cfDNA is obtained, induced, or amplified from a cell line or patient sample.

[0025] The observed variant allele frequencies may be lower than expected. This can be due to one or more biases, such as capture bias or alignment bias. This observation is commonly illustrated in Figure 2. Generally, because libraries (e.g., panels) target a reference allele, alternative alleles will have a mismatch to the probe, resulting in capture bias. In some cases, if the reads contain significant differences from the reference allele (e.g., in the context of sequencing errors), they are more difficult to align with the alignment algorithm (e.g., BWA), resulting in alignment bias. Both capture bias and alignment bias may favor the reference allele over alternatives and tend to be stricter for larger edit distances (e.g., single nucleotide polymorphism (SNP) < short indel < long indel < structural variant (SV)).

[0026] In some embodiments, if the sequences of a genome containing variant sequences (genomic variants) are known, they can be incorporated into polynucleotides. For example, if the sequence of a cancer genome is already known, it can be incorporated to avoid probe mismatch or sequence mismatch. In some examples, the cancer includes MRD. In some examples, the library not only targets the variant site but also incorporates the variant sequence into the probe itself. In some examples, this design can be used to avoid reference allele bias, as described herein.

[0027] This specification provides a library of polynucleotides containing a predetermined variant sequence (e.g., a variant). In some examples, the polynucleotide library contains at least 1, 5, 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000, or at least 2000 variants. In some examples, the polynucleotide library contains about 1, 5, 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000, or about 2000 variants. In some examples, the polynucleotide library includes variants of 1, 5, 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000 or less, or 2000 or less. In some examples, the polynucleotide library includes variants of 1-500, 5-500, 10-500, 10-2000, 10-150, 15-500, 20-1000, 50-500, 50-750, 50-1000, 100-1000, 100-500, 100-750, 250-800, 400-1000, or 400-2000. In some examples, polynucleotide libraries are designed to contain approximately 100, 150, or 200 targets, each having about 1, 2, 3, 4, 5, or 6 variants per tissue origin.

[0028] The polynucleotides provided herein may be tiled across nucleic acid regions. In some examples, tiling refers to the design of polynucleotides (or their complements or reverse complements) that cover or span a target region (e.g., a variant). In some examples, tiling results in increased detection sensitivity for a probe targeting a variant, or in the design of a corresponding standard, control, or reference. This may be beneficial for regions with low abundance or regions containing sequences that are difficult to sequence (repetitive, high / low GC, or other difficulties). In some examples, each tiled polynucleotide for a target region is distinct from each other tiled polynucleotide for the target region. In some examples, such a tiling design contains approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 27, 30, 32, 35, 40, 45, or approximately 50 polynucleotides tiled across a region (e.g., a variant). In some examples, such a tiling design contains at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, 35, 40, 45, or at least 50 polynucleotides tiled across a region. In some examples, such a tiling design contains 10–100, 5–50, 2–50, 25–50, 30–40, or 30–60 polynucleotides tiled across a region. In some examples, the tiled polynucleotides contain at least one overlapping region with another polynucleotide. In some examples, both the 5' and 3' ends of the tiled polynucleotides overlap with the ends of the adjacent tiled polynucleotides. In some examples, one or more tiled polynucleotides are tiled with a certain offset value such that the first polynucleotide starts at a different position than the next tiled polynucleotide. In some examples, the offset is approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 17, 20, 25, or 30 bases.In some examples, the offset is 1–30, 1–20, 1–10, 1–8, or 2–5 bases. In some examples, the length of at least some of the polynucleotides is 20–500, 50–500, 75–500, 100–200, 100–500, 200–500, 100–250, 100–200, 100–1000, 250–500, or 250–1000. In some examples, the length of at least some of the polynucleotides is approximately 50, 75, 100, 125, 150, 155, 160, 165, 170, 175, 180, 190, 200, or 225 bases. In some examples, the length of at least 80% of the polynucleotides is 20–500, 50–500, 75–500, 100–200, 100–500, 200–500, 100–250, 100–200, 100–1000, 250–500, or 250–1000. In some examples, the length of at least 80% of the polynucleotides is approximately 50, 75, 100, 125, 150, 155, 160, 165, 170, 175, 180, 190, 200, or 225 bases. In some examples, the length of at least 90% of the polynucleotides is 20–500, 50–500, 75–500, 100–200, 100–500, 200–500, 100–250, 100–200, 100–1000, 250–500, or 250–1000. In some examples, the length of at least 90% of the polynucleotides is approximately 50, 75, 100, 125, 150, 155, 160, 165, 170, 175, 180, 190, 200, or 225 bases. In some examples, at least some of the polynucleotides are double-stranded. In some examples, at least 50%, 60%, 70%, 75%, 80%, 90%, 95%, or at least 98% of the polynucleotides are double-stranded. In some cases, multiple polynucleotides containing at least one variant sequence associated with a low-variant-frequency allele (e.g., MRD) may be approximately 150 bases or less, 170 bases or less, or 200 bases or less in length.

[0029] Variant sequences may exist at predetermined frequencies relative to other variant sequences in the library (e.g., sample library). In some examples, at least 80% of at least one variant sequence exist at frequencies that differ from the expected frequencies for the uniformly pooled variants by only 20%, 15%, 12%, 10%, 8% or less, or 5% or less. In some examples, at least 90% of at least one variant sequence exist at frequencies that differ from the expected frequencies for the uniformly pooled variants by only 20%, 15%, 12%, 10%, 8% or less, or 5% or less. In some examples, at least 95% of at least one variant sequence exist at frequencies that differ from the expected frequencies for the uniformly pooled variants by only 20%, 15%, 12%, 10%, 8% or less, or 5% or less. In some examples, at least 99% of at least one variant sequence exist at frequencies that differ by only 20%, 15%, 12%, 10%, 8% or less, or 5% or less compared to the expected frequencies for the uniformly pooled variants.

[0030] The compositions (libraries) described herein may comprise a plurality of polynucleotides, each containing at least one variant sequence associated with minimal residual disease (MRD). In some examples, the at least one variant sequence is located within 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 bases of the center in each of the plurality of polynucleotides. The center of the sequence may generally include a position in the sequence that is as far as possible from the ends of the sequence. In some examples, the at least one variant sequence is located within at least 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 bases of the center in each of the plurality of polynucleotides. In some examples, the at least one variant sequence is located within up to 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 bases of the center in each of the plurality of polynucleotides. In some examples, at least one variant sequence is 2-5, 2-10, 2-15, 2-20, 2-25, 2-30, 2-35, 2-40, 2-45, 2-50, 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 10-15, 10-20, 10-25, 10-30, 10-35, 10-40, 10-45 , 10-50, 15-20, 15-25, 15-30, 15-35, 15-40, 15-45, 15-50, 20-25, 20-30, 20-35, 20-40, 20-45, 20-50, 25-30, 25-35, 25-40, 25-45, 25-50, 30-35, 30-40, 30-45, 30-50, 35-40, 35-45, 35-50, 40-45, 40-50, or 45-50 bases. In some examples, at least one variant sequence is within 2%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, or 50% of the center of each sequence of multiple polynucleotides. In some examples, at least one variant sequence is located within at least 2%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the center of each of the multiple polynucleotides.In some examples, at least one variant sequence is located within 2%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the center of each of the multiple polynucleotides. In some examples, at least one variant sequence is located within 2%-5%, 2%-10%, 2%-15%, 2%-20%, 2%-25%, 2%-30%, 2%-40%, 2%-50%, 5%-10%, 5%-15%, 5%-20%, 5%-25%, 5%-30%, 5%-40%, 5%-50%, 10%-15%, or 10% of the center of each of the multiple polynucleotide sequences. The percentages are within ~20%, 10%~25%, 10%~30%, 10%~40%, 10%~50%, 15%~20%, 15%~25%, 15%~30%, 15%~40%, 15%~50%, 20%~25%, 20%~30%, 20%~40%, 20%~50%, 25%~30%, 25%~40%, 25%~50%, 30%~40%, 30%~50%, or 40%~50%. In some examples, the position of each variant sequence in each sequence of multiple polynucleotides includes a distribution that includes the mean. In some examples, the mean is the center of each sequence. In some examples, the mean is within 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 bases from the center of each sequence. In some examples, the mean is within at least 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 bases of the center of each sequence. In some examples, the mean is within at most 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 bases of the center of each sequence.In some examples, the mean is 2-5, 2-10, 2-15, 2-20, 2-25, 2-30, 2-35, 2-40, 2-45, 2-50, 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 10-15, 10-20, 10-25, 10-30, 10-35, 10-40, 10-45, 10-50, 15-20, 1 It lies within 5-25, 15-30, 15-35, 15-40, 15-45, 15-50, 20-25, 20-30, 20-35, 20-40, 20-45, 20-50, 25-30, 25-35, 25-40, 25-45, 25-50, 30-35, 30-40, 30-45, 30-50, 35-40, 35-45, 35-50, 40-45, 40-50, or 45-50 bases. In some examples, the mean lies within 2%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, or 50% of the center of each sequence. In some examples, the mean is within at least 2%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, or 50% of the center of each sequence. In some examples, the mean is within up to 2%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, or 50% of the center of each sequence. In some examples, the mean is within 2%-5%, 2%-10%, 2%-15%, 2%-20%, 2%-25%, 2%-30%, 2%-40%, 2%-50%, 5%-10%, 5%-15%, 5%-20%, 5%-25%, 5%-30%, 5%-40%, 5%-50%, 10%-15%, 10%-20%, 10%-25%, 1 The percentages are within the following ranges: 0%-30%, 10%-40%, 10%-50%, 15%-20%, 15%-25%, 15%-30%, 15%-40%, 15%-50%, 20%-25%, 20%-30%, 20%-40%, 20%-50%, 25%-30%, 25%-40%, 25%-50%, 30%-40%, 30%-50%, or 40%-50%. In some examples, the distribution is normal.

[0031] The compositions described herein may include a backgrounded set (or library) of polynucleotides. In some examples, the background set mimics background cfDNA present in a patient sample. In some examples, the background polynucleotides are mixed with sample polynucleotides (e.g., polynucleotides containing variant sequences, variant polynucleotide library) to generate a reference standard or control. In some examples, the standard or control includes a variant sequence having 0%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, 2%, 5%, 10%, 15%, or 20% VAF relative to the wild-type genome sequence. In some examples, the background polynucleotide set includes a wild-type region corresponding to the position of at least one variant sequence. In some examples, the wild-type sequence is derived from a reference database or sample. In some examples, the background polynucleotide set contains wild-type regions corresponding to the positions of at least 1, 2, 5, 10, 15, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, or at least 500 variants. In some examples, the wild-type regions appear within 30%, 25%, 20%, 15%, 12%, 10%, 9%, 8%, 7%, or 5% of the variant frequencies in the variant set (e.g., the sample set). In some examples, the background set contains low levels of mutation. In some examples, at least one background polynucleotide contains a variant present at a frequency of 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some examples, at least 1% of the background polynucleotides contain variant sequences present at frequencies of 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some examples, the background set is synthesized from a given sequence. In some examples, the given sequence reflects the desired variant frequencies.In some examples, synthetic background sets are used to calibrate instruments or methods by providing control over variant frequencies. In some examples, synthetic background sets are configured to mimic variant frequencies corresponding to a specific sample or disease state.

[0032] In some cases, the background set contains background polynucleotides. In some cases, the background set contains background polynucleotides substantially derived from wild-type sequences. In some cases, the background set is derived from or isolated from a healthy individual. In some cases, the healthy individual is male. In some cases, the healthy individual is female. In some cases, the healthy individual is 40, 35, 30, 25 years or younger; 20 years or 15 years old. In some cases, the background set is obtained from a biological sample. In some cases, the biological sample includes blood, plasma, or another nucleic acid source. In some cases, the background set contains cfDNA. In some cases, the background set contains at least 2, 5, 10, 100, 200, 500, 1,000, 10,000, 100,000, 500,000 polynucleotides, 1 million, 5 million, 10 million, 50 million, 100 million, 200 million, or more than 500 million polynucleotides. In some examples, the highest abundance of polynucleotides in the background set is 100–500, 50–500, 75–250, 50–750, 50–300, 100–300, 100–200, 125–300, 150–175, 150–185, or 125–200 bases in length. In some examples, at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or at least 97% of the polynucleotides in the background set are mononucleosomes or dinucleosomes. In some examples, the ratio of mononucleosome polynucleotides to dinucleosome polynucleotides is 50:50–90:10, 60:40–90:10, 60:40–95:5, 70:30–95:5, 70:30–90:10, or 80:20–95:5.

[0033] The polynucleotide libraries described herein may be mixed to form a standard (reference). In some examples, the standard (reference) includes both a sample (variant) polynucleotide set and a control polynucleotide set. In some examples, the standard, which includes both a sample polynucleotide set and a control polynucleotide set, further includes a liquid buffer. In some examples, the buffer includes TE or TBE buffer. In some examples, the standard includes 50%, 40%, 30%, 25%, 20%, less than or equal to 15%, or less than or equal to 10%, of the sample polynucleotide relative to the background polynucleotide. In some examples, the standard includes a variant sequence having a VAF of 0%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some examples, the standard is subjected to one or more quality control operations, including one or more of the following: fluorescence / UV DNA quantification, electrophoretic size analysis, sequencing, ddPCR analysis, or other analytical techniques. In some cases, the sample polynucleotide set is subjected to one or more quality control operations, including one or more of the following: fluorescence / UV DNA quantification, electrophoretic size analysis, sequencing, ddPCR analysis, or other analytical techniques, before being mixed with the background polynucleotide set. In some cases, the sample polynucleotide is ligated to an adapter containing a unique molecular identifier (UM1) as described herein.

[0034] This specification provides methods for preparing polynucleotide libraries as described herein. In some examples, the polynucleotide library may be used to detect variant sequences having low variant allele frequencies. In some examples, the polynucleotide library can be used to detect MRDs. In some examples, the method includes a step of providing at least one variant sequence. In some cases, at least one variant sequence may be associated with an MRD. In some examples, the method further includes a step of synthesizing a plurality of polynucleotides containing at least one variant sequence. In some examples, the method further includes a step of providing a background set as described herein. In some examples, the method further includes a step of mixing the background set with a plurality of polynucleotides containing at least one variant sequence. In some examples, the step of mixing the background set with the plurality of polynucleotides includes mixing the background set with the plurality of polynucleotides such that at least one variant sequence is present at a frequency of 0%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some examples, the synthesis step includes chemical synthesis. In some examples, the synthesis step involves synthesis on a surface. In some examples, the synthesis step involves coupling of nucleoside phosphoramidites. In some examples, the method further includes a step of sequencing the polynucleotide library. In some examples, the method further includes ddPCR measurement of the polynucleotide library. In some examples, the method further includes fluorescence / UV DNA quantification and size distribution of the polynucleotide library.

[0035] Synthetic libraries containing variant sequences (e.g., sample libraries, sample sets, variant sets) may have fewer contaminants (less contamination) than libraries derived from biological samples. In some cases, lower levels of contaminants result in improved performance as a reference standard. In some cases, contamination includes, but is not limited to, cellular components, lipids, RNA, proteins, or other biomolecules derived from biological sources. In some cases, biological sources include plasma, cells, blood, or other nucleic acid sources. In some cases, synthetic libraries are prepared or stored in buffer. In some cases, synthetic libraries are free of biological contaminants by at least 95%, 96%, 97%, 98%, 99%, 99.5%, or at least 99.7%.

[0036] Genome variant

[0037] Genetic variants between individual populations ("variants" in nucleic acid sequences) can provide information regarding disease risk, individual identification, response to drug treatment, or susceptibility to environmental factors such as toxins. This specification describes compositions and methods involving the synthesis of polynucleotide libraries containing such variant sequences. In some examples, variant sequences include single nucleotide polymorphisms (SNPs), single nucleotide variations (SNVs), indels, copy number variations, translocations, fusions, inversions, or structural variants. In some examples, SNPs differ between individuals in the same population. In some cases, SNPs differ between individuals in different populations. In some examples, SNVs contain a single nucleotide mutation with no limitation of frequency. In some examples, the polynucleotide libraries described herein (e.g., probe libraries) are used to identify such variants after sequencing. In some examples, the polynucleotide libraries are configured to enrich nucleic acid molecules (e.g., genomic fragments) containing variant sequences. In some examples, such nucleic acid molecules are captured using the polynucleotide library and sequenced for variant calling. In some cases, variant calls may be evaluated by comparing them to known variant sequences using metrics such as reproducibility and / or accuracy for one or all of the variant sequences. In some cases, the SNP or SNV is heterozygous. In some cases, the SNP or SNV is homozygous. In some cases, the SNP or SNV is homozygous in matching with a reference sequence. In some cases, the variant sequence is homozygous for states other than those observed in the human reference genome. In some cases, the variant sequence is identified after sequencing by comparison with a reference database. In some cases, the reference database includes GiAB, dbSNP, DoGSD, dbGaP, clinvar, ncbi, refseq, refSNP, COSMIC, or any other database containing known variants.In some cases, variant sequences include insertions, deletions, fusions, duplications, frameshifts, repeat extensions, or substitutions. In some cases, variant sequences include copy number variants (CNVs), microsatellite instability, loss of heterozygosity (LOH), DNA methylation, early stop codons, trinucleotide repeats, translocations, somatic rearrangements, allelomorphisms, single nucleotide variants (SNVs), indels, splice variants, regulatory variants, copy number variants, or fusions. In some cases, indels are 1–50, 1–25, 1–20, 1–15, 2–20, 5–25, 5–15, or 5–10 nucleotides long. In some cases, indels are 1, 2, 3, 5, 7, 8, 10, 12, 15, 17, 20, 25 nucleotides or less, or 50 nucleotides or less. In some cases, the variants described herein are located within a gene. In some examples, the libraries described herein include variant sequences found in at least 2, 5, 10, 15, 20, 25, 30, 50, 60, 75, 100, 125, 150, 200, 250, 300, 400, or at least 500 genes. In some examples, the libraries described herein include variant sequences found in 5-500, 5-100, 5-50, 10-200, 10-100, 25-500, 25-250, 25-150, 50-150, 50-250, 50-500, or 75-500 genes.

[0038] In some embodiments, variant sequence identification is achieved using imputed data. In some examples, the identification of variant sequences near known or detected variant sequences provides identity information for variant sequences that are not measured or lack sequencing data to be accurately named. In some examples, unmeasured (or unknown) genomic variants are within 100, 500, 1,000, 10,000, 100,000, or 1,000,000 bases, or more, of measured (or identified) genomic variants, depending on linkage disequilibrium (the non-random association of alleles to different variants within a population). In some examples, linkage disequilibrium can be inferred by utilizing information about recombination rates observed in the genome or population. In some examples, recombination rates, gene distance maps, and variants themselves may vary between different populations.

[0039] Variants can exist at different frequencies in a population of individuals, a single individual, a tissue, or other groups, such as within a genome. In some cases, genomic variants co-occur in less than 0.001, 0.01, 0.1, 0.5, 1, 1.5, 2, 5, 10, 20, 25, 50%, or less than 75% of individuals in a group. In some cases, genomic variants co-occur in more than 0.001, 0.01, 0.1, 0.5, 1, 1.5, 2, 5, 10, 20, 25, 50%, or more than 75% of individuals in a group. In some cases, genomic variants co-occur in approximately 0.001, 0.01, 0.1, 0.5, 1, 1.5, 2, 5, 10, 20, 25, 50%, or about 75% of individuals in a group. In some cases, genomic variants co-occur in 0.1–10%, 0.001–10%, 0.01–10%, 0.01–1%, 0.001–1%, 0.1–25%, 0.1–10%, or 0.1–5% of individuals in a group. In some cases, the occurrence of variants is called the variant allele frequency (VAF).

[0040] This specification describes variant sequences for detecting diseases or illnesses. In some cases, the diseases or illnesses are proliferative disorders. In some cases, the diseases or illnesses are associated with low variant allele frequencies, such as minimal residual disease (MRD). In some cases, the diseases or illnesses are viral or bacterial diseases or illnesses. In some cases, the diseases or illnesses are cancers. In some cases, the variant sequences are present in oncogenes or tumor suppressor genes. In some examples, variant sequences are found in genes ABL1, ABL2, AKT1, ALK, APC, AR, ARAF, ARID1A, ATM, ATR, BAP1, BRAF, BRCA1, BRCA2, CCND1, CDC6, CDH1, CDK12, CDK4, CDX2, CTNNB1, DDR2, EGFR, EML4, ERBB2, ERBB3, ERG, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FLT3, FOXA1, FOXL2, GATA3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, ID It exists in one or more of the following: H2, JAK2, KDM5C, KDM6A, KIF5B, KIT, KRAS, MAP2K1, MAPK1, MET, MIR4728, ERBB2, MLH1, MPL, MYCN, MYD88, NCOA4, NF1, NF2, NFE2L2, NOTCH1, NPM1, NRAS, PBRM1, PDGFRA, PIK3CA, PTEN, PTPN11, RET, RHEB, RHOA, RIT1, ROS1, SETD2, SMAD4, SMO, SPOP, TERT, TMPRSS2, TP53, TPR, TSC1, and VHL.In some examples, the variant sequences are ABL1, ABL2, AKT1, ALK, APC, AR, ARAF, ARID1A, ATM, ATR, BAP1, BRAF, BRCA1, BRCA2, CCND1, CDC6, CDH1, CDK12, CDK4, CDX2, CTNNB1, DDR2, EGFR, EML4, ERBB2, ERBB3, ERG, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FLT3, FOXA1, FOXL2, GATA3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, KDM 5C, KDM6A, KIF5B, KIT, KRAS, MAP2K1, MAPK1, MET, MIR4728, ERBB2, MLH1, MPL, MYCN, MYD88, NCOA4, NF1, NF2, NFE2L2, NOTCH1, NPM1, NRAS, PBRM1, PDGFRA, PIK3CA, PTEN, PTPN11, RET, RHEB, RHOA, RIT1, ROS1, SETD2, SMAD4, SMO, SPOP, TERT, TMPRSS2, TP53, TPR, TSC1, and VHL are present in 1, 2, 3, 5, 7, 10, 20, 25, or more. In some cases, multiple variant sequences are present within a single gene. In some cases, variant sequences are present in 1, 2, 3, 5, 7, 10, 20, 25, or more of the genes. In some cases, variant sequences are present in one, two, three, five, seven, ten, twenty, twenty-five, or more genes associated with a disease or disorder.

[0041] In some cases, the disease or illness is breast cancer. In some cases, the variant sequence is present in one or more of the genes TP53, PIK3CA, ERBB2, MYC, FGFR1 / ZNF703, GATA3, CCND1, and CHD1 (e.g., CDH1*).

[0042] In some cases, the disease or illness is lung cancer. In some cases, the variant sequence is located in one or more of the genes KRAS (e.g., K117N), EGFR, ROS, ALK, and BRAF.

[0043] In some cases, the disease or illness is colorectal cancer. In some cases, the variant sequence is found in one or more of the genes TP53 APC, KRAS, BRAF, PIK3CA, SMAD4, FBXW7 (e.g., R465C), and NF1.

[0044] In some cases, the disease or illness is bladder cancer. In some cases, the variant sequence is located in one or more of the following: TP53, FGFR3 (e.g., S249C), ARID1A, and KDM6A.

[0045] In some cases, the disease or illness is prostate cancer. In some cases, the variant sequence is found in one or more of the following genes: ETS (e.g., ETS-TMPRSS2), SPOP (e.g., F133V), TP53, FOXA1 (e.g., R219), and PTEN.

[0046] In some cases, the disease or illness is renal cancer. In some cases, the variant sequence is located in one or more of the genes PBRM1, SETD2, BAP1, KDM5C, MTOR, VHL, MET, NF2, KDM6A, SMARCB1, FH, and CDKN2A.

[0047] In some cases, the disease or illness is melanoma. In some cases, the variant sequence is found in one or more of the genes NRAS, BRAF, PTEN, CDKN2A, MAP2K1, MAP2K2, GNAQ, GNA11, BAP (e.g., W196X).

[0048] In some examples, the variant sequence is one of the variants listed in Table 1 below.

[0049] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5]

[0050] In some examples, the variant sequence is one of the variants listed in Table 2 below.

[0051] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9] [Table 2-10] [Table 2-11] [Table 2-12] Table 2-13 Table 2-14 Table 2-15 Table 2-16 Table 2-17 Table 2-18 Table 2-19 Table 2-20 Table 2-21 Table 2-22 Table 2-23 Table 2-24 Table 2-25 Table 2-26 Table 2-27 Table 2-28 Table 2-29

[0052] In some examples, the variant sequence is one of the variants listed in Table 3 below.

[0053] [Table 3]

[0054] In some examples, the variant sequence is one of the variants listed in Table 4 below.

[0055] [Table 4-1] [Table 4-2]

[0056] In some examples, the variant sequence is one of the variants listed in Table 5 below. [Table 5]

[0057] In some examples, the variant sequence is one of the variants listed in Table 6 below.

[0058] [Table 6-1] [Table 6-2] [Table 6-3]

[0059] In some examples, the variant sequence is one of the variants listed in Tables 1-6 above.

[0060] Variant sequences (e.g., genomic variants) can be detected from a sample (e.g., a genomic sample) with varying degrees of reproducibility and precision. In some cases, the upper limit of detection is determined by the performance of a reference standard described herein. In some cases, the reference standard has a pre-selected variant frequency for comparison with a patient sample. In some cases, reproducibility represents the number of variant sequences detected from all variants expected to be detectable. In some cases, precision represents the number of variant sequences correctly called out of all detected variants. In some cases, variant sequences are detected with at least 30%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or at least 99% reproducibility. In some cases, variant sequences are detected with approximately 30%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or approximately 99% reproducibility. In some cases, variant sequences are detected with recall rates of approximately 10%–99%, 25–99%, 30–90%, 45–80%, 50–99%, 75–99%, or 90–99%. In some cases, variant sequences are detected with accuracy of at least 30%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or at least 99%. In some cases, variant sequences are detected with accuracy of approximately 30%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or approximately 99%. In some cases, variant sequences are detected with accuracy of approximately 10%–99%, 25–99%, 30–90%, 45–80%, 50–99%, 75–99%, or 90–99%.

[0061] A polynucleotide library may be designed to contain sequences that are identical to or complementary to (target, hybridize with) one or more variant sequences. In some examples, at least some of the polynucleotides are configured to hybridize to genomic regions containing at least two variant sequences. In some examples, at least some of the polynucleotides are configured to hybridize to genomic regions containing at least one, two, three, four, five, six, or more than six variant sequences. In some examples, at least some of the polynucleotides are configured to hybridize to genomic regions containing one to four variant sequences. In some examples, at least some of the polynucleotides are configured to hybridize to genomic regions containing one to two or three variant sequences. In some examples, at least 50% of the polynucleotides are configured to hybridize to genomic regions containing at least two variant sequences. In some examples, at least 50% of the polynucleotides are configured to hybridize to genomic regions containing at least one, two, three, four, five, six, or more than six variant sequences. In some examples, at least 50% of the polynucleotides are configured to hybridize to genomic regions containing 1–4 variant sequences. In some examples, at least 50% of the polynucleotides are configured to hybridize to genomic regions containing 1–2 or 3 variant sequences. In some examples, at least 25% of the polynucleotides are configured to hybridize to genomic regions containing at least 2 variant sequences. In some examples, at least 25% of the polynucleotides are configured to hybridize to genomic regions containing at least 1, 2, 3, 4, 5, 6, or more than 6 variant sequences. In some examples, at least 25% of the polynucleotides are configured to hybridize to genomic regions containing 1–4 variant sequences. In some examples, at least 25% of the polynucleotides are configured to hybridize to genomic regions containing 1–2 or 3 variant sequences.In some examples, at least 5% of each polynucleotide is configured to hybridize into genomic regions containing at least 2 variant sequences. In some examples, at least 5% of each polynucleotide is configured to hybridize into genomic regions containing at least 1, 2, 3, 4, 5, 6, or more than 6 variant sequences. In some examples, at least 5% of each polynucleotide is configured to hybridize into genomic regions containing 1–4 variant sequences. In some examples, at least 5% of each polynucleotide is configured to hybridize into genomic regions containing 1–2 or 3 variant sequences.

[0062] Polynucleotide libraries can be configured to bind to a large number of variant sequences. In some examples, polynucleotide libraries are collectively configured to bind to genomic regions containing approximately 50, 100, 200, 500, 800, 1000, 2000, 5000, 8000, 10,000, 20,000, 50,000, 80,000, 100,000, 250,000, 500,000, 750,000, 1 million, 1.5 million, 2 million, 2.5 million, 3 million, 3.5 million, 4 million, 4.5 million, or approximately 5 million variant sequences. In some examples, polynucleotide libraries are collectively configured to bind to genomic regions containing at least 50, 100, 200, 500, 800, 1000, 2000, 5000, 8000, 10,000, 20,000, 50,000, 80,000, 100,000, 250,000, 500,000, 750,000, 1 million, 1.5 million, 2 million, 2.5 million, 3 million, 3.5 million, 4 million, 4.5 million, or at least 5 million variant sequences. In some examples, polynucleotide libraries are collectively constructed to bind to regions of the genome containing variant sequences of 100–1000, 50–100, 50–500, 50–5000, 50–10,000, 100,000–5,000,000, 250,000–3,000,000, 500,000–2,000,000, 750,000–4,000,000, 1,000,000–5,000,000, 1,000,000–3,000,000, 1,000,000–4,000,000, or 4,000,000–6,000,000.

[0063] Polynucleotide libraries for identifying variant sequences can be optimized. In some cases, the library is uniform (each unique polynucleotide is represented equally). In some cases, the library is not uniform. In some cases, polynucleotides are represented in amounts no more than 1.5 times the mean representation of the polynucleotide library. In some cases, polynucleotides are represented in amounts no more than 2 times the mean representation of the polynucleotide library. In some cases, polynucleotides are represented in amounts no more than 1.2 times the mean representation of the polynucleotide library. In some cases, polynucleotides are represented in amounts no more than 1.7 times the mean representation of the polynucleotide library. In some cases, at least 80% of polynucleotides are represented in amounts no more than 1.5 times the mean representation of the polynucleotide library. In some cases, at least 80% of polynucleotides are represented in amounts no more than 2 times the mean representation of the polynucleotide library. In some cases, at least 80% of polynucleotides are represented in amounts no more than 1.7 times the mean representation of the polynucleotide library. In some examples, at least 80% of polynucleotides are displayed in amounts no more than approximately twice the average displayed amount in the polynucleotide library. In some examples, at least 90% of polynucleotides are displayed in amounts no more than approximately 1.5 times the average displayed amount in the polynucleotide library. In some examples, at least 90% of polynucleotides are displayed in amounts no more than approximately twice the average displayed amount in the polynucleotide library. In some examples, at least 80% of polynucleotides are displayed in amounts no more than approximately 1.7 times the average displayed amount in the polynucleotide library. In some examples, at least 90% of polynucleotides are displayed in amounts no more than approximately twice the average displayed amount in the polynucleotide library. In some examples, at least 95% of polynucleotides are displayed in amounts no more than approximately 1.5 times the average displayed amount in the polynucleotide library.In some examples, at least 95% of polynucleotides are displayed in amounts no more than approximately twice the average display in the polynucleotide library. In some examples, at least 95% of polynucleotides are displayed in amounts no more than approximately 1.7 times the average display in the polynucleotide library. In some examples, at least 95% of polynucleotides are displayed in amounts no more than approximately twice the average display in the polynucleotide library. In some examples, the polynucleotide library contains at least several polynucleotides, each containing an overlapping region with another polynucleotide in the library. In some examples, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% of the polynucleotides each contain an overlapping region with another polynucleotide in the library. In some examples, approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or approximately 90% of the polynucleotides each contain an overlapping region with another polynucleotide in the library. In some examples, 10%–90%, 10–80%, 10–75%, 25%–50%, 25–90%, 50–90%, 1–35%, or 80–99% of the polynucleotides contain overlapping regions with other polynucleotides in the library, respectively. In some examples, at least some amounts of polynucleotides in the library are 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library. In some examples, at least 1% of the polynucleotides in the library are 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library. In some cases, the amount of at least 2% of polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library.In some cases, the amount of polynucleotides in the library of at least 5% is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library. In some cases, the amount of polynucleotides in the library of 5% or less is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library. In some cases, the amount of polynucleotides in the library of 10% or less is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library. In some cases, the amount of at least 1% to 10% of polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average display of the polynucleotide library. In some cases, the amount of at least 1% to 20% of polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average display of the polynucleotide library. In some cases, the relative amount of the polynucleotide library is adjusted based on high or low GC content.

[0064] A polynucleotide library for identifying variant sequences can collectively target a desired number of bases (bait territories). In some examples, a polynucleotide library contains bait territories of at least 5 million, 10 million, 15 million, 20 million, 25 million, 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million bases, or at least 100 million bases. In some examples, a polynucleotide library contains bait territories of approximately 5 million, 10 million, 15 million, 20 million, 25 million, 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million bases, or at least about 100 million bases. In some examples, polynucleotide libraries contain bait territories of 5 million, 10 million, 15 million, 20 million, 25 million, 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million bases or less, or 100 million bases or less.

[0065] Unique molecular identifier

[0066] This specification describes adapters that include a unique molecular identifier (UMI). In some examples, the adapter includes structure 1000 in Figure 21. In some examples, the adapter includes a universal adapter. In some examples, the adapter includes a Y-annealing region (annealed to form a yoke), one or more Y-step non-annealing regions, a first index region 1001a, a second index region 1001b, a first UMI (index) region 1002a, a second UMI (index) region 1002b, and one or more regions outside the index. In some examples, the adapter 1000 is ligated 1004 to a sample polynucleotide 1003 to form an adapter-ligated polynucleotide 1005. After denaturation 1005 (Figure 21A), top 1007a and bottom 1007b chain ligation products are formed. In some examples, each chain is labeled with a different UMI. Following amplification 1009 with forward primer 1008a and backward primer 1008b, PCR products of the top strand 1010a and bottom strand 1010b are generated. In some examples, the adapter ligation polynucleotides generated using the universal adapter are further amplified using barcoded primers. In some examples, the adapter described herein includes an "inline" UMI in which at least one of the 5' or 3' UMIs is not complementary to the other corresponding strand of the adapter (1001a and 1001b are not complementary). In some examples, the adapter described herein includes a "duplex" UMI in which at least one of the 5' or 3' UMIs is complementary to the other corresponding strand of the adapter (1001a and 1001b are complementary).

[0067] Adapter-ligated libraries containing unique molecular identifiers can be used to distinguish “true” mutations from polynucleotide sample libraries from artifacts generated during the preparation of the sequencing library (e.g., PCR errors, sequencing errors, or other incorrect base calls). In some examples, the workflow shown in Figure 22 is used to analyze a library of adapter-ligated sample polynucleotides 1101. Each adapter-ligated sample polynucleotide 1101 contains two distinct UMI1101b represented by letters (A-F; for brevity, six combinations of barcodes are shown) and is coupled to sample polynucleotide 1101c. After sequencing 1106, forward and reverse read pairs 1102 from sequencing are sorted into read pair group 1102a. Potential PCR base errors are designated as “*” and true polymorphisms as “+”. The read pairs 1103 are then grouped by barcode and barcode position (1107). Next, single-stranded consensus sequences 1104 are generated from each group of barcode-grouped read pairs (1108). Errors from DC and FE are identified, but errors in AB remain. Finally, duplex consensus sequences 1105 are generated by comparing each set of single-stranded consensus sequences (1109). Errors in AB can be identified, and true mutations EF can be confirmed. In some cases, errors include substitutions, deletions, or insertions. In some cases, errors are present in the sample polynucleotide portion of the adapter-ligated polynucleotide. In some cases, errors are present in barcodes configured to identify the sample origin (e.g., an index) or to uniquely identify the sample polynucleotide. In some cases, errors are present in the UMI. In some cases, errors are present in the sample index. The compositions and methods described herein are used in some cases to identify such errors.

[0068] In this specification, a set of UMIs is described, which has certain properties. In some examples, a UMI set contains multiple distinct polynucleotides having unique sequences. In some examples, a UMI set has 8, 12, 16, 20, 24, 30, 32, 36, 39, 48, or 64 unique sequences. In some examples, the sequences of a UMI set differ by a Hamming distance of 1, 2, 3, 4 or less, or 5 or less. In some examples, the sequences of a UMI set differ by at least 1, 2, 3, 4, or 5 Hamming distances. In some examples, the sequences of a UMI set differ by at least 2 Hamming distances. In some examples, the sequences of a UMI set differ by at least 1 Hamming distance.

[0069] UMI can be of any length depending on the desired application. In some examples, UMI is 15, 12, 10, 8, 7, 6, 5, 4 bases or less, or 3 bases or less. In some examples, UMI is approximately 15, 12, 10, 8, 7, 6, 5, 4 bases, or approximately 3 bases. In some examples, UMI is approximately 3–12, 3–10, 3–8, 4–12, 4–10, 4–8, 6–12 bases, or 8–12 bases. A set of UMI may contain one or more lengths. In some examples, 10, 20, 25, 30, 40, 50, 60, or 70 percent of the UMI in a set is the first length, and 90, 80, 75, 70, 60, 50, 40, or 30 percent is the second length. In some examples, the first length is 3–5 bases, and the second length is 3–5 bases. In some examples, UMIs contain 5 or 6 base pairs in length.

[0070] After adding a UMI-containing adapter to a sample polynucleotide, at least a portion of the sample polynucleotide can be uniquely labeled. In some examples, at least 30%, 50%, 75%, 80%, 90%, 95%, or at least 98% of the sample polynucleotide is ligated to the UMI-containing adapter. In some examples, at least 1%, 2%, 5%, 10%, 15%, 20%, 30%, 50%, 75%, 80%, 90%, 95%, or at least 98% of the sample polynucleotide is labeled with a unique UMI sequence. In some examples, 1%, 2%, 5%, 10%, 15%, 20%, 30%, 50%, 75%, 80%, 90%, 95%, or less than 98% of the sample polynucleotide is labeled with a unique UMI sequence. In some cases, at least 1%, 2%, 5%, 10%, 15%, 20%, 30%, 50%, 75%, 80%, 90%, 95%, or at least 98% of the sample polynucleotides are uniquely identifiable after labeling with UMI.

[0071] The UMIs described herein include, in some examples, one or more sequences from among AAGGA, ACAAC, ATACG, CACTG, CATGA, CGATA, CGTGT, GCCAT, GCTGT, GTCAC, GTCGT, TACGA, TCCTA, TCGTG, TTGCG, TTGGC, AACAC, AATGC, ACTAG, AGCAT, AGTAC, ATCTC, CAGAC, CAGTA, CGAAT, CGGTT, CTTGG, GCATA, GCTAA, GTGAG, GTGTC, and TGTGC. UMIs described herein include, in some examples, two or more sequences from among AAGGA, ACAAC, ATACG, CACTG, CATGA, CGATA, CGTGT, GCCAT, GCTGT, GTCAC, GTCGT, TACGA, TCCTA, TCGTG, TTGCG, TTGGC, AACAC, AATGC, ACTAG, AGCAT, AGTAC, ATCTC, CAGAC, CAGTA, CGAAT, CGGTT, CTTGG, GCATA, GCTAA, GTGAG, GTGTC, and TGTGC. UMIs described herein include, in some examples, five or more sequences from among AAGGA, ACAAC, ATACG, CACTG, CATGA, CGATA, CGTGT, GCCAT, GCTGT, GTCAC, GTCGT, TACGA, TCCTA, TCGTG, TTGCG, TTGGC, AACAC, AATGC, ACTAG, AGCAT, AGTAC, ATCTC, CAGAC, CAGTA, CGAAT, CGGTT, CTTGG, GCATA, GCTAA, GTGAG, GTGTC, and TGTGC. The UMIs described herein include, in some examples, 10 or more sequences from among AAGGA, ACAAC, ATACG, CACTG, CATGA, CGATA, CGTGT, GCCAT, GCTGT, GTCAC, GTCGT, TACGA, TCCTA, TCGTG, TTGCG, TTGGC, AACAC, AATGC, ACTAG, AGCAT, AGTAC, ATCTC, CAGAC, CAGTA, CGAAT, CGGTT, CTTGG, GCATA, GCTAA, GTGAG, GTGTC, and TGTGC.

[0072] UMI can be displayed as a pre-selected percentage within the UMI library. In some examples, at least 90% of UMI exists in fractions of 1-5%. In some examples, at least 90% of UMI exists in fractions of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 7%, or 8%. In some examples, at least 90% of UMI exists in fractions of 0.5-8%, 1-7%, 1.5-7%, 2-7%, 2.5-6%, 3-8%, 3-6%, 1-5%, 0.5-5.5%, 1-4%, 1-6%, or 1-8%.

[0073] Any amount of sample polynucleotide (e.g., input DNA or other nucleic acid) can be ligated to the adapter described herein. In some examples, the amount of sample polynucleotide is about 1, 5, 8, 10, 15, 20, 25, 30, 50, 75 ng, or about 100 ng. In some examples, the amount of sample polynucleotide is less than or equal to 1, 5, 8, 10, 15, 20, 25, 30, 50, 75 ng, or less than or equal to 100 ng. In some examples, the amount of sample polynucleotide is at least 1, 5, 8, 10, 15, 20, 25, 30, 50, 75 ng, or at least 100 ng. In some examples, the amount of sample polynucleotide is 1-10 ng, 1-100 ng, 3-10 ng, 5-100 ng, 5-75 ng, 5-50 ng, 10-100 ng, 10-50 ng, 25-100 ng, or 25-75 ng.

[0074] This specification provides methods for generating adapters containing UMIs. A first method of adapter synthesis involves synthesizing a top chain and a complementary bottom chain of an adapter containing at least one UMI. After annealing the top and bottom adapter chains, an adapter containing the structure of adapter 1000 is formed (Figure 21C). A second method of adapter synthesis involves synthesizing the top chain without UMIs and the bottom chain containing a complementary region and UMIs (Figure 21D). After annealing, complementary UMIs are generated on the top chain using PCR, and a T is added to the 3' end of the top chain by terminal transferase to generate adapter 1000. A third method of synthesis involves synthesizing a top chain without UMIs and a bottom chain containing UMIs, a restriction site, and a 5' overhang (Figure 21E). After annealing, the top chain is extended by PCR, and parts of the 3' top chain and 5' bottom chain are cleaved using a restriction endonuclease to generate adapter 1000. In the fourth method of adapter synthesis, two complementary chains (3' top chain, 5' bottom chain), each containing a UMI, restriction site, and overhang, are synthesized, annealed, and cleaved with restriction enzymes to produce adapter 1000. Multiple UMIs may be present for each adapter. In some examples, an adapter contains 1, 2, 3, 4, 5, or more UMIs. In some examples, an adapter contains a first UMI and a second UMI. In some examples, the first UMI and the second UMI are complementary. In some examples, an adapter contains a first UMI and a second UMI. In some examples, the first UMI and the second UMI are not complementary. In some examples, adapters are combined into an adapter library. In some examples, adapters in the library contain UMIs. In some examples, adapters in the library contain unique combinations of the first UMI and the second UMI.

[0075] Universal adapter

[0076] Universal adapters are provided herein. In some examples, the universal adapter includes one or more unique molecular identifiers. In some examples, the universal adapters disclosed herein may include a universal polynucleotide adapter comprising a first strand and a second strand. In some examples, the first strand includes a first primer-binding region, a first non-complementary region, and a first yoke region. In some examples, the second strand includes a second primer-binding region, a second non-complementary region, and a second yoke region. In some examples, the primer-binding region enables PCR amplification of the polynucleotide adapter. In some examples, the primer-binding region enables PCR amplification of the polynucleotide adapter and the simultaneous addition of one or more barcodes to the polynucleotide adapter. In some examples, the first yoke region is complementary to the second yoke region. In some examples, the first non-complementary region is not complementary to the second non-complementary region. In some examples, the universal adapter is a Y-shaped or fork-shaped adapter. In some examples, one or more yoke regions contain a nucleic acid base analog that increases the Tm between the first and second yoke regions. The primer-binding regions described herein may be in the form of a polynucleotide terminal adapter region. In some examples, the universal adapter contains one index sequence. In some examples, the universal adapter contains one unique molecular identifier. In some examples, the universal adapter is configured for use with a barcoded primer, and after ligation, the barcoded primer is attached via PCR.

[0077] Universal (polynucleotide) adapters can be shortened compared to typical barcoded adapters (e.g., full-length "Y adapters"). For example, universal adapter strands are 20–45 base pairs long. In some examples, universal adapter strands are 25–40 base pairs long. In some examples, universal adapter strands are 30–35 base pairs long. In some examples, universal adapter strands are 50 bases or less, 45 bases or less, 40 bases or less, 35 bases or less, 30 bases or less, or 25 bases or less. In some examples, universal adapter strands are approximately 25, 27, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58 bases, or approximately 60 bases. In some examples, universal adapter strands are approximately 60 base pairs long. In some examples, universal adapter strands are approximately 58 base pairs long. In some examples, the universal adapter strand is approximately 52 base pairs long. In other examples, the universal adapter strand is approximately 33 base pairs long.

[0078] The universal adapter may be modified to facilitate ligation with the sample polynucleotide. For example, the 5' end is phosphorylated. In some examples, the universal adapter contains one or more non-native nucleic acid base bonds, such as a phosphorothioate bond. For example, the universal adapter contains a phosphorothioate between the 3' terminal base and the base adjacent to the 3' terminal base. The sample polynucleotide, in some examples, contains nucleic acids from various sources, such as DNA or RNA of human, bacterial, plant, animal, fungal, or viral origin. Adapter-ligated sample polynucleotides, in some examples, contain a sample polynucleotide (e.g., sample nucleic acid) having an adapter universal adapter ligated to both the 5' and 3' ends of the sample polynucleotide to form an adapter-ligated polynucleotide. Duplex sample polynucleotides contain both a first strand (forward) and a second strand (reverse).

[0079] A universal adapter may contain any number of various nucleic acid bases (DNA, RNA, etc.), nucleic acid base analogs, or non-nucleobase linkers or spacers. For example, the adapter may involve hybridization between two strands of the adapter (T m The adapter contains one or more nucleic acid base analogs or other groups that enhance the adapter. In some examples, the nucleic acid base analog is located in the yoke region of the adapter. Examples of nucleic acid base analogs and other groups include, but are not limited to, locked nucleic acid (LNA), bicyclic nucleic acid (BNA), C5-modified pyrimidine bases, 2'-O-methyl-substituted RNA, peptide nucleic acid (PNA), glycol nucleic acid (GNA), threose nucleic acid (TNA), xenon nucleic acid (XNA), morpholino-modified bases, minor globe binder (MGB), spermine, G-clamps, or anthraquinone (Uaq) caps.

[0080] The universal adapter is for desired hybridization T m Depending on the context, the adapter may contain any number of nucleic acid base analogs (such as LNA or BNA). For example, the adapter contains 1 to 20 nucleic acid base analogs. In some examples, the adapter contains 1 to 8 nucleic acid base analogs. In some examples, the adapter contains at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or at least 12 nucleic acid base analogs. In some examples, the adapter contains approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or approximately 16 nucleic acid base analogs. In some examples, the number of nucleic acid base analogs is expressed as a percentage of the total bases in the adapter. For example, the adapter contains at least 1%, 2%, 5%, 10%, 12%, 18%, 24%, 30%, or more than 30% nucleic acid base analogs. In some examples, the adapters described herein (e.g., the universal adapter) contain methylated nucleic acid bases such as methylated cytosine.

[0081] barcode

[0082] Polynucleotide primers may contain a defined sequence, such as a barcode (or index). The adapter contains one or more barcodes in some examples. In some examples, the adapter contains at least one index barcode and at least one unique molecular identifier barcode. The barcodes can be attached to a universal adapter, for example, using PCR and barcoding primers, to generate a barcoding adapter ligation sample polynucleotide. Primer binding sites, such as universal primer binding sites, facilitate the simultaneous amplification of all members or subpopulations of members of a barcoding primer library. In some examples, the primer binding site includes a region that binds to a flow cell or other solid support during next-generation sequencing. In some examples, the barcoding primer contains a P5 sequence with the nucleic acid sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 52) or a P7 sequence with the nucleic acid sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 53). In some examples, the primer binding site is configured to bind to the universal adapter sequence to facilitate the amplification and generation of the barcoding adapter. In some examples, the barcoding primer is 60 bases or less in length. In some examples, the barcoding primer is 55 bases or less in length. In some examples, the barcoding primer is 50 to 60 bases in length. In some examples, the barcoding primer is approximately 60 bases in length. In some examples, the barcodes described herein include methylated nucleic acid bases such as methylated cytosine.

[0083] The number of unique barcodes available in a barcode set (a collection or combination of unique barcodes configured to be used together to uniquely define a sample) may depend on the length of the barcodes. In some examples, the Hamming distance is defined by the number of base differences between any two barcodes. In some cases, the Levenshtein distance is defined by the number of changes (insertions, substitutions, or deletions) required to transform one barcode into another. In some examples, the barcode sets described herein include at least 2, 3, 4, 5, 6, 7, or at least 8 Levenshtein distances. In some examples, the barcode sets described herein include at least 2, 3, 4, 5, 6, 7, or at least 8 Hamming distances.

[0084] Barcodes may be misassociated with samples other than those to which they are assigned. In some cases, inaccurate barcodes result from PCR errors (e.g., substitutions) during library amplification. In some cases, entire barcodes are "hopped" or transferred from one sample polynucleotide to another. Such transfers in some cases are due to cross-contamination of free adapters or primers during the library generation workflow. In some cases, a set of barcodes is selected to minimize "barcode hopping." In some cases, barcode hopping (for a single barcode) for the barcode sets described herein is 7%, 5%, 4%, 3%, 2%, 1%, less than or equal to 0.5%, or less than or equal to 0.1%. In some cases, barcode hopping (for a single barcode) for the barcode sets described herein is 0.1–6%, 0.1–5%, 0.2–5%, 0.5–5%, 1–7%, 1–5%, or 0.5–7%. In some examples, the barcode hopping (for two barcodes) of the barcode sets described herein is 0.7%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.05% or less, or 0.1% or less. In some examples, the barcode hopping (for two barcodes) of the barcode sets described herein is 0.01–0.6%, 0.01–0.5%, 0.02–0.5%, 0.05–0.5%, 0.1–0.7%, 0.1–0.5%, or 0.05–0.7%.

[0085] Barcoding primers contain one or more barcodes. In some examples, the barcodes are added to a universal adapter by a PCR reaction. A barcode is a nucleic acid sequence that allows for the identification of several features of the polynucleotide to which the barcode is associated. In some examples, the barcode includes an index sequence. In some examples, the index sequence allows for the identification of a specific source of the sample or nucleic acid being sequenced. In some examples, a barcode or combination of barcodes identifies a particular patient. In some examples, a barcode or combination of barcodes identifies a specific sample from a particular patient among other samples from the same patient. After sequencing, the barcode (or barcode region) provides an index for identifying features associated with the coding region or sample source. Barcodes can be designed with lengths suitable for enabling a sufficient degree of identification, for example, at least approximately 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 bases, or longer. Multiple barcodes, such as approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, or longer barcodes, may be used on the same molecule and may be separated by non-barcode sequences of any choice. In some examples, the barcodes are placed on the 5' and 3' ends of the sample polynucleotide. In some examples, each barcode in a group of barcodes differs from all other barcodes at at least three base positions, such as approximately 3, 4, 5, 6, 7, 8, 9, 10, or more. The use of barcodes enables pooling and simultaneous processing of multiple libraries for downstream applications such as sequencing (multiplexing). In some examples, at least 4, 8, 16, 32, 48, 64, 128, or more than 512 barcoded libraries are used.In some examples, at least 400, 500, 800, 1000, 2000, 5000, 10,000, 12,000, 15,000, 18,000, 20,000, or at least 25,000 barcodes are used. Barcoded primers or adapters may include unique molecular identifiers (UMIs). Such UMIs, in some examples, uniquely tag all nucleic acids in the sample. In some examples, at least 60%, 70%, 80%, 90%, 95%, or more than 95% of the nucleic acids in the sample are tagged with UMIs. In some examples, at least 85%, 90%, 95%, 97%, or at least 99% of the nucleic acids in the sample are tagged with unique barcodes or UMIs. Barcoded primers, in some examples, include an index sequence and one or more UMIs. UMI allows for internal measurement of initial sample concentration or stoichiometry before downstream sample processing (e.g., PCR or enrichment steps) that may introduce bias. In some cases, UMI includes one or more barcode sequences. In some examples, each strand (forward vs. reverse) of the adapter-ligated sample polynucleotide has one or more unique barcodes. Such barcodes are optionally used to uniquely tag each strand of the sample polynucleotide. In some examples, the barcoded primers include an index barcode and a UMI barcode. In some examples, after amplification with at least two barcoded primers, the resulting amplicon includes two index sequences and two UMIs. In some examples, after amplification with at least two barcoded primers, the resulting amplicon includes two index barcodes and one UMI barcode. In some examples, each strand of the universal adapter-sample polynucleotide duplex is tagged with a unique barcode such as a UMI or index barcode.

[0086] The barcoded primers in the library contain regions complementary to the primer-binding region on the universal adapter. For example, the universal adapter binding region is complementary to the primer region of the universal adapter, and the universal adapter binding region is complementary to the primer region of the universal adapter. Such arrangement promotes the extension of the universal adapter during PCR and allows the barcoded primer to adhere. In some cases, the Tm between the primer and the primer-binding region is 40–65°C. In some cases, the Tm between the primer and the primer-binding region is 42–63°C. In some cases, the Tm between the primer and the primer-binding region is 50–60°C. In some cases, the Tm between the primer and the primer-binding region is 53–62°C. In some cases, the Tm between the primer and the primer-binding region is 54–58°C. In some cases, the Tm between the primer and the primer-binding region is 40–57°C. In some cases, the Tm between the primer and the primer-binding region is 40–50°C. In some examples, the Tm between the primer and the primer-binding region is approximately 40, 45, 47, 50, 52, 53, 55, 57, 59, 61, or 62°C.

[0087] Hybridization blockers

[0088] A blocker may contain any number of different nucleic acid bases (DNA, RNA, etc.), nucleic acid base analogs (non-standard), or non-nucleobase linkers or spacers. In some examples, a blocker includes a universal blocker. Such blockers may, in some examples, be described as a “set,” and the set includes two or more blockers configured to prevent unwanted interactions with the same adapter sequence. In some examples, a universal blocker prevents adapter-adapter interactions independently of one or more barcodes present in at least one of the adapters. For example, a blocker may prevent hybridization (T) between the blocker and the adapter. m) contains one or more nucleic acid base analogs or other groups that enhance ). In some examples, the blocker enhances hybridization (T) between the blocker and the adapter. m It contains one or more nucleic acid bases (e.g., "universal" bases) that reduce the hybridization (T) between the blocker and the adapter. In some examples, the blockers described herein include one or more nucleic acid bases (e.g., "universal" bases) that reduce the hybridization (T) between the blocker and the adapter. m ) increases one or more nucleic acid bases, and hybridization (T) between the blocker and adapter. m It contains both one or more nucleic acid bases that reduce )

[0089] This specification describes hybridization blockers comprising one or more regions (e.g., adapters) that enhance binding to a target sequence, and one or more regions (e.g., adapters) that reduce binding to the target sequence. In some examples, each region is tuned for a given desired level of off-bait activity in a targeted enrichment application. In some examples, each region can be modified with one or more types of chemical modifications / parts to increase or decrease the overall affinity of the molecule to the target sequence. In some cases, the melting temperatures of all individual members of the blocker set are maintained above a certain temperature (e.g., by the addition of parts such as LNA and / or BNA). In some examples, a given set of blockers improves off-bait performance independently of index length, independently of the index sequence, and independently of the number of adapter indices present in the hybridization.

[0090] Blockers may include portions that increase and / or decrease affinity for target sequencing, such as adapters. In some examples, such specific regions can be thermodynamically tuned to a particular melting temperature to avoid or increase affinity for specific target sequences. Such combinations of modifications are sometimes designed to help increase the affinity of the blocker molecule to specific and unique adapter sequences and decrease the affinity of the blocker molecule to repeating adapter sequences (e.g., the Y-stem annealing portion of the adapter). In some examples, blockers include portions that reduce the binding of the blocker to the Y-stem region of the adapter. In some examples, blockers include portions that reduce the binding of the blocker to the Y-stem region of the adapter and portions that increase the binding of the blocker to the non-Y-stem region of the adapter.

[0091] Blockers (e.g., universal blockers) and adapters can form several different populations during hybridization. Population "A" may include blockers that bind precisely to the non-indexed regions of the adapter. Population "B" may have regions of the blocker that bind to the "yoke" region of the adapter, but the rest of the blocker does not bind to the adjacency regions of the adapter. Population "C" may have two blockers that unproductively dimerize. Population "D" may have no binding to any other nucleic acids at all. In some cases, as the number of DNA modifications that reduce affinity in the Y-stem annealing region of the blocker increases, populations "A" and "D" become dominant, resulting in either the desired effect or the minimal effect. In some cases, as the number of DNA modifications that reduce affinity in the Y-stem annealing region of the blocker decreases, populations "B" and "C" become dominant, resulting in undesirable effects such as annealing to daisy chaining or other adapters ("B"), or sequestering the blocker so that it cannot function properly ("C").

[0092] In both single-index adapter designs and dual-index adapter designs, the index may be partially or completely covered by a universal blocker extended with DNA modifications specifically designed to cover the adapter index bases. In some examples, such modifications include portions that reduce annealing to the index, such as universal bases. In some examples, the index of a dual-index adapter is partially covered (or overlapped) by one or more blockers. In some examples, the index of a dual-index adapter is completely covered by one or more blockers. In some examples, the index of a single-index adapter is partially covered by one or more blockers. In some examples, the index of a single-index adapter is completely covered by one or more blockers. In some examples, the blocker overlaps the index sequence by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 bases, or more than 20 bases. In some examples, the blocker overlaps the index sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 bases or less, or 25 bases or less. In some examples, the blocker overlaps the index sequence by approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 bases, or approximately 30 bases. In some examples, the blocker overlaps the index sequence by 1-5, 1-3, 2-5, 2-8, 2-10, 3-6, 3-10, 4-10, 4-15, 1-4 bases, or 5-7 bases. In some examples, the region of the blocker overlapping the index sequence contains at least one 2-deoxyinosine or 5-nitroindole nucleic acid base.

[0093] One or two blockers may overlap the index sequence present on the adapter. In some examples, a combination of one or two blockers overlaps the index sequence by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 bases, or more than 20 bases. In some examples, a combination of one or two blockers overlaps the index sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 bases or less, or 20 bases or less. In some examples, a combination of one or two blockers overlaps the index sequence by approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 bases, or approximately 20 bases. In some examples, one or two combined blockers overlap bases 1–5, 1–3, 2–5, 2–8, 2–10, 3–6, 3–10, 4–10, 4–15, 1–4, or 5–7 of the index sequence. In some examples, the region of the blocker overlapping the index sequence contains at least one 2-deoxyinosine or 5-nitroindole nucleic acid base.

[0094] In the first configuration, the length of the adapter index overhang can vary. When designed from one side, the adapter index overhang can be modified to cover 0 to n adapter index bases from either side of the index. This gives the ability to design such adapter blockers for both single-index adapter systems and dual-index adapter systems.

[0095] In the second configuration, the adapter index bases are covered from both sides. When the adapter index bases are covered from both sides, the length of the coverage area of ​​each blocker can be chosen so that a single pair of blockers can interact with a certain range of adapter index lengths while still covering a significant portion of the total number of index bases. As an example, take two blockers designed with a 3 bp overhang to cover the adapter index. In the context of 6 bp, 8 bp, or 10 bp adapter index lengths, these blockers would leave 0 bp, 2 bp, or 4 bp exposed, respectively, during hybridization.

[0096] In the third configuration, modified nucleic acid bases are selected to cover the index adapter bases. Examples of these modifications currently available commercially include degenerate bases (i.e., A, T, C, G mixed bases), 2'-deoxyinosine, and 5-nitroindole.

[0097] In the fourth configuration, a blocker with an adapter index overhang is coupled to either the sense (i.e., "top") or antisense (i.e., "bottom") strand of the next-generation sequencing library.

[0098] In the fifth configuration, the blocker is further extended to cover other polynucleotide sequences (e.g., poly-A tails added in preceding biochemical steps to facilitate ligation or other methods for introducing the defined adapter sequence, unique molecular identifiers for bioinformatics assignment after sequencing) in addition to standard adapter index bases of defined length and composition. These types of sequences can be placed at multiple positions on the adapter, in which case the most widely used case (i.e., unique molecular indexes adjacent to genomic inserts) is presented. Other positions of unique molecular identifiers (e.g., adjacent to adapter index bases) can also be addressed with a similar approach.

[0099] In the sixth configuration, all of the aforementioned configurations are used in various combinations to satisfy the target performance metric for off-bait performance during target enrichment under specified conditions.

[0100] Blockers may contain parts such as nucleonucleotide analogs. Examples of nucleonucleotide analogs and other groups include, but are not limited to, locked nucleic acid (LNA), bicyclic nucleic acid (BNA), C5-modified pyrimidine bases, 2'-O-methyl-substituted RNA, peptide nucleic acid (PNA), glycol nucleic acid (GNA), threose nucleic acid (TNA), inosine, 2'-deoxyinosine, 3-nitropyrrole, 5-nitroindole, xenon nucleic acid (XNA), morpholino-modified bases, minor globe binders (MGB), spermine, G-clamps, or anthraquinone (Uaq) caps. In some examples, the nucleonucleotide analog contains a universal base, and the nucleonucleotide base has a lower Tm for binding to congener nucleonucleotide bases. In some examples, the universal base contains 5-nitroindole or 2'-deoxyinosine. In some examples, the blocker contains a spacer element that connects two polynucleotide chains. In some cases, the blocker contains one or more nucleic acid base analogs. In some cases, such nucleic acid base analogs are the T of the blocker. m It is added to control the desired hybridization T. mDepending on the situation, it may contain any number of nucleobase analogs (such as LNA or BNA). For example, the blocker contains 20 to 40 nucleobase analogs. In some examples, the blocker contains 8 to 16 nucleobase analogs. In some examples, the blocker contains at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or at least 12 nucleobase analogs. In some examples, the blocker contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleobase analogs. In some examples, the number of nucleobase analogs is expressed as a percentage of the total bases in the blocker. For example, the blocker contains at least 1%, 2%, 5%, 10%, 12%, 18%, 24%, 30%, or more than 30% nucleobase analogs. In some examples, for a blocker containing nucleobase analogs, T m is increased in the range of about 2 °C to about 8 °C for each nucleobase analog. In some examples, T mThe temperature rises by at least or about 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 12°C, 14°C, or 16°C for each nucleic acid base analog. In some examples, such blockers are configured to bind to the top or "sense" strand of the adapter. In some examples, the blockers are configured to bind to the bottom or "antisense" strand of the adapter. In some examples, a set of blockers includes sequences configured to bind to both the top and bottom strands of the adapter. In some examples, additional blockers are configured on the complementary, inverse, forward, or reverse complementary strands of the adapter sequence. In some examples, a set of blockers targeting the top strand (binding to the top) or the bottom strand (or both) is designed and tested, followed by optimizations such as replacing the top blocker with a bottom blocker or the bottom blocker with a top blocker. In some examples, the blockers are configured to completely or partially overlap the index or barcode bases on the adapter. A set of blockers, in some examples, includes at least one blocker that overlaps with the adapter index sequence. A set of blockers, in some examples, includes at least one blocker that overlaps with the adapter index sequence and at least one blocker that does not overlap with the adapter sequence. A set of blockers, in some examples, includes at least one blocker that does not overlap with the yoke region sequence. A set of blockers, in some examples, includes at least one blocker that does not overlap with the yoke region sequence and at least one blocker that overlaps with the yoke region sequence. A set of blockers, in some examples, includes 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 blockers.

[0101] The blocker is the adapter size or hybridization T mDepending on the context, the length can be any length. For example, a blocker is 20–50 bases long. In some examples, blockers are 25–45 bases, 30–40 bases, 20–40 bases, or 30–50 bases long. In some examples, blockers are 25–35 bases long. In some examples, blockers are at least 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 bases, or at least 35 bases long. In some examples, blockers are 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 bases or less, or 35 bases or less. In some examples, blockers are approximately 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 bases, or approximately 35 bases long. In some examples, blockers are approximately 50 bases long. A set of blockers targeting an adapter-tagged genome library fragment may, in some examples, include blockers of multiple lengths. Two blockers may, in some cases, be tethered together by a linker. Various linkers are well known in the art and, in some examples, include alkyl groups, polyether groups, amine groups, amide groups, or other chemical groups. In some examples, the linker includes separate linker units that are integrally linked (or bound to the blocker polynucleotide) via a backbone such as a phosphate, thiophosphate, amide, or other backbone. In an exemplary configuration, the linker spans an index region between a first blocker, each targeting the 5' end of the adapter sequence, and a second blocker, each targeting the 3' end of the adapter sequence. In some examples, a capping group is added to the 5' or 3' end of the blocker to prevent downstream amplification. The capping group may vary, including polyethers, polyalcohols, alkanes, or other non-hybridizing groups that hinder amplification. Such groups may, in some examples, be linked via a phosphate, thiophosphate, amide, or other backbone. In some examples, one or more blockers are used. In some examples, at least four non-identical blockers are used.In some examples, the first blocker spans the first 3' end of the adapter sequence, the second blocker spans the first 5' end of the adapter sequence, the third blocker spans the second 3' end of the adapter sequence, and the fourth blocker spans the second 5' end of the adapter sequence. In some examples, the first blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 bases long, or at least 35 bases long. In some examples, the second blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 bases long, or at least 35 bases long. In some examples, the third blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 bases long, or at least 35 bases long. In some examples, the fourth blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 bases long, or at least 35 bases long. In some examples, the first, second, third, or fourth blocker contains a nucleic acid base analog. In some examples, the nucleic acid base analog is LNA.

[0102] The blocker design involves the desired hybridization T to the adapter array. m It may be affected by the T of the blocker. In some cases, the T of the blocker m To increase or decrease the T of the blocker, non-standard nucleic acids (e.g., lock nucleic acids, crosslinked nucleic acids, or other non-standard nucleic acids or analogs) are inserted into the blocker. In some examples, the T of the blocker is inserted into the blocker. m T is a polynucleotide containing non-standard amino acids. m It is calculated using a tool specific to that calculation. In some examples, T m This is calculated using the Exiqon® online prediction tool. In some examples, the T of the blockers described herein is calculated. m It is calculated in silico. In some examples, the T of the blocker mIt is calculated in silico and correlated with experimental in vitro conditions. Although not bound by theory, the experimentally determined T m This may be further affected by experimental parameters such as salt concentration, temperature, presence of additives, or other factors. In some examples, the T described herein m T was calculated in silico. m This is used to design or optimize blocker performance. In some examples, T m The value is predicted, evaluated, or determined from melting curve analysis experiments. In some examples, the blocker T m The temperature range is 70°C to 99°C. In some examples, the T of the blocker m The temperature range is 75°C to 90°C. In some examples, the T of the blocker m It is at least 85°C. In some examples, the T of the blocker m The temperature is at least 70, 72, 75, 77, 80, 82, 85, 88, 90, or 92°C. In some examples, the T of the blocker m These are approximately 70, 72, 75, 77, 80, 82, 85, 88, 90, 92, or 95°C. In some examples, the T of the blocker m The temperature range is 78°C to 90°C. In some examples, the T of the blocker m The temperature range is 79°C to 90°C. In some cases, the T of the blocker m The temperature is 80°C to 90°C. In some examples, the T of the blocker m The temperature range is 81°C to 90°C. In some examples, the T of the blocker m The temperature range is 82°C to 90°C. In some cases, the T of the blocker m The temperature range is 83°C to 90°C. In some cases, the T of the blocker m The temperature range is 84°C to 90°C. In some examples, the average T of a set of blockers is... m The temperature range is 78°C to 90°C. In some examples, the average T of a set of blockers is... m The temperature is 80°C to 90°C. In some examples, the average T of a set of blockers is m The temperature is at least 80°C. In some examples, the average T of a pair of blockers mIt is at least 81°C. In some examples, the average T of a pair of blockers m It is at least 82°C. In some examples, the average T of a pair of blockers m It is at least 83°C. In some examples, the average T of a pair of blockers m The temperature is at least 84°C. In some examples, the average T of a pair of blockers m It is at least 86°C. Blocker's T m In some cases, this is modified as a result of other components described herein, such as the use of fast hybridization buffers and / or hybridization accelerators.

[0103] The molar ratio of blocker to adapter target may affect the off-bait (and subsequently off-target) rate during hybridization. The more efficiently the blocker binds to the target adapter, the fewer blockers are required. In some examples, the blockers described herein achieve sequencing results of less than 20% off-target reads at a molar ratio of less than 20:1 (blocker:target). In some examples, less than 20% off-target reads are achieved at a molar ratio of less than 10:1 (blocker:target). In some examples, less than 20% off-target reads are achieved at a molar ratio of less than 5:1 (blocker:target). In some examples, less than 20% off-target reads are achieved at a molar ratio of less than 2:1 (blocker:target). In some examples, less than 20% off-target reads are achieved at a molar ratio of less than 1.5:1 (blocker:target). In some examples, less than 20% off-target reads are achieved at a molar ratio of less than 1.2:1 (blocker:target). In some cases, off-target reads of less than 20% are achieved with a molar ratio of less than 1.05:1 (blocker:target).

[0104] The universal blocker may be used with a panel library of varying size. In some embodiments, this panel library includes at least or about 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 1.0, 2.0, 4.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0, 30.0, 40.0, 50.0, 60.0 megabases (Mb), or more than 60.0 Mb.

[0105] Blockers as described herein may improve on-target performance. In some embodiments, on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95%. In some embodiments, on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95% across various index designs. In some embodiments, on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95% across various panel sizes.

[0106] De novo synthesis of small polynucleotide populations for amplification reactions

[0107] Methods for synthesizing polynucleotides from a surface, e.g., a plate, are described herein (Figure 15). In some examples, the polynucleotide library includes a sample polynucleotide library. In some examples, the polynucleotides are synthesized and released on clusters of loci for polynucleotide elongation and then subjected to an amplification reaction, e.g., PCR. An exemplary workflow for the synthesis of polynucleotides from clusters is shown in Figure 15. A silicon plate 201 contains several clusters 203. Within each cluster, there are several loci 221. Polynucleotides are de novo synthesized on plate 201 from clusters 203 (207). The polynucleotides are cleaved (211) and removed from the plate (213) to form a population 215 of released polynucleotides. The population 215 of released polynucleotides is then amplified (217) to form a library 219 of amplified polynucleotides.

[0108] This specification provides a method for amplifying polynucleotides synthesized on clusters, which provides enhanced control over polynucleotide expression compared to amplifying polynucleotides across the entire surface of a structure that does not have such a clustered arrangement. In some examples, amplification of polynucleotides synthesized from a surface having a clustered arrangement of loci for polynucleotide elongation is provided to overcome the negative effects on expression resulting from the repeated synthesis of large polynucleotide populations. Exemplary negative effects on expression resulting from the repeated synthesis of large polynucleotide populations include, but are not limited to, amplification biases due to high / low GC content, repeating sequences, trailing adenines, secondary structures, affinity for target sequence binding, or modified nucleotides in the polynucleotide sequence.

[0109] Cluster amplification, in contrast to amplifying polynucleotides across the entire plate without clustering, can result in a tighter distribution around the mean. For example, if 100,000 reads are randomly sampled, an average of 8 reads per sequence will produce a library with a distribution of approximately 1.5 times above the mean. In some cases, single-cluster amplification can result in up to approximately 1.5 times, 1.6 times, 1.7 times, 1.8 times, 1.9 times, or 2.0 times above the mean. In other cases, single-cluster amplification can result in at least approximately 1.0 times, 1.2 times, 1.3 times, 1.5 times, 1.6 times, 1.7 times, 1.8 times, 1.9 times, or 2.0 times above the mean.

[0110] Compared to amplification across the entire plate, the cluster amplification methods described herein may result in polynucleotide libraries requiring less sequencing for equivalent sequence representation. In some cases, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% less sequencing is required. In some cases, up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 60%, up to 70%, up to 80%, up to 90%, or up to 95% less sequencing is required. Compared to amplification across the entire plate, up to 30% less sequencing may be required after cluster amplification. In some cases, the sequencing of polynucleotides is validated by high-throughput sequencing such as next-generation sequencing. Sequencing of a sequencing library may be performed using any suitable sequencing technique, including, but not limited to, single-molecule real-time (SMRT) sequencing, Polony sequencing, ligation sequencing, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electron sequencing, pyrosequencing, Maxum-Gilbert sequencing, chain termination reaction (e.g., Sanger) sequencing, +S sequencing, or synthesis sequencing. The number of times a single nucleotide or polynucleotide is identified or "read" is defined as sequencing depth or read depth. In some cases, read depth is referred to as fold coverage, e.g., 55x (or 55×) coverage, and optionally, the percentage of bases is stated.

[0111] In some cases, amplification from clustered arrangements results in less dropout or sequences that are not detected after sequencing of the amplified product, compared to amplification across the entire plate. Dropout can be AT and / or GC. In some cases, the number of dropouts is up to about 1%, 2%, 3%, 4%, or 5% of the polynucleotide population. In some cases, the number of dropouts is zero.

[0112] The clusters described herein comprise a collection of discrete, non-overlapping loci for polynucleotide synthesis. A cluster may contain approximately 50–1000, 75–900, 100–800, 125–700, 150–600, 200–500, or 300–400 loci. In some examples, each cluster contains 121 loci. In some examples, each cluster contains approximately 50–500, approximately 50–200, or approximately 100–150 loci. In some examples, each cluster contains at least approximately 50, 100, 150, 200, 500, 1000, or more loci. In some examples, a single plate contains 100, 500, 10000, 20000, 30000, 50000, 100000, 500000, 700000, 1000000, or more loci. Loci can be spots, wells, microwells, channels, or posts. In some examples, each cluster has at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, or more redundancy of other features supporting the elongation of polynucleotides having identical sequences.

[0113] Generation of polynucleotide libraries with controlled sequence content stoichiometry

[0114] In some cases, polynucleotide libraries (such as sample polynucleotide sets for variant detection) are synthesized with a specified distribution of desired polynucleotide sequences. In some cases, tailoring the polynucleotide library to enrich specific desired sequences results in improved downstream application outcomes.

[0115] One or more specific sequences may be selected based on their evaluation in downstream applications. In some examples, the evaluation is binding affinity to a target sequence for amplification, enrichment, or detection, stability, melting temperature, biological activity, ability to assemble into larger fragments, or other properties of polynucleotides. In some examples, the evaluation is empirical or predicted from prior experiments and / or computer algorithms. Illustrative applications include augmenting sequences in a probe library corresponding to regions of genomic targets with read depths smaller than the average read depth.

[0116] The selected sequences in a polynucleotide library may represent at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or may exceed 95%. In some examples, the selected sequences in a polynucleotide library represent up to 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or up to 100% of the sequence. In some cases, the selected sequences represent approximately 5–95%, 10–90%, 30–80%, 40–75%, or 50–70% of the sequence.

[0117] Polynucleotide libraries can be tuned to the frequency of each selected sequence. In some cases, a polynucleotide library is favored to a larger number of selected sequences. For example, a library might be designed where the polynucleotide frequency increase for selected sequences ranges from approximately 40% to approximately 90%. In some cases, a polynucleotide library might contain a smaller number of selected sequences. For example, a library might be designed so that the polynucleotide frequency increase for selected sequences ranges from approximately 10% to approximately 60%. Libraries can be designed to favor higher and lower frequencies of selected sequences. In some cases, a library might prefer a uniform sequence representation. For example, the polynucleotide frequencies are uniform in the range of approximately 10% to approximately 90% with respect to the frequency of selected sequences. In some cases, a library might contain polynucleotides with selected sequence frequencies ranging from approximately 10% to approximately 95% of the sequences.

[0118] The generation of polynucleotide libraries with specifically selected sequence frequencies is sometimes achieved by combining at least two polynucleotide libraries having different selected sequence frequency content. In some examples, at least 2, 3, 4, 5, 6, 7, 10, or more than 10 polynucleotide libraries are combined to generate a population of polynucleotides with a specified selected sequence frequency. In some cases, 2 or fewer, 3 or fewer, 4 or fewer, 5 or fewer, 6 or fewer, 7 or fewer, or 10 or fewer polynucleotide libraries are combined to generate a population of non-identical polynucleotides with a specified selected sequence frequency.

[0119] In some examples, the selected sequence frequencies are regulated by synthesizing fewer or more polynucleotides per cluster. For example, at least approximately 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than 1000 non-identical polynucleotides are synthesized on a single cluster. In some cases, approximately 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or fewer non-identical polynucleotides are synthesized on a single cluster. In some examples, 50 to 500 non-identical polynucleotides are synthesized on a single cluster. In some examples, 100 to 200 non-identical polynucleotides are synthesized on a single cluster. In some examples, approximately 100, approximately 120, approximately 125, approximately 130, approximately 150, approximately 175, or approximately 200 non-identical polynucleotides are synthesized on a single cluster.

[0120] In some cases, the selected sequence frequencies are regulated by synthesizing non-identical polynucleotides of varying lengths. For example, the length of each non-identical polynucleotide may be at least or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500, 2000 nucleotides, or more. The length of the synthesized non-identical polynucleotides may be at most or at most about 2000, 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 nucleotides, or less. The lengths of each non-identical polynucleotide can range from 10–2000, 10–500, 9–400, 11–300, 12–200, 13–150, 14–100, 15–50, 16–45, 17–40, 18–35, and 19–25.

[0121] Use of polynucleotide libraries as standards or for detection

[0122] This specification provides methods for using polynucleotide libraries to improve the sensitivity and accuracy of nucleic acid variant detection. In some examples, the methods include preparing nucleic acid samples useful for determining the detection limits of genomic variants. In some examples, the methods include one or more of the following steps: providing a polynucleotide library described herein (e.g., a reference standard); obtaining at least one sample from a patient suspected of having a disease or illness; detecting the presence or absence of one or more variant sequences in the library; and detecting the presence or absence of one or more variant sequences in at least one sample. In some examples, the detection step includes sequencing. In some examples, the detection step includes next-generation sequencing (NGS). In some examples, the sequencing includes synthetic sequencing, nanopore sequencing, SMRT sequencing, or other sequencing methods described herein. In some examples, the detection step includes ddPCR or specific hybridization to an array.

[0123] Methods for using a polynucleotide library to detect variant sequences in a sample are also provided herein. The polynucleotides in the library themselves may contain variants. A polynucleotide library may be used to detect variant sequences having low variant allele frequencies. At least one variant is present in the sample at a frequency of about 0.001% to 0.1%. In some examples, a polynucleotide library may be used to detect MRD. In some examples, the method includes the step of providing at least one variant sequence. The variant sequence may be associated with MRD. In some examples, a polynucleotide library is used to detect minimal residual disease (MRD) in a sample. In some examples, the method includes the step of providing a library provided herein. The polynucleotides in the library may contain at least one variant for detection (e.g., a variant associated with MRD). In some examples, the method includes the step of contacting the library with a sample. In some examples, the method includes the step of detecting the presence or absence of one or more variant sequences associated with a disease or illness, such as MRD, in the sample. In some cases, the recall of one or more variant sequences is at least 5% greater than that of multiple polynucleotides that do not contain one or more variant sequences. In some cases, the recall of one or more variant sequences is 5% to 10% greater than that of multiple polynucleotides that do not contain one or more variant sequences.

[0124] In some examples, the method further includes the step of ligating a sequencing adapter to at least some polynucleotides in the test sample, library, or both. In some examples, the method further includes the step of amplifying at least some polynucleotides in the sample, library, or both.

[0125] The method may include the step of obtaining a sample from an individual. In some cases, the individual has been previously treated, is currently being treated, or has had a clinical diagnosis of cancer. In some cases, the sample may include a liquid biopsy. In some cases, the sample may contain circulating tumor DNA (ctDNA). In some cases, the sample is obtained from blood. In some cases, the sample is a biological sample obtained from the kidney, lung, breast, CRC, or melanoma. In some cases, the sample is substantially cell-free. In some cases, the polynucleotides may contain variant sequences corresponding to somatic variants found in diseases such as cancer, e.g., somatic variants found in breast, lung, CRC, melanoma, or renal cell carcinoma.

[0126] The sample (test sample) can be obtained from any source. In some cases, the source is human. In some cases, the source is a human (or patient) suspected of having the disease or illness. In some cases, the test sample includes a liquid biopsy. In some cases, the test sample includes circulating tumor DNA (ctDNA). In some cases, the test sample includes circulating tumor DNA (ctDNA). In some cases, the test sample is obtained from blood. In some cases, the test sample is substantially cell-free. In some cases, multiple test samples are analyzed sequentially or in parallel. In some cases, at least 1, 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, 1000, or more than 2000 test samples are analyzed. In some cases, the method further includes the detection of minimal residual disease (MRD). In some cases, the patient is suspected of having the disease or illness. In some cases, the disease or illness is a proliferative disorder. In some cases, the disease or illness is cancer. In some cases, the patient has previously undergone cancer treatment, is currently undergoing cancer treatment, or has received a clinical diagnosis of cancer. In some cases, the method further includes ligating a sequencing adapter to at least several polynucleotides in the sample, library, or both. In some cases, the method further includes a step of amplifying at least several polynucleotides in the sample, library, or both. In some cases, if one or more variant sequences are not detected in the library, the results obtained from at least one sample are discarded or re-analyzed.

[0127] kit

[0128] This specification provides kits containing libraries of polynucleotides. In some examples, the kit includes one or more reference standards (controls), the reference standards include a sample polynucleotide set and a background set; instructions for use of the kit contents; and packaging for holding and describing the kit contents. In some examples, the kit includes at least two standards selected from sample polynucleotides having VAFs of 0%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some examples, the kit includes five standards, each having VAFs of 0%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some examples, the kit includes instructions for using the reference standards with one or more sequencing instruments or other instruments configured to measure genomic variants. In some examples, the reference standards are packaged in buffer. In some examples, the reference standards are packaged in tubes. In some cases, the reference standards are not packaged in plasma-like form. In some examples, the reference standards contain 500 ng to 5 micrograms of total DNA.

[0129] This specification provides kits for detecting minimal residual disease (MRD) in a sample. In some examples, the kit includes a library of multiple polynucleotides containing at least one variant associated with minimal residual disease (MRD). In some examples, the kit includes instructions for using the kit. In some examples, the kit includes packaging configured to hold and describe the contents of the kit. In some examples, the kit further includes a second library containing at least one variant associated with minimal residual disease (MRD). In some examples, the second library contains at least one variant sequence at different frequencies compared to the said library. In some examples, the second library contains variant sequences different from those in the said library. In some examples, the kit includes a library having VAFs of 0%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some cases, the kit includes five standards, each with VAFs of 0%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence. In some cases, one or more libraries in the kit are packaged in buffer. In some cases, one or more libraries are packaged in tubes. In some cases, one or more libraries are not packaged in plasma-like form. In some cases, one or more libraries contain 500 ng to 5 micrograms of total DNA.

[0130] Application of next-generation sequencing

[0131] Downstream applications of polynucleotide libraries (such as sample polynucleotide sets or reference standards) may include next-generation sequencing. For example, enrichment of target sequences with a controlled stoichiometric polynucleotide probe library results in more efficient sequencing. The performance of a polynucleotide library for capturing or hybridizing targets may be defined by many different metrics that describe efficiency, accuracy, and precision. For example, Picard metrics include: HS library size (the number of unique molecules in the library corresponding to the target region, calculated from read pairs), average target coverage (the percentage of bases that reach a specific coverage level), coverage depth (the number of reads containing a given nucleotide), fold enrichment (the total sample length / target length multiplied by the number of sequence reads that uniquely map to the target / reads that map to the total sample), off-bait base percentage (percentage of bases that do not correspond to the probe / bait bases), off-target percentage (percentage of bases that do not correspond to the target bases), usable bases on the target, AT or GC dropout rate, fold-80 base penalty (fold over-coverage required to bring 80 percent of non-zero targets to the average coverage level), percentage zero coverage targets, PF reads (number of reads that passed the quality filter), and percentage of selected bases (on-bait bases). This includes variables such as the total number of bases and near-bait bases divided by the total number of aligned bases, the overlap percentage, or other variables consistent with the specification.

[0132] Read depth (sequencing depth, or sampling) represents the total number of times a sequence yields sequenced nucleic acid fragments ("reads"). Theoretical read depth is defined as the expected number of times the same nucleotide is read, assuming reads are perfectly distributed across an ideal genome. Read depth is expressed as a function of coverage rate (or coverage width). For example, 10 million reads for a perfectly distributed 1 million base pair genome theoretically yield a read depth 10 times that of 100% of the sequence. In practice, more reads (higher theoretical read depth, or oversampling) may be needed to obtain the desired read depth for a percentage of the target sequence. Enriching the target sequence using a controlled stoichiometric probe library improves the efficiency of downstream sequencing because fewer total reads are needed to obtain a result with an acceptable number of reads for the desired percentage of the target sequence. For example, in some cases, a theoretical read depth of 55 times the target sequence yields at least 30 times coverage of at least 90% of the sequence. In some cases, a theoretical read depth of 55 times or less the target sequence yields at least 30 times read depth of at least 80% of the sequence. In some cases, a theoretical read depth of 55 times or less the target sequence yields at least 30 times read depth of at least 95% of the sequence. In some cases, a theoretical read depth of 55 times or less the target sequence yields at least 10 times read depth of at least 98% of the sequence. In some cases, a theoretical read depth of 55 times the target sequence yields at least 20 times read depth of at least 98% of the sequence. In some cases, a theoretical read depth of 55 times or less the target sequence yields at least 5 times read depth of at least 98% of the sequence. Increasing probe enrichment during hybridization with the target can increase read depth. In some cases, the probe concentration is increased by at least 1.5 times, 2.0 times, 2.5 times, 3 times, 3.5 times, 4 times, 5 times, or more than 5 times.In some examples, increasing the probe concentration results in an increase of at least 1000% in read depth, or increases of 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 750%, 1000%, or more than 1000%. In some examples, tripling the probe concentration results in a 10000% increase in read depth. In some examples, sequencing is performed to achieve theoretical read depths of at least 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, or at least 1000x. In some examples, sequencing is performed to achieve a theoretical read depth of approximately 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, or approximately 1000x. In some examples, sequencing is performed to achieve a theoretical read depth of 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x or less, or 1000x or less. In some examples, sequencing is performed to achieve an actual read depth of at least 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, or at least 1000x. In some examples, sequencing is performed to achieve an actual read depth of 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x or less, or 1000x or less. In some examples, sequencing is performed to achieve an actual read depth of approximately 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, or approximately 1000x.

[0133] The on-target rate represents the percentage of sequencing reads that match the desired target sequence. In some examples, controlled stoichiometric polynucleotide probe libraries yield on-target rates of at least 30%, or at least 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or at least 90%. Increasing the concentration of the polynucleotide probe during contact with the target nucleic acid increases the on-target rate. In some examples, the probe concentration is increased by at least 1.5 times, 2.0 times, 2.5 times, 3 times, 3.5 times, 4 times, 5 times, or more than 5 times. In some embodiments, increasing the probe concentration increases on-target binding by at least 20%, or by 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, or at least 500%. In some cases, a threefold increase in probe concentration leads to a 20% increase in on-target rates.

[0134] Coverage uniformity is sometimes calculated as read depth as a function of target sequence identity. Higher coverage uniformity results in fewer sequencing reads required to obtain the desired read depth. For example, characteristics of the target sequence can affect read depth, such as high or low GC or AT content, repetitive sequences, trailing adenines, secondary structure, affinity for target sequence binding (for amplification, enrichment, or detection), stability, melting temperature, bioactivity, ability to assemble into larger fragments, sequences containing modified nucleotides or nucleotide analogs, or any other characteristics of polynucleotides. Enriching target sequences with a controlled stoichiometric polynucleotide probe library results in higher coverage uniformity after sequencing. In some examples, 95% of sequences have read depths within approximately 1x the average library read depth, or within approximately 0.05, 0.1, 0.2, 0.5, 0.7, 1, 1.2, 1.5, 1.7, or approximately 2x the average library read depth. In some examples, 80%, 85%, 90%, 95%, 97%, or 99% of the sequences have a read depth that is within 1x the average.

[0135] Enrichment of target nucleic acids using polynucleotide probe libraries

[0136] The probe libraries described herein can be used to enrich target polynucleotides present in a population of sample polynucleotides for various downstream applications. In some examples, the sample is obtained from one or more sources, and a population of sample polynucleotides is isolated. The sample is obtained from biological sources such as saliva, blood, tissue, skin, or a completely synthetic source (as an example). Multiple polynucleotides obtained from the sample are fragmented, end-repaired, and adenylated to form double-stranded sample nucleic acid fragments. In some examples, end repair is achieved by treatment with one or more enzymes such as T4 DNA polymerase, Klenow enzyme, and T4 polynucleotide kinase in appropriate buffer. Nucleotide overhangs to facilitate ligation to the adapter are added in some examples, along with a 3'-5' exo-minus Klenow fragment and dATP.

[0137] Adapters (such as universal adapters) can be ligated to both ends of a sample polynucleotide fragment using a ligase such as T4 ligase to generate a library of adapter-tagged polynucleotide strands, which are then amplified using primers such as universal primers. In some examples, the adapter is a Y-shaped adapter containing one or more primer binding sites, one or more grafting regions, and one or more index (or barcode) regions. In some examples, one or more index regions are present on each strand of the adapter. In some examples, the grafting regions are complementary to the flow cell surface to facilitate next-generation sequencing of the sample library. In some examples, the Y-shaped adapter contains partially complementary sequences. In some examples, the Y-shaped adapter contains a single thymidine overhang that hybridizes to the overhang guadenine of the double-stranded adapter-tagged polynucleotide strand. The Y-shaped adapter may contain modified nucleic acids that are resistant to cleavage. For example, a phosphorothioate backbone is used to attach the overhang thymidine to the 3' end of the adapter. When universal primers are used, library amplification is performed to attach barcoded primers to the adapters. The library of double-stranded adapter-tagged polynucleotide strands is contacted with polynucleotide probes to form hybrid pairs. Such pairs are separated from non-hybridized fragments and isolated from the probes to produce an enriched library. The enriched library can then be sequenced.

[0138] Next, a library of double-stranded sample nucleic acid fragments is denatured in the presence of an adapter blocker. The adapter blocker minimizes off-target hybridization of the probe to the adapter sequence present on the adapter-tagged polynucleotide strand (instead of the target sequence) and / or prevents intermolecular hybridization of the adapter (i.e., "daisy chaining"). Denaturation is carried out at 96°C in some examples, or at approximately 85°C, 87°C, 90°C, 92°C, 95°C, 97°C, 98°C, or 99°C. The polynucleotide-targeted library (probe library) is denatured at 96°C in some examples, or at approximately 85°C, 87°C, 90°C, 92°C, 95°C, 97°C, 98°C, or 99°C in the hybridization solution. The denatured adapter-tagged polynucleotide library and hybridization solution are incubated for an appropriate time and temperature to allow the probes to hybridize with their complementary target sequences. In some examples, a suitable hybridization temperature is about 45–80°C, or at least 45°C, at least 50°C, at least 55°C, at least 60°C, at least 65°C, at least 70°C, at least 75°C, at least 80°C, at least 8°C, or at least 90°C. In some examples, the hybridization temperature is 70°C. In some examples, a suitable hybridization time is 16 hours, or at least 4 hours, at least 6 hours, at least 8 hours, at least 10 hours, at least 12 hours, at least 14 hours, at least 16 hours, at least 18 hours, at least 20 hours, at least 22 hours, or 22 hours or more, or about 12–20 hours. Subsequently, the binding buffer is added to the hybridized adapter-tagged polynucleotide probe, and the hybridized adapter-tagged polynucleotide probe is selectively bound using a solid support containing the capture portion. After washing the solid support with buffer to remove unbound polynucleotides, elution buffer is added to release the concentrated tagged polynucleotide fragment from the solid support.In some cases, the solid support is washed twice, or once, twice, three times, four times, five times, or six times. The enriched library of adapter-tagged polynucleotide fragments is amplified, and the enriched library is sequenced.

[0139] Multiple nucleic acids (i.e., genome sequences) may be obtained from a sample, fragmented, optionally repaired at the ends, and adenylated. Adapters are ligated to both ends of polynucleotide fragments to generate a library of adapter-tagged polynucleotide chains, and the adapter-tagged polynucleotide library is amplified. Subsequently, the adapter-tagged polynucleotide library is denatured at a high temperature, preferably 96°C, in the presence of an adapter blocker. The polynucleotide-targeted library (probe library) is denatured at a high temperature, preferably about 90-99°C, in a hybridization solution and combined with the tagged polynucleotide library denatured at about 45-80°C for about 10-24 hours in a hybridization solution. Then, a binding buffer is added to the hybridized tagged polynucleotide probe, and the hybridized adapter-tagged polynucleotide probe is selectively bound using a solid support containing a capture portion. The solid support is washed with buffer one or more times, preferably about two and about five times, to remove unbound polynucleotides, and then elution buffer is added to release the concentrated adapter-tagged polynucleotide fragments from the solid support. The concentrated library of adapter-tagged polynucleotide fragments is amplified, and then the library is sequenced. Alternative variables such as incubation time, temperature, reaction volume / concentration, number of washes, or other variables consistent with this specification may also be employed in this method.

[0140] In any of the examples, the detection or quantification analysis of oligonucleotides can be achieved by sequencing. Subunits or the entire synthesized oligonucleotide can be detected by complete sequencing of all oligonucleotides by any suitable method known in the art, e.g., Illumina sequencing by synthesis, PacBio nanopore sequencing, or BGI / MGI nanoball sequencing, including the sequencing methods described herein.

[0141] Sequencing can be performed through classical Sanger sequencing methods well known in the art. Sequencing may also be achieved using high-throughput systems, some of which allow for the detection of sequenced nucleotides immediately after or as they are incorporated into a growing strand, i.e., the detection of sequences at the time read or substantially in real time. In some cases, high-throughput sequencing generates at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, or at least 500,000 sequence reads per hour; each read is at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, or at least 150 base pairs per read.

[0142] In some examples, high-throughput sequencing involves using techniques available through Illumina's Genome Analyzer IIX, MiSeq personal sequencer, or HiSeq systems, such as HiSeq 2500, HiSeq 1500, HiSeq 2000, HiSeq 1000, iSeq 100, Mini Seq, MiSeq, NextSeq 550, NextSeq 2000, NextSeq 550, or NovaSeq 6000. These machines use reversible terminator-based sequencing via synthetic chemistry. These machines can generate 6000 Gb or more of reads in 13–44 hours. Smaller systems may be available for runs within 3 days, 2 days, 1 day, or less. Shorter synthetic cycles may be used to minimize the time required to obtain sequencing results.

[0143] In some cases, high-throughput sequencing involves the use of techniques available through the ABI Solid System. This gene analysis platform enables ultra-parallel sequencing of clonely amplified DNA fragments ligated to beads. The sequencing methodology is based on sequential ligation with dye-labeled oligonucleotides.

[0144] Next-generation sequencing may include ion-semiconductor sequencing (e.g., using technology from Life Technologies (Ion Torrent)). Ion-semiconductor sequencing can utilize the fact that ions can be released when nucleotides are incorporated into the DNA chain. To perform ion-semiconductor sequencing, a high-density array of microfabricated wells can be formed. Each well can hold a single DNA template. There may be an ion-sensitive layer beneath the wells, and an ion sensor beneath the ion-sensitive layer. When nucleotides are added to the DNA, H+ ions are released, which can be measured as a change in pH. The H+ ions can be converted into a voltage and recorded by the semiconductor sensor. The array chip may be continuously filled (flooded) with one nucleotide at a time. Scanning, light, or cameras are not required. In some cases, the IONPROTON® Sequencer is used to sequence nucleic acids. In some cases, the IONPGM® Sequencer is used. The Ion Torrent Personal Genome Machine (PGM) can perform 10 million reads in 2 hours.

[0145] In some embodiments, high-throughput sequencing involves the use of techniques available from Helicos BioSciences Corporation (Cambridge, Mass.), such as Single Molecule Sequencing by Synthesis (SMSS) methods. SMSS is unique because it enables sequencing of the entire human genome in up to 24 hours. Ultimately, SMSS is powerful because, unlike MW techniques, it does not require a pre-amplification step before hybridization. In fact, SMSS does not require amplification at all.

[0146] In some embodiments, high-throughput sequencing involves the use of technologies available from 454 Lifesciences, Inc. (Branford, Conn.), such as Pico Titer Plate instruments, which include a fiber optic plate that transmits a chemiluminescent signal generated by the sequencing reaction, recorded by a CCD camera within the instrument. Using this fiber optic, detection of a minimum of 20 million base pairs is possible in 4.5 hours.

[0147] A method for using bead amplification followed by fiber optics detection is described in Marguiles et al., “Genome sequencing in microfabricated high-density picolitre reactors”, Nature, 2005, vol. 437, pages 376-380.

[0148] In some cases, high-throughput sequencing is performed using synthetic sequencing (SBS) that utilizes reversible terminator chemistry, such as the Clonal Single Molecule Array (Solexa, Inc.) or Constans, “Beyond Sanger: toward the $1,000 genome: new technologies promise faster and cheaper whole-genome sequencing”, The Scientist, 2003, vol. 17, issue 13, page 36+. High-throughput sequencing of oligonucleotides can be achieved using any suitable sequencing method known in the art, such as those commercially available from Pacific Biosciences, Complete Genomics, Genia Technologies, Halcyon Molecular, Oxford Nanopore Technologies, etc. Overall, such systems involve sequencing a target oligonucleotide molecule with multiple bases by timely addition of bases via polymerization reactions measured on the oligonucleotide molecule, i.e., tracking the activity of nucleic acid polymerases on the template oligonucleotide molecule being sequenced in real time. Subsequently, the sequence can be estimated by identifying which bases are incorporated into the growing complementary strand of the target oligonucleotide based on the catalytic activity of the nucleic acid polymerase at each step in the base addition sequence. Polymerase on the target oligonucleotide molecular complex is provided at an appropriate position to move along the target oligonucleotide molecule and extend the oligonucleotide primer at the active site. Multiple labeled nucleotide analogs are provided very close to the active site, and each identifiable nucleotide analog is complementary to a different nucleotide in the target oligonucleotide sequence.The growing oligonucleotide chain is extended by adding a nucleotide analog to the oligonucleotide chain at the active site using polymerase, where the added nucleotide analog is complementary to the nucleotide of the target oligonucleotide at the active site. The nucleotide analog added to the oligonucleotide primer as a result of the polymerization step is identified. The steps of providing the labeled nucleotide analog, polymerizing the growing oligonucleotide chain, and identifying the added nucleotide analog are repeated, resulting in further extension of the oligonucleotide chain and measurement of the target oligonucleotide sequence.

[0149] Next-generation sequencing technology includes real-time sequencing (SMRT®) technology from Pacific Biosciences. In SMRT, each of the four DNA bases may be bound to one of four different fluorescent dyes. These dyes can be phosphor-bound. A single DNA polymerase can be immobilized on a single molecule of template single-stranded DNA at the bottom of a zero-mode waveguide (ZMW). The ZMW can be a constraint structure that allows observation of the incorporation of a single nucleotide by DNA polymerase against a background of fluorescent nucleotides that can rapidly diffuse (in microseconds) in and out of the ZMW. Incorporation of a nucleotide into a growing strand may take several milliseconds. During this time, the fluorescent label can be excited to generate a fluorescent signal, and the fluorescent tag can be cleaved. The ZMW can be illuminated from below. Attenuated light from the excitation beam can penetrate the bottom 20-30 nm of each ZMW. A microscope with a detection limit of 20 zeptoliters (10" liters) can be constructed. This small detection volume can improve background noise reduction by a factor of 1000. Detecting the corresponding fluorescence of the dye can indicate which bases have been incorporated. This process can be repeated.

[0150] In some cases, next-generation sequencing is nanopore sequencing. For example, Soni et al., “Progress toward ultrafast DNA sequencing using solid-state nanopores”, Clin Chem., 2007, vol. 53, pages 1996-2001. Nanopores can be tiny holes with a diameter of about 1 nanometer. Immersion of nanopores in a conductive fluid and the application of a potential across them can result in a small current due to the conduction of ions through the nanopores. The amount of current flowing can be sensitive to the size of the nanopore. As DNA molecules pass through a nanopore, each nucleotide on the DNA molecule can block the nanopore to a different degree. Therefore, changes in the current through the nanopore as DNA molecules pass through it can represent reads of the DNA sequence. Nanopore sequencing technology can originate from Oxford Nanopore Technologies, for example, the GridION system. Single nanopores can be inserted into a polymer membrane spanning the top of microwells. Each microwell may have an electrode for individual sensing. Microwells can be assembled into array chips with more than 100,000 microwells per chip (e.g., 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or more than 1,000,000). Instruments (or nodes) can be used to analyze the chips. Data can be analyzed in real time. One or more instruments can be operated simultaneously. Nanopores can be protein nanopores, e.g., protein alpha-hemolysin, heptameric protein pores. Nanopores are made of solid-state nanopores, e.g., synthetic membranes (e.g., SiN). xThese can be nanometer-sized holes formed in (or SiO2). Nanopores can be hybrid pores (e.g., integration of protein pores into solid-state membranes). Nanopores can be equipped with integrated sensors (e.g., tunneling electrode detectors, capacitance detectors, or graphene-based nanogap or edge state detectors) (e.g., Garaj et al., “Graphene as a subnanometre trans-electrode membrane, Nature, 2010, vol. 67, pages (See 190-193). Nanopores can be functionalized to analyze specific types of molecules (e.g., DNA, RNA, or proteins). Nanopore sequencing can include "strand sequencing," in which a complete DNA polymer is passed through a protein nanopore and the DNA can be sequenced in real time as it moves through the protein nanopore. Enzymes can separate the strands of double-stranded DNA and pass the strands through the nanopore. The DNA may have a hairpin at one end, and the system can read both strands. In some cases, nanopore sequencing is "exonuclease sequencing," in which individual nucleotides can be cleaved from the DNA strand by a forward-moving exonuclease, and the nucleotides can pass through the protein nanopore. The nucleotides can transiently bind to molecules within the pore (e.g., cyclodextran). The characteristic discontinuities of the current can be used to identify the bases.

[0151] GENIA's nanopore sequencing technology may be used. Manipulated protein pores can be embedded in a lipid bilayer membrane. Using "active control" techniques, effective nanopore-membrane assembly can be enabled, allowing control of DNA movement through channels. In some cases, the nanopore sequencing technology is from NABsys. Genomic DNA can be fragmented into strands with an average length of approximately 100kb. 100kb fragments can be single-stranded and then hybridized with 6-mer probes. Genomic fragments with probes can pass through nanopores, creating current-vs-time traces. Current traces can provide probe locations on each genomic fragment. Genomic fragments can be aligned to create probe maps for the genome. This process can be performed in parallel with a library of probes. Genomic length probe maps can be generated for each probe. Errors can be corrected using a process called "moving window sequencing by hybridization (mwSBH)". In some examples, the nanopore sequencing technology is from IBM / Roche. Using an electron beam, nanopore-sized apertures can be fabricated in microchips. An electric field can be used to attract or twist DNA through the nanopore. DNA transistor devices in nanopores can include nanometer-sized layers of alternating metals and dielectrics. Separate charges within the DNA backbone can be confined inside the DNA nanopore by the electric field. The DNA sequence can be read by switching the gate voltage on and off.

[0152] Next-generation sequencing may include DNA nanoball sequencing (for example, as performed by Complete Genomics (see, e.g., Drmanac et al., “Human genome sequencing using unchained base reads on self-assembling DNA nanoarrays”, Science, 2010, vol. 327, pages 78-81)). DNA can be isolated, fragmented, and size-selected. For example, DNA can be fragmented to an average length of about 500 bp (e.g., by sonication). Adapters (Adl) can be attached to the ends of the fragments. Adapters can be used to hybridize to anchors for sequencing reactions. DNA with adapters attached to each end can be PCR amplified. Adapter sequences can be modified so that complementary single-stranded ends join to form circular DNA. DNA can be methylated to protect it from cleavage by IIS-type restriction enzymes used in subsequent steps. Adapters (e.g., right adapter) may have restriction recognition sites, which may remain unmethylated. Unmethylated restriction recognition sites in the adapter can be recognized by restriction enzymes (e.g., Acul), and the DNA can be cleaved by Acul 13 bp to the right of the right adapter to form linear double-stranded DNA. The right and left adapters (Ad2) of the second round can be ligated to either end of the linear DNA, and all DNA to which both adapters are bound can be PCR-amplified (e.g., by PCR). The Ad2 sequences can be modified so that they can join together to form circular DNA. The DNA can be methylated, but the restriction enzyme recognition sites may remain unmethylated on the left Ad1 adapter. Restriction enzymes (e.g., Acul) can be applied, and the DNA can be cleaved 13 bp to the left of Ad1 to form linear DNA fragments.The right and left adapters (Ad3) of the third round can be ligated to the right and left sides of linear DNA, and the resulting fragments can be amplified by PCR. The adapters can be modified so that they can bind to each other to form circular DNA. A type III restriction enzyme (e.g., EcoP15) may also be added; EcoP15 can cleave the DNA at 26 bp on the left side of Ad3 and 26 bp on the right side of Ad2. This cleavage removes large segments of DNA, allowing the DNA to be linearized again. The right and left adapters (Ad4) of the fourth round can be ligated to the DNA, and the DNA can be amplified (e.g., by PCR) and modified so that they bind to each other to form the completed circular DNA template.

[0153] Small fragments of DNA can be amplified using rolling circle replication (e.g., using Phi29 DNA polymerase). The four adapter sequences may contain palindrome sequences that can hybridize, and the single strands can fold over themselves to form DNA nanoballs (DNBs™), which may have an average diameter of approximately 200-300 nanometers. The DNA nanoballs can be attached to a microarray (sequencing flow cell) (e.g., by adsorption). The flow cell may be a silicon wafer coated with silicon dioxide, titanium, and hexamethyldisilazane (HMDS), as well as photoresist material. Sequencing can be performed by non-linking sequencing by ligating a fluorescent probe to the DNA. The fluorescence color at the interrogated position can be visualized by a high-resolution camera. The identity of nucleotide sequences between adapter sequences can be determined.

[0154] The polynucleotide population may be concentrated before adapter ligation. In one example, a polynucleotide population is obtained from a sample, fragmented, optionally end-repaired, and denatured at a high temperature, preferably 90-99°C. The polynucleotide-targeted library (probe library) is denatured in a hybridization solution at a high temperature, preferably about 90-99°C, and combined with a tagged polynucleotide library denatured in a hybridization solution at about 45-80°C for about 10-24 hours. Subsequently, a binding buffer is added to the hybridized tagged polynucleotide probe, and the hybridized adapter-tagged polynucleotide probe is selectively bound using a solid support containing the capture portion. The solid support is washed with buffer one or more times, preferably about two and about five times, to remove unbound polynucleotides, and then elution buffer is added to release the concentrated adapter-tagged polynucleotide fragments from the solid support. Subsequently, the enriched polynucleotide fragments are polyadenylated, and adapters are ligated to both ends of the polynucleotide fragments to generate a library of adapter-tagged polynucleotide strands, and the adapter-tagged polynucleotide library is amplified. Then, the adapter-tagged polynucleotide library is sequenced.

[0155] Furthermore, polynucleotide targeting libraries can also be used to filter out undesirable sequences from multiple polynucleotides by hybridizing them to undesirable fragments. For example, multiple polynucleotides are obtained from a sample, fragmented, optionally repaired at the ends, and adenylated. Adapters are ligated to both ends of the polynucleotide fragments to generate a library of adapter-tagged polynucleotide chains, and the adapter-tagged polynucleotide library is amplified. Alternatively, the adenylation and adapter ligation steps are performed after enrichment of the sample polynucleotide. The adapter-tagged polynucleotide library is then denatured at a high temperature, preferably 90-99°C, in the presence of an adapter blocker. A polynucleotide filtering library (probe library) designed to remove undesirable non-target sequences is denatured in a hybridization solution at a high temperature, preferably about 90-99°C, and combined with the tagged polynucleotide library denatured in a hybridization solution at about 45-80°C for about 10-24 hours. Subsequently, a binding buffer is added to the hybridized tagged polynucleotide probe, and the hybridized adapter-tagged polynucleotide probe is selectively bound using a solid support containing the capture moiety. The solid support is washed with buffer one or more times, preferably about once and about five times, to elute the unbound adapter-tagged polynucleotide fragments. The enriched library of unbound adapter-tagged polynucleotide fragments is amplified, and then the amplified library is sequenced.

[0156] Highly parallel de novo nucleic acid synthesis

[0157] This specification describes a platform approach that leverages miniaturization, parallelization, and vertical integration of an end-to-end process from polynucleotide synthesis to gene assembly in nanowells on silicon to create an innovative synthesis platform. The apparatus described herein has the same footprint as a 96-well plate, and the silicon synthesis platform can increase throughput by 100 to 1,000 times compared to conventional synthesis methods, producing up to approximately 1,000,000 polynucleotides in a single highly parallelized run. In some examples, a single silicon plate described herein provides the synthesis of approximately 6,100 non-identical polynucleotides. In some examples, each non-identical polynucleotide is located within a cluster. A cluster may contain 50 to 500 non-identical polynucleotides.

[0158] The methods described herein provide for the synthesis of polynucleotide libraries, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library includes sequences encoding mutations at least one codon, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are generated by a standard translation process. The synthesized specific changes in the nucleic acid sequence can be introduced by incorporating the nucleotide changes into duplicate or blunt-end polynucleotide primers. Alternatively, a population of polynucleotides may collectively encode a long nucleic acid (e.g., a gene) and its variants. In this arrangement, the population of polynucleotides can be hybridized and subjected to standard molecular biology techniques to form a long nucleic acid (e.g., a gene) and its variants. When the long nucleic acid (e.g., a gene) and its variants are expressed in cells, different protein libraries are generated. Similarly, methods for synthesizing variant libraries encoding RNA sequences (e.g., miRNA, shRNA, and mRNA) or DNA sequences (e.g., enhancer, promoter, UTR, and terminator regions) are provided herein. Furthermore, downstream applications of variants selected from libraries synthesized using the methods described herein are also provided herein. Downstream applications include the identification of variant nucleic acids or protein sequences with enhanced biologically relevant functions, such as altered biochemical affinity, enzymatic activity, or cellular activity, and for the purpose of treating or preventing disease conditions.

[0159] substrate

[0160] A substrate comprising multiple clusters is provided herein, where each cluster comprises multiple sacs supporting the binding and synthesis of polynucleotides. As used herein, the term “sac” refers to a distinct structural region supporting a single, predetermined sequence of polynucleotides extending from a surface. In some examples, a sac lies on a two-dimensional surface, e.g., a substantially flat surface. In some examples, a sac refers to a distinct raised or recessed site on a surface, e.g., a well, microwell, channel, or post. In some examples, the surface of a sac contains an actively functionalized substance that binds to at least one nucleotide for polynucleotide synthesis, or preferably, a group of identical nucleotides for the synthesis of a group of polynucleotides. In some examples, polynucleotide refers to a group of polynucleotides encoding the same nucleic acid sequence. In some examples, the surface of the apparatus encompasses one or more surfaces of the substrate.

[0161] Structures are provided herein that may include a surface supporting the synthesis of multiple polynucleotides having different predetermined sequences at addressable positions on a common support. In some examples, the apparatus may include 5,000;10,000;20,000;30,000;50,000;75,000;100,000;200,000;300,000;400,000;500,000;600,000;700,000;800,000;900,000;1,000,000;1,200,000;1 It provides support for the synthesis of non-identical polynucleotides exceeding 400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; and 10,000,000. In some examples, the device has separate arrays of CODE, 2,000; 5,000; 10,000; 20,000; 30,000; 50,000; 75,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1 It provides a support for the synthesis of polynucleotides exceeding 200,000; 1,400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; and 10,000,000 or more. In some examples, at least a portion of the polynucleotides have the same sequence or are configured to be synthesized with the same sequence.

[0162] This specification provides methods and apparatus for producing and growing polynucleotides having a length of approximately 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 bases. In some examples, the length of the polynucleotide formed is approximately 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or 225 base pairs. A polynucleotide can be at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 base pairs long. A polynucleotide can be 10–225 base pairs, 12–100 base pairs, 20–150 base pairs, 20–130 base pairs, or 30–100 base pairs long.

[0163] In some examples, polynucleotides are synthesized at separate loci on the substrate, where each locus supports the synthesis of a population of polynucleotides. In some examples, each locus supports the synthesis of a population of polynucleotides having a different sequence from the population of polynucleotides grown on another locus. In some examples, the loci of the apparatus are located within multiple clusters. In some examples, the apparatus contains at least 10, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 30000, 40000, 50000, or more clusters. In some examples, the device is 2,000; 5,000; 10,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1,100,000; 1,200,000; 1,300,000; 1,400,000; 1,500,000; 1,600,000; 1,700,000; 1,800,000; 1,900,000; 2,0 00,000;300,000;400,000;500,000;600,000;700,000;800,000;900,000;1,000,000;1,200,000;1,400,000;1,600,000;1,800,000;2,000,000;2,500,000;3,000,000;3,500,000;4,000,000;4,500,000;greater than 5,000,000;or containing 10,000,000 or more distinct situations. In some examples, the device contains approximately 10,000 distinct situations. The number of situations within a single cluster varies in different examples. In some examples, each cluster contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 130, 150, 200, 300, 400, 500, 1000, or more seats. In some examples, each cluster contains approximately 50 to 500 seats. In some examples, each cluster contains approximately 100 to 200 seats. In some examples, each cluster contains approximately 100 to 150 seats.In some examples, each cluster contains approximately 109, 121, 130, or 137 loci. In other examples, each cluster contains approximately 19, 20, 61, 64, or more loci.

[0164] The number of distinct polynucleotides synthesized on the instrument may depend on the number of distinct loci available in the substrate. In some examples, the density of loci within the instrument cluster is 1 mm 2 At least or approximately 1 seat per unit, 1 mm 2 At least or about 10 seats per unit, 1 mm 2 At least or approximately 25 seats per 1mm 2 At least or approximately 50 spots per unit, 1 mm 2 At least or approximately 65 seats per 1mm 2 At least or approximately 75 seats per 1mm 2 At least or approximately 100 seats per unit, 1 mm 2 At least or approximately 130 seats per 1mm 2 At least or approximately 150 seats per unit, 1 mm 2 At least or approximately 175 seats per unit, 1 mm 2 At least or approximately 200 seats per 1mm 2 At least or approximately 300 seats per 1mm 2 At least or approximately 400 seats per 1mm 2 At least or approximately 500 seats per unit, 1 mm 2 There are at least or about 1,000 seats per unit, or more. In some cases, the device is 1 mm 2 Approximately 10 to 500 seats per unit, 1 mm 2 Approximately 25 to 400 seats per unit, 1 mm 2 Approximately 50 to 500 seats per unit, 1 mm 2 Approximately 100 to 500 seats per unit, 1 mm 2 Approximately 150 to 500 seats per unit, 1 mm 2 Approximately 10 to 250 seats per unit, 1 mm 2 Approximately 50 to 250 seats per unit, 1 mm 2 Approximately 10 to 200 seats per unit, 1 mm 2Each cluster contains approximately 50 to 200 loci. In some examples, the distance from the centers of two adjacent loci within a cluster is approximately 10 μm to 500 μm, 10 μm to 200 μm, or 10 μm to 100 μm. In some examples, the distance from the centers of two adjacent loci is longer than approximately 10 μm, 20 μm, 30 μm, 40 μm, 50 μm, 60 μm, 70 μm, 80 μm, 90 μm, or 100 μm. In some examples, the distance from the centers of two adjacent loci is approximately 200 μm, 150 μm, 100 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, or less than 10 μm. In some examples, the width of each locus is approximately 0.5 μm, 1 μm, 2 μm, 3 μm, 4 μm, 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, 20 μm, 30 μm, 40 μm, 50 μm, 60 μm, 70 μm, 80 μm, 90 μm, or 100 μm. In some examples, the width of each locus is approximately 0.5 μm to 100 μm, approximately 0.5 μm to 50 μm, approximately 10 μm to 75 μm, or approximately 0.5 μm to 50 μm.

[0165] In some cases, the cluster density within the device is 100 mm 2 At least or approximately 1 cluster per 10mm 2 At least or approximately 1 cluster per 5mm 2 At least or approximately 1 cluster per 4mm 2 At least or approximately 1 cluster per 3mm 2 At least or approximately 1 cluster per 2mm 2 At least or approximately 1 cluster per 1 mm 2 At least or approximately 1 cluster per 1 mm 2 At least or approximately 2 clusters per 1 mm 2 At least or approximately 3 clusters per 1 mm 2 At least or approximately 4 clusters per 1 mm 2 At least or approximately 5 clusters per 1 mm 2 At least or approximately 10 clusters per 1 mm 2There are at least or about 50 clusters per unit, or more. In some examples, the device is 10 mm 2 Approximately 1 cluster per unit ~ 1 mm 2 Each cluster contains approximately 10 clusters. In some examples, the distance from the center of two adjacent clusters is less than approximately 50 μm, less than approximately 100 μm, less than approximately 200 μm, less than approximately 500 μm, less than approximately 1000 μm, less than approximately 2000 μm, or less than approximately 5000 μm. In some examples, the distance from the center of two adjacent clusters is approximately 50 μm to 100 μm, approximately 50 μm to 200 μm, approximately 50 μm to 300 μm, approximately 50 μm to 500 μm, and approximately 100 μm to 2000 μm. In some examples, the distance from the center of two adjacent clusters is approximately 0.05 mm to 50 mm, 0.05 mm to 10 mm, 0.05 mm to 5 mm, 0.05 mm to 4 mm, 0.05 mm to 3 mm, 0.05 mm to 2 mm, 0.1 mm to 10 mm, 0.2 mm to 10 mm, 0.3 mm to 10 mm, 0.4 mm to 10 mm, 0.5 mm to 10 mm, 0.5 mm to 5 mm, or 0.5 mm to 2 mm. In some examples, each cluster has a diameter or width along one dimension of approximately 0.5 to 2 mm, 0.5 to 1 mm, or 1 to 2 mm. In some examples, each cluster has a diameter or width along one dimension of approximately 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mm. In some examples, each cluster has an inner diameter or width along one dimension of approximately 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.15, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mm.

[0166] The device is approximately the size of a standard 96-well plate, for example, about 100 - 200 mm × about 50 - 150 mm. In some examples, the diameter of the device is about 1000 mm or less, about 500 mm or less, about 450 mm or less, about 400 mm or less, about 300 mm or less, about 250 mm or less, about 200 mm or less, about 150 mm or less, about 100 mm or less, or about 50 mm or less. In some examples, the diameter of the device is about 25 mm - 1000 mm, about 25 mm - about 800 mm, about 25 mm - about 600 mm, about 25 mm - about 500 mm, about 25 mm - about 400 mm, about 25 mm - about 300 mm, or about 25 mm - about 200 mm. Non-limiting examples of the size of the device include about 300 mm, about 200 mm, about 150 mm, about 130 mm, about 100 mm, about 76 mm, about 51 mm, about 25 mm. In some examples, the device is at least about 100 mm 2 ; 200 mm 2 ; 500 mm 2 ; 1,000 mm 2 ; 2,000 mm 2 ; 5,000 mm 2 ; 10,000 mm 2 ; 12,000 mm 2 ; 15,000 mm 2 ; 20,000 mm 2 ; 30,000 mm 2 ; 40,000 mm 2 ; 50,000 mm 2, or having a planar surface area of ​​, or greater. In some examples, the thickness of the device is approximately 50 mm to 2000 mm, approximately 50 mm to 1000 mm, approximately 100 mm to 1000 mm, approximately 200 mm to 1000 mm, or approximately 250 mm to 1000 mm. Non-limiting examples of device thickness include 275 mm, 375 mm, 525 mm, 625 mm, 675 mm, 725 mm, 775 mm, and 925 mm. In some examples, the thickness of the device varies with diameter and depends on the composition of the substrate. For example, a device containing a material other than silicon will have a different thickness than a silicon device of the same diameter. The thickness of the device is determined by the mechanical strength of the material used and must be sufficient to support its own weight without cracking during handling. In some examples, the structure includes multiple devices as described herein.

[0167] surface material

[0168] Apparatuses including surfaces are provided herein, where the surfaces are modified to support polynucleotide synthesis at predetermined locations, resulting in low error rates, low dropout rates, high yields, and high oligo-expression. In some examples, the surfaces of the apparatuses for polynucleotide synthesis provided herein are made from a variety of modifiable materials to support de novo polynucleotide synthesis reactions. In some cases, the apparatus is sufficiently conductive, for example, to form a uniform electric field over all or part of the apparatus. Apparatuses described herein may also include flexible materials. Typical flexible materials include, but are not limited to, modified nylon, unmodified nylon, nitrocellulose, and polypropylene. Apparatuses described herein may also include rigid materials. Typical rigid materials include, but are not limited to, glass, quartz glass (fuse silica), silicon, silicon dioxide, silicon nitride, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and mixtures thereof), and metals (e.g., gold, platinum). The apparatus disclosed herein may be made from materials including silicon, polystyrene, agarose, dextran, cellulosic polymers, polyacrylamide, polydimethylsiloxane (PDMS), glass, or any combination thereof. In some cases, the apparatus disclosed herein may be manufactured from the materials listed herein or other suitable combinations of materials known in the art.

[0169] A list of typical tensile strengths of materials described herein is as follows: nylon (70 MPa), nitrocellulose (1.5 MPa), polypropylene (40 MPa), silicon (268 MPa), polystyrene (40 MPa), agarose (1-10 MPa), polyacrylamide (1-10 MPa), and polydimethylsiloxane (PDMS) (3.9-10.8 MPa). Solid supports described herein may have tensile strengths of 1-300, 1-40, 1-10, 1-5, or 3-11 MPa. Solid supports described herein may have tensile strengths of approximately 1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 20, 25, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 270, or higher MPa. In some examples, the apparatus described herein includes a solid support for polynucleotide synthesis, which may take the form of a flexible material that can be stored in a continuous loop or reel, such as a tape or flexible sheet.

[0170] Young's modulus measures a material's resistance to elastic (recoverable) deformation under load. A list of typical Young's moduli for stiffness of materials described herein is as follows: nylon (3 GPa), nitrocellulose (1.5 GPa), polypropylene (2 GPa), silicon (150 GPa), polystyrene (3 GPa), agarose (1–10 GPa), polyacrylamide (1–10 GPa), and polydimethylsiloxane (PDMS) (1–10 GPa). Solid supports described herein may have Young's moduli of 1–500, 1–40, 1–10, 1–5, or 3–11 GPa. The solid supports described herein may have a Young's modulus of approximately 1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 20, 25, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 400, 500 GPa or higher. Since the relationship between flexibility and rigidity is inverse, flexible materials have a low Young's modulus and change their shape significantly under load.

[0171] In some cases, the apparatus disclosed herein includes a silicon dioxide base and a silicon oxide surface. Alternatively, the apparatus may have a silicon oxide base. The surface of the apparatus provided herein may be textured to increase the overall surface area for polynucleotide synthesis. The apparatus disclosed herein may contain at least 5%, 10%, 25%, 50%, 80%, 90%, 95%, or 99% silicon. The apparatus disclosed herein may be made from silicon-on-insulator (SOI) wafers.

[0172] Surface architecture

[0173] Apparatus having raised and / or recessed features is provided herein. One advantage of having such features is the increased surface area supporting polynucleotide synthesis. In some examples, apparatus having raised and / or recessed features is called a three-dimensional substrate. In some examples, the three-dimensional apparatus includes one or more channels. In some examples, one or more loci include channels. In some examples, the channels are available for the deposition of reagents by a deposition apparatus such as a polynucleotide synthesizer. In some examples, the reagents and / or fluids accumulate in a larger well that is fluid-communicated with one or more channels. For example, the apparatus includes multiple channels corresponding to multiple loci having clusters, and the multiple channels are fluid-communicated with one well of the cluster. In some methods, a library of polynucleotides is synthesized at multiple loci of the cluster.

[0174] In some examples, the structure is configured to allow controlled flow and mass transfer pathways for polynucleotide synthesis on a surface. In some examples, the configuration of the apparatus allows for control and uniform distribution of mass transfer pathways, chemical exposure time, and / or washing effects during polynucleotide synthesis. In some examples, the configuration of the apparatus allows for increased efficiency of transfer (sweep) by providing sufficient volume for growing polynucleotides, for example, so that the volume shut out by the growing polynucleotides does not exceed 50, 45, 40, 35, 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1% or less of the initially available volume or the volume suitable for growing polynucleotides. In some examples, the three-dimensional structure allows for control of fluid flow to enable rapid exchange of chemical exposure.

[0175] This specification provides methods for synthesizing amounts of DNA of 1 fM, 5 fM, 10 fM, 25 fM, 50 fM, 75 fM, 100 fM, 200 fM, 300 fM, 400 fM, 500 fM, 600 fM, 700 fM, 800 fM, 900 fM, 1 pM, 5 pM, 10 pM, 25 pM, 50 pM, 75 pM, 100 pM, 200 pM, 300 pM, 400 pM, 500 pM, 600 pM, 700 pM, 800 pM, 900 pM, or more. In some examples, a polynucleotide library can be approximately 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% the length of the gene. The gene can vary by up to approximately 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 100%.

[0176] A non-identical polynucleotide can collectively encode sequences for at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 100% of a gene. In some examples, the polynucleotide can encode 50%, 60%, 70%, 80%, 85%, 90%, 95%, or more of the sequence of a gene. In some examples, the polynucleotide can encode 80%, 85%, 90%, 95%, or more of the sequence of a gene.

[0177] In some examples, isolation is achieved by physical structure. In some examples, isolation is achieved by differential functionalization of a surface that generates active and passive regions for polynucleotide synthesis. Differential functionalization is also achieved by alternating hydrophobicity across the device surface, thereby creating an effect on the water contact angle that causes beading or wetting of the deposited reagent. By utilizing larger structures, splashing and cross-contamination of separate polynucleotide synthesis sites with reagents in adjacent spots can be reduced. In some examples, a device such as a polynucleotide synthesizer is used to deposit reagents at separate polynucleotide synthesis sites. A substrate having three-dimensional features can synthesize a large number of polynucleotides (e.g., more than about 10,000) with a low error rate (e.g., less than about 1:500, less than about 1:1000, less than about 1:1500, less than about 1:2,000, less than about 1:3,000, less than about 1:5,000, or less than about 1:10,000). In some examples, the device contains features at a density of about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, or 500, or more per mm 2 and can contain features at a density of about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, or 500, or more per mm, or more.

[0178] The wells of the apparatus may have the same or different width, height, and / or volume as other wells of the substrate. The channels of the apparatus may have the same or different width, height, and / or volume as other channels of the substrate. In some examples, the cluster widths are approximately 0.05 mm to 50 mm, approximately 0.05 mm to 10 mm, approximately 0.05 mm to 5 mm, approximately 0.05 mm to 4 mm, approximately 0.05 mm to 3 mm, approximately 0.05 mm to 2 mm, approximately 0.05 mm to 1 mm, approximately 0.05 mm to 0.5 mm, approximately 0.05 mm to 0.1 mm, approximately 0.1 mm to 10 mm, approximately 0.2 mm to 10 mm, approximately 0.3 mm to 10 mm, approximately 0.4 mm to 10 mm, approximately 0.5 mm to 10 mm, approximately 0.5 mm to 5 mm, or approximately 0.5 mm to 2 mm. In some examples, the widths of the wells containing clusters are approximately 0.05 mm to 50 mm, 0.05 mm to 10 mm, 0.05 mm to 5 mm, 0.05 mm to 4 mm, 0.05 mm to 3 mm, 0.05 mm to 2 mm, 0.05 mm to 1 mm, 0.05 mm to 0.5 mm, 0.05 mm to 0.1 mm, 0.1 mm to 10 mm, 0.2 mm to 10 mm, 0.3 mm to 10 mm, 0.4 mm to 10 mm, 0.5 mm to 10 mm, 0.5 mm to 5 mm, or 0.5 mm and 2 mm. In some examples, the cluster width is less than 5mm, 4mm, 3mm, 2mm, 1mm, 0.5mm, 0.1mm, 0.09mm, 0.08mm, 0.07mm, 0.06mm, or 0.05mm. In some examples, the cluster width is approximately 1.0 to 1.3mm. In some examples, the cluster width is approximately 1.150mm. In some examples, the well width is less than 5mm, 4mm, 3mm, 2mm, 1mm, 0.5mm, 0.1mm, 0.09mm, 0.08mm, 0.07mm, 0.06mm, or 0.05mm. In some examples, the well width is approximately 1.0 to 1.3mm. In some examples, the well width is approximately 1.150mm. In some examples, the cluster width is approximately 0.08mm. In some examples, the well width is approximately 0.08mm. The cluster width can refer to clusters in a two-dimensional or three-dimensional substrate.

[0179] In some examples, the well heights are approximately 20 μm to 1000 μm, 50 μm to 1000 μm, 100 μm to 1000 μm, 200 μm to 1000 μm, 300 μm to 1000 μm, 400 μm to 1000 μm, or 500 μm to 1000 μm. In some examples, the well heights are less than approximately 1000 μm, less than approximately 900 μm, less than approximately 800 μm, less than approximately 700 μm, or less than approximately 600 μm.

[0180] In some examples, the device includes multiple channels corresponding to multiple loci within a cluster, with channel heights or depths ranging from approximately 5 μm to 500 μm, 5 μm to 400 μm, 5 μm to 300 μm, 5 μm to 200 μm, 5 μm to 100 μm, 5 μm to 50 μm, or 10 μm to 50 μm. In some examples, channel heights are less than 100 μm, less than 80 μm, less than 60 μm, less than 40 μm, or less than 20 μm.

[0181] In some examples, the diameters of channels, seats (e.g., in a substantially planar substrate), or both channels and seats (e.g., in a three-dimensional structural apparatus where the seat corresponds to a channel) are approximately 1 μm to approximately 1000 μm, approximately 1 μm to approximately 500 μm, approximately 1 μm to approximately 200 μm, approximately 1 μm to approximately 100 μm, approximately 5 μm to approximately 100 μm, or approximately 10 μm to approximately 100 μm, e.g., approximately 90 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, or 10 μm. In some examples, the diameters of channels, seats, or both channels and seats are approximately 100 μm, 90 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, or less than 10 μm. In some examples, the distance from the center of two adjacent channels, loci, or channels and loci is approximately 1 μm to 500 μm, approximately 1 μm to 200 μm, approximately 1 μm to 100 μm, approximately 5 μm to 200 μm, approximately 5 μm to 100 μm, approximately 5 μm to 50 μm, or approximately 5 μm to 30 μm, for example, approximately 20 μm.

[0182] surface modification

[0183] In various examples, surface modification is used for chemical and / or physical modification of a surface by adding or removing processes that change one or more chemical and / or physical properties of the device surface or selected parts or regions of the device surface. For example, surface modification includes, but is not limited to, (1) changing the wettability of the surface; (2) functionalizing the surface, i.e., providing, modifying or substituting surface functional groups; (3) defunctionalizing the surface, i.e., removing surface functional groups; (4) otherwise changing the chemical composition of the surface, for example by etching; (5) increasing or decreasing the surface roughness; (6) providing a coating on the surface, e.g., a coating that exhibits a different wettability from the surface; and / or (7) depositing particles on the surface.

[0184] In some cases, the addition of a chemical layer on the surface (called an adhesion promoter) facilitates the structured patterning of the sacs on the substrate surface. Typical surfaces for the application of adhesion promoters include, without limitation, glass, silicon, silicon dioxide, and silicon nitride. In some cases, the adhesion promoter is a chemical with a high surface energy. In some cases, a second chemical layer is deposited on the substrate surface. In some cases, the second chemical layer has a low surface energy. In some cases, the surface energy of the chemical layer coated on the surface supports droplet localization. Depending on the arrangement of the selected patterns, the area of ​​sac proximity and / or fluid contact at the sacs can be modified.

[0185] In some examples, the device surface or decomposed site on which nucleic acids or other parts are deposited, for example, for polynucleotide synthesis, may be smooth, substantially planar (e.g., two-dimensional), or have irregularities such as raised or recessed features (e.g., three-dimensional features). In some examples, the device surface is modified with one or more different layers of compounds. Such modifying layers in subject to application include, but are not limited to, inorganic and organic layers such as metals, metal oxides, polymers, and small organic molecules. In some examples, the device surface is modified with one or more different layers of compounds. Such modifying layers in subject to application include, but are not limited to, inorganic and organic layers such as metals, metal oxides, polymers, and small organic molecules. Non-limited polymer layers include peptides, proteins, nucleic acids or their mimics (e.g., peptide nucleic acids), polysaccharides, phospholipids, polyurethanes, polyesters, polycarbonates, polyureas, polyamides, polyethyleneamines, polyarylene sulfides, polysiloxanes, polyimides, polyacetates, and other suitable compounds described herein or otherwise known in the art. In some cases, the polymer is a heteropolymer. In some cases, the polymer is a homopolymer. In some cases, the polymer contains or is bonded to functional moieties.

[0186] In some examples, the decomposed seat of the apparatus is functionalized with one or more parts that increase and / or decrease the surface energy. In some examples, the parts are chemically inert. In some examples, some parts are configured to support one or more processes in a desired chemical reaction, e.g., polynucleotide synthesis. The surface energy, i.e., hydrophobicity, of the surface is a factor for determining the affinity of nucleotides to adhere to the surface. In some examples, a method for functionalizing the apparatus includes (a) providing an apparatus having a surface comprising silicon dioxide; and (b) silane-treating the surface using a suitable silanizing agent described herein or otherwise known in the art, e.g., an organofunctionalized alkoxysilane molecule.

[0187] In some examples, the organofunctional alkoxysilane molecules include dimethylchloro-octadecyl-silane, methyldichloro-octadecyl-silane, trichloro-octadecyl-silane, trimethyl-octadecyl-silane, triethyl-octadecyl-silane, or any combination thereof. In some examples, the surface of the device is functionalized with polyethylene / polypropylene (functionalized by gamma ray irradiation or chromic acid oxidation and reduction to a hydroxyalkyl surface), highly crosslinked polystyrene-divinylbenzene (derivatized by chloromethylation and aminated to a benzylamine functional surface), nylon (the terminal aminohexyl groups are directly reactive), or etched with reduced polytetrafluoroethylene. Other methods and functionalizing agents are described in U.S. Patent No. 5,474,796, which is hereby incorporated by reference in its entirety.

[0188] In some examples, the surface of the device is functionalized by contact with a derivatization composition containing a mixture of silanes under reaction conditions effective to bind the silanes to the surface of the device via reactive hydrophilic moieties typically present on the surface of the device. The silane treatment generally covers the surface with organofunctional alkoxysilane molecules by self-assembly.

[0189] As currently known in the art, for example, various siloxane functionalization reagents may further be used to decrease or increase surface energy. The organofunctional alkoxysilanes can be classified according to their organic functionality.

[0190] This specification provides an apparatus that may include patterning of agents that can be coupled to nucleosides. In some examples, the apparatus may be coated with an active agent. In some examples, the apparatus may be coated with a passive agent. Exemplary active agents included in the coating materials described herein include, without limitation, N-(3-triethoxysilylpropyl)-4-hydroxybutylamide (HAPS), 11-acetoxyundecyltriethoxysilane, n-decyltriethoxysilane, (3-aminopropyl)trimethoxysilane, (3-aminopropyl)triethoxysilane, 3-glycidoxypropyltrimethoxysilane (GOPS), 3-iodo-propyltrimethoxysilane, butyl-aldehyde-trimethoxysilane, dimeric secondary aminoalkylsiloxane, (3-aminopropyl)-diethoxy-methylsilane, (3-aminopropyl)-dimethyl-ethoxysilane, and (3-aminopropyl)-trimethoxysilane, (3-glycidoxypropyl)-dimethyl-ethoxysilane, glycidoxy-trimethoxysilane, (3-mercaptopropyl)-trimethoxysilane, 3-4 epoxycyclohexyl-ethyltrimethoxysilane, and (3-mercaptopropyl)-methyl-dimethoxysilane, allyltrichlorochlorosilane, 7-oct-1-enyltrichlorochlorosilane, or bis(3-trimethoxysilylpropyl)amine.

[0191] Typical passive agents for inclusion in the coating materials described herein include, but are not limited to, perfluorooctyltrichlorosilane; tridecafluoro-1,1,2,2-tetrahydrooctyl)trichlorosilane; 1H,1H,2H,2H-fluorooctyltriethoxysilane (FOS); trichloro(1H,1H,2H,2H -Perfluorooctyl)silane; tert-butyl-[5-fluoro-4-(4,4,5,5-tetramethyl-1,3,2-dioxaborolan-2-yl)indole-1-yl]-dimethylsilane; CYTOP(trademark); Fluorinert(trademark); perfluorooctyltrichlorosilane (PFOTCS); perfluorooctyldimethylchlorosilane (PFODCS); perfluorodecyltriethoxysilane (PFDTES); pentafluorophenyl-dimethylpropylchlorosilane (PFPTES); perfluorooctyltriethoxysilane; perfluorooctyltrimethoxysilane; octylchlorosilane; dimethylchloro-octodecyl-silane; methyldichloro-octodecyl-silane; trichloro-octodecyl-silane; trimethyl-octodecyl-silane; triethyl-octodecyl-silane; or octadecyltrichlorosilane.

[0192] In some examples, the functionalizing agent includes hydrocarbon silanes such as octadecyltrichlorosilane. In some examples, the functionalizing agent includes 11-acetoxyundecyltriethoxysilane, n-decyltriethoxysilane, (3-aminopropyl)trimethoxysilane, (3-aminopropyl)triethoxysilane, glycidyloxypropyl / trimethoxysilane, and N-(3-triethoxysilylpropyl)-4-hydroxybutylamide.

[0193] Polynucleotide synthesis

[0194] The methods of this disclosure for polynucleotide synthesis may include processes comprising phosphoramidite chemistry. In some examples, polynucleotide synthesis includes coupling a base with a phosphoramidite. Polynucleotide synthesis may also include coupling a base by depositing a phosphoramidite under coupling conditions, where the same base is optionally deposited with the phosphoramidite more than once, i.e., double coupling is performed. Polynucleotide synthesis may include capping of unreacted sites. In some examples, capping is optional. Polynucleotide synthesis may also include oxidation or an oxidation step. Polynucleotide synthesis may include deblocking, detritylation, and sulfidation. In some examples, polynucleotide synthesis includes either oxidation or sulfidation. In some examples, during one step in the polynucleotide synthesis reaction or between each step, the apparatus is washed with, for example, tetrazole or acetonitrile. The time frame for any one step in the phosphoramidite synthesis method may be about 2 minutes, 1 minute, 50 seconds, 40 seconds, 30 seconds, 20 seconds, and less than 10 seconds.

[0195] Polynucleotide synthesis using the phosphoramidite method may later include adding phosphoramidite building blocks (e.g., nucleoside phosphoramidites) to the growing polynucleotide chain to form a phosphate triester linkage. Phosphoramidite polynucleotide synthesis proceeds in the 3' to 5' direction. Phosphoramidite polynucleotide synthesis allows for the controlled addition of one nucleotide to the growing nucleic acid chain per synthetic cycle. In some examples, each synthetic cycle includes a coupling step. Phosphoramidite linkage involves the formation of a phosphate triester linkage between an activated nucleoside phosphoramidite and a nucleoside bonded to a substrate, for example, by a linker. In some examples, the nucleoside phosphoramidite is supplied to the activated apparatus. In some examples, the nucleoside phosphoramidite is supplied to the apparatus in an activator. In some examples, the nucleoside phosphoramidite is supplied to the apparatus in an excess amount of 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100 times, or more, compared to the nucleoside coupled to the substrate. In some examples, the addition of the nucleoside phosphoramidite is carried out in an anhydrous environment, for example, in anhydrous acetonitrile. Following the addition of the nucleoside phosphoramidite, the apparatus is optionally washed. In some examples, the coupling step is optionally repeated one or more additional times, along with the washing step between additions of nucleoside phosphoramidite to the substrate. In some examples, the polynucleotide synthesis method used herein includes one, two, three, or more consecutive coupling steps. Before coupling, in many cases, the nucleoside coupled to the apparatus is deprotected by the removal of a protecting group, where the protecting group functions to prevent polymerization. A common protecting group is 4,4'-dimethoxytrityl (DMT).

[0196] After coupling, the phosphoramidite polynucleotide synthesis method optionally includes a capping step. In the capping step, the growing polynucleotide is treated with a capping agent. The capping step is useful for blocking the 5'-OH group bound to the unreacted substrate after coupling from further chain elongation, preventing the formation of polynucleotides with internal base deletions. Furthermore, phosphoramidites activated with 1H-tetrazole may react slightly with the O6 position of guanosine. Without being bound by theory, oxidation with I2 / water may also result in depurination of this byproduct, possibly via O6-N7 migration. The depurination site will be cleaved during the final deprotection of the polynucleotide, thus potentially reducing the yield of the full-length product. The O6 modification can be removed by treatment with a capping reagent before oxidation with I2 / water. In some examples, including a capping step during polynucleotide synthesis reduces the error rate compared to synthesis without capping. As an example, the capping step involves treating the polynucleotides bound to the substrate with a mixture of acetic anhydride and 1-methylimidazole. Following the capping step, the apparatus is optionally cleaned.

[0197] In some cases, the growing nucleic acids bound to the apparatus are oxidized after the addition of nucleoside phosphoramidites and, optionally, after capping and one or more washing steps. The oxidation step involves oxidizing the phosphate triester to a tetracoordinate phosphate triester, which is a protected precursor of the nucleoside bond of naturally occurring phosphate diesters. In some cases, the oxidation of the growing polynucleotide is achieved by treatment with iodine and water, optionally in the presence of a weak base (e.g., pyridine, lutidine, colidine). Oxidation can be carried out under anhydrous conditions using, for example, tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, a capping step is carried out following oxidation. A second capping step allows the apparatus to be dried because residual water from the potentially persistent oxidation can inhibit subsequent coupling. After oxidation, the apparatus and the growing polynucleotide are optionally washed. In some examples, the oxidation step is replaced by a sulfurization step to obtain polynucleotide phosphorothioates, and an optional capping step may be performed after sulfurization. Many reagents, including, but not limited to, 3-(dimethylaminomethylidene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide (also known as Beaucage's reagent), and N,N,N'N'-tetraethylthiuram disulfide (TETD), can perform efficient sulfur transfer.

[0198] For the subsequent nucleoside incorporation cycle to occur via coupling, the protected 5' end of the growing polynucleotide bound to the apparatus is removed, allowing the primary hydroxyl group to react with the next nucleoside phosphoramidite. In some examples, the protecting group is DMT, and deblocking occurs with trichloroacetic acid in dichloromethane. Detritylation over a long period of time, or with a stronger detritylation than the recommended acid solution, can increase the depurination of the polynucleotide bound to the solid support, thus reducing the yield of the desired full-length product. The methods and compositions of this disclosure described herein provide controlled deblocking conditions that limit undesirable depurination reactions. In some examples, the polynucleotide bound to the apparatus is washed after deblocking. In some examples, efficient washing after deblocking contributes to the synthesized polynucleotide having a low error rate.

[0199] Methods for the synthesis of polynucleotides typically involve a series of iterating steps: applying a protected monomer to an actively functionalized surface (e.g., a seat) for binding to either an activated surface, a linker, or a previously deprotected monomer; deprotecting the applied monomer to react with a subsequently applied protected monomer; and applying another protected monomer for binding. One or more intermediate steps involve oxidation or sulfurization. In some examples, one or more washing steps precede or follow one or all of the steps.

[0200] Methods for the synthesis of phosphoramidite-based polynucleotides involve a series of chemical steps. In some examples, one or more steps of the synthesis method involve the cycling of reagents, where one or more steps of the method involve applying reagents useful for the process to the apparatus. For example, reagents are circulated by a series of liquid deposition and vacuum drying steps. In the case of a substrate having three-dimensional features such as wells, microwells, and channels, the reagents optionally pass through one or more areas of the apparatus via the wells and / or channels.

[0201] The methods and systems described herein relate to polynucleotide synthesizers for the synthesis of polynucleotides. Synthesis can be carried out in parallel. For example, at least or at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 1000, 10000, 50000, 75000, 100000, or more polynucleotides can be synthesized in parallel. The total number of polynucleotides that can be synthesized in parallel may be between 2–100,000, 3–50,000, 4–10,000, 5–1,000, 6–900, 7–850, 8–800, 9–750, 10–700, 11–650, 12–600, 13–550, 14–500, 15–450, 16–400, 17–350, 18–300, 19–250, 20–200, 21–150, 22–100, 23–50, 24–45, 25–40, and 30–35. Those skilled in the art will recognize that the total number of polynucleotides synthesized in parallel may be within any range (e.g., 25–100) constrained by any of these values. The total number of polynucleotides synthesized in parallel may be within any range defined by any of the values ​​that act as endpoints of the range. The total molar mass of polynucleotides synthesized in the apparatus, or the molar mass of each polynucleotide, may be at least or at least about 10, 20, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 25000, 50000, 75000, 100000 picomoles, or more. The length of each polynucleotide in the apparatus, or the average length of polynucleotides, may be at least or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500 nucleotides, or more.The length of each polynucleotide in the device, or the average length of the polynucleotides, may be at most or approximately at most about 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, or 10 nucleotides. The length of each polynucleotide in the device, or the average length of the polynucleotides, may be between 10 and 500, 9 and 400, 11 and 300, 12 and 200, 13 and 150, 14 and 100, 15 and 50, 16 and 45, 17 and 40, 18 and 35, or 19 and 25. A person skilled in the art will recognize that the length of each polynucleotide in the device, or the average length of the polynucleotides, may be within any range (e.g., 100 to 300) constrained by any of these values. The length of each polynucleotide within the device, or the average length of the polynucleotides, can be within any range defined by either of the values ​​that serve as the range endpoints.

[0202] The surface-based polynucleotide synthesis methods provided herein enable high-speed synthesis. For example, at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 125, 150, 175, 200 nucleotides, or more, are synthesized per hour. The nucleotides include adenine, guanine, thymine, cytosine, uridine building blocks, or their analogs / modified versions. In some examples, libraries of polynucleotides are synthesized in parallel on the substrate. For example, an instrument containing approximately or at least approximately 100;1,000;10,000;30,000;75,000;100,000;1,000,000;2,000,000;3,000,000;4,000,000; or 5,000,000 degraded loci can assist in the synthesis of at least the same number of distinct polynucleotides, where the polynucleotides encoding distinct sequences are synthesized at the degraded loci. In some examples, libraries of polynucleotides are synthesized on the instrument in approximately 3 months, 2 months, 1 month, 3 weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, less than 24 hours, or less, with the low error rates described herein. In some examples, large nucleic acids assembled from polynucleotide libraries synthesized with low error rates using the substrates and methods described herein can be prepared in approximately 3 months, 2 months, 1 month, 3 weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, less than 24 hours, or less.

[0203] In some examples, the methods described herein result in the generation of a library of polynucleotides containing different variant polynucleotides at multiple codon sites. In some examples, the polynucleotides may have one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, thirty, forty, fifty, or more variant codon sites.

[0204] In some examples, one or more mutant codon sites may be adjacent. One or more mutant codon sites may not be adjacent and may be separated by codons 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some examples, a polynucleotide may contain multiple mutant codon sites, where all mutant codon sites are adjacent to each other to form a stretch of mutant codon sites. In some examples, a polynucleotide may contain multiple mutant codon sites, where none of the mutant codon sites are adjacent to each other. In some examples, a polynucleotide may contain multiple mutant codon sites, where some of the mutant codon sites are adjacent to each other to form a stretch of mutant codon sites, and some of the mutant codon sites are not adjacent to each other.

[0205] Large polynucleotide libraries with low error rates

[0206] The average error rate of polynucleotides synthesized in a library using the provided system and method is less than 1 in 1000, less than 1 in 1250, less than 1 in 1500, less than 1 in 2000, less than 1 in 3000, or often lower. In some examples, the average error rate of polynucleotides synthesized in a library using the provided system and method is less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1250, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, or 1 / 3000, or lower. In some examples, the average error rate of polynucleotides synthesized in a library using the provided system and method is less than 1 / 1000.

[0207] In some examples, the total error rate of polynucleotides synthesized in the library using the provided system and method is less than or equal to 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1250, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, or 1 / 3000 compared to a predetermined sequence. In some examples, the total error rate for polynucleotides synthesized in the library using the provided system and method is less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, or 1 / 1000. In some examples, the total error rate of polynucleotides synthesized in the library using the provided system and method is less than 1 / 1000.

[0208] In some cases, error-correcting enzymes can be used with oligonucleotides synthesized in a library using the provided methods and systems. In some cases, the total error rate for oligonucleotides with error correction may be less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, or 1 / 3000 compared to a predetermined sequence. In some cases, the total error rate for oligonucleotides synthesized in a library using the provided systems and methods may be less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, or 1 / 1000. In some cases, the overall error rate with error correction for oligonucleotides synthesized in the library using the provided system and method may be less than 1 / 1000.

[0209] The error rate can limit the value of gene synthesis for the production of a library of gene variants. At an error rate of 1 / 300, approximately 0.7% of clones in a 1500-base-pair gene will be correct. Since most errors from oligonucleotide synthesis result in frameshift mutations, more than 99% of clones in such a library will not produce full-length proteins. By reducing the error rate by 75%, the fraction of correct clones will increase 40-fold. The methods and compositions of this disclosure enable rapid de novo synthesis of large polynucleotide and gene libraries at lower error rates than commonly observed gene synthesis methods, thanks to both improved synthesis quality and the applicability of error correction methods enabled in a massively parallel and time-efficient manner. Thus, libraries can be synthesized across the entire library, or at 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, and 99.98% of the library. 99.99% or more of the data consists of base insertions, deletions, substitutions, or 1 / 300, 1 / 400, 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 9000, 1 / 10000, 1 / 12000, 1 / 15000, 1 / 20000, 1 / 2 The total error rates can be combined to be less than or equal to 5000, 1 / 30000, 1 / 40000, 1 / 50000, 1 / 60000, 1 / 70000, 1 / 80000, 1 / 90000, 1 / 100000, 1 / 125000, 1 / 150000, 1 / 200000, 1 / 300000, 1 / 400000, 1 / 500000, 1 / 600000, 1 / 700000, 1 / 800000, 1 / 900000, or 1 / 1000000.The methods and compositions of the present disclosure further relate to libraries of large synthetic oligonucleotides and genes with low error rates associated with at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or higher of polynucleotides or genes in at least a subset of libraries associated with error-free sequences compared to predetermined / pre-selected sequences. In some examples, at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more of the polynucleotides or genes in the isolated amounts within the library have the same sequence. In some cases, at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.9%, or more of the polynucleotides or genes have the same sequence, relating to similarity or identity of 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more. In some cases, the error rate relating to a given locus on a polynucleotide or gene is optimized.Therefore, given loci of one or more polynucleotides or genes as part of a large library, each locus is divided into 1 / 300, 1 / 400, 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 9000, 1 / 10000, and 1 / 120. It may have an error rate of 00, 1 / 15000, 1 / 20000, 1 / 25000, 1 / 30000, 1 / 40000, 1 / 50000, 1 / 60000, 1 / 70000, 1 / 80000, 1 / 90000, 1 / 100000, 1 / 125000, 1 / 150000, 1 / 200000, 1 / 300000, 1 / 400000, 1 / 500000, 1 / 600000, 1 / 700000, 1 / 800000, 1 / 900000, less than 1 / 1000000, or lower. In various examples, the number of optimized seats for such errors is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 8 This may include seats 00, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 30000, 50000, 75000, 100000, 500000, 1000000, 2000000, 3000000, or more. The error-optimized loci may be distributed across at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 300000, 75000, 100000, 500000, 1000000, 2000000, 3000000, or more polynucleotides or genes.

[0210] The error rate can be achieved with or without error correction. The error rate can be achieved across the entire library or across 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more of the library.

[0211] Computer system

[0212] Any system described herein can be operably coupled to a computer and can be automated locally or remotely via the computer. In various examples, the methods and systems of the present disclosure may further include software programs on a computer system and their use. Thus, computer control for synchronization of dispensing / decompression / refilling functions, such as orchestrating and synchronizing the operation of a material deposition device, dispensing actions, and decompression operations, is within the scope of the present disclosure. The computer system is programmed to interfere between the base sequence specified by the user and the position of the material deposition device in order to deliver the correct reagent to the specified area of the substrate.

[0213] The computer system 1200 illustrated in Figure 16 is understood as a logical device capable of reading instructions from a medium 1211 and / or a network port 1205, and may optionally be connected to a server 1209 having a fixed medium 1212. A system such as that shown in Figure 16 may include a CPU 1201, a disk drive 1203, optional input devices such as a keyboard 1215 and / or a mouse 1216, and an optional monitor 1207. Data communication may be achieved through a communication medium indicated to the server at a local or remote location. The communication medium may include any means for transmitting and / or receiving data. For example, the communication medium may be a network connection, a wireless connection, or an Internet connection. Such a connection may provide communication over the World Wide Web. Data relating to this disclosure may be transmitted by such a network or connection for acceptance and / or consideration by Party 1222, as illustrated in Figure 16.

[0214] Figure 17 is a block diagram showing a first exemplary architecture of a computer system 1300 that may be used in connection with an exemplary instance of the present disclosure. As shown in Figure 17, an example of a computer system may include a processor 1302 for processing instructions. Examples of processors, not limited to: Intel Xeon® processors, AMD Opteron® processors, Samsung 32-bit RISC ARM 1176JZ(F)-S v1.0® processors, ARM Cortex-A8 Samsung S5PC100® processors, ARM Cortex-A8 Apple A4® processors, Marvell PXA 930® processors, or functionally equivalent processors. Execution of multiple threads may be used for parallel processing. In some examples, multiple processors, or processors with multiple cores, may also be used, whether in a single computer system, in a cluster, or distributed across a network system including multiple computers, mobile phones, and / or personal digital assistants.

[0215] As illustrated in Figure 17, the high-speed cache 1304 can be connected to or incorporated into the processor 1302 to provide high-speed memory for recently or frequently used instructions or data by the processor 1302. The processor 1302 is connected to the northbridge 1306 by the processor bus 1308. The northbridge 1306 is connected to the random access memory (RAM) 1310 by the memory bus 1312, which manages the processor 1302's access to the RAM 1310. The northbridge 1306 is also connected to the southbridge 1314 by the chipset bus 1316. The southbridge 1314 is connected to the peripheral bus 1318. The peripheral bus can be, for example, PCI, PCI-X, PCI Express, or other peripheral buses. The northbridge and southbridge are often referred to as the processor chipset and manage data transfer between the processor, RAM, and peripheral components on the peripheral bus 1318. In some alternative architectures, the functionality of the northbridge can be incorporated into the processor instead of using a separate northbridge chip. In some examples, system 1300 may include an accelerator card 1322 mounted on peripheral bus 1318. The accelerator may include a field-programmable gate array (FPGA) or other hardware to facilitate specific processing. For example, the accelerator may be used for reconstructing adaptive data or for evaluating algebraic expressions used in extended configuration processing.

[0216] Software and data are stored in external storage device 1324 and can be loaded into RAM 1310 and / or cache 1304 for use by the processor. System 1300 includes an operating system for managing system resources; examples of operating systems, not limited to these, include: Linux, Windows®, MACOS®, BlackBerry OS®, iOS®, and other functionally equivalent operating systems, as well as application software that runs on the operating system to manage data storage and optimization in accordance with the examples of this disclosure. In this example, system 1300 also includes external storage devices such as network-attached storage (NAS), and network interface cards (NICs) 1320 and 1321 connected to a peripheral bus to provide network interfaces to other computer systems that may be used for distributed parallel processing.

[0217] Figure 18 is a schematic diagram showing a network 1400 including multiple computer systems 1402a and 1402b, multiple mobile phones and personal digital assistants 1402c, and network-attached storage (NAS) 1404a and 1404b. In this example, systems 1402a, 1402b, and 1402c can manage data storage and optimize data access to data stored in network-attached storage (NAS) 1404a and 1404b. A mathematical model can be used on this data and evaluated using distributed parallel processing across computer systems 1402a and 1402b, as well as mobile phones and personal digital assistant systems 1402c. Computer systems 1402a and 1402b, as well as mobile phones and personal digital assistant systems 1402c, can also provide parallel processing for adaptive data reconstruction of data stored in network-attached storage (NAS) 1404a and 1404b. Figure 18 shows only one example, and various other computer architectures and systems may be used with the various examples of this disclosure. For example, a blade server may be used to provide parallel processing. Processor blades may be connected through a backplane to provide parallel processing. Storage may also be connected to the backplane through a separate network interface or as network-attached storage (NAS). In some examples, processors may maintain a separate memory space and transmit data through a network interface, a backplane, or other connectors for parallel processing by other processors. In other examples, some or all of the processors may be able to use a shared virtual address memory space.

[0218] Figure 19 is a block diagram of a multiprocessor computer system 1500 that uses a shared virtual address memory space according to an example of an embodiment. The system includes a plurality of processors 1502a-f that can access a shared memory subsystem 1504. The system incorporates a plurality of programmable hardware memory algorithm processors (MAPs) 1506a-f into the memory subsystem 1504. Each MAP 1506a-f may comprise a memory 1508a-f and one or more field-programmable gate arrays (FPGAs) 1510a-f. The MAPs provide configurable functional units, and specific algorithms or parts of algorithms can be provided to the FPGAs 150a-f for processing in close cooperation with their respective processors. For example, a MAP can be used to evaluate algebraic expressions relating to a data model and to perform adaptive data reconstruction in the example. In this example, each MAP is globally accessible by all processors for such purposes. In one configuration, each MAP can use direct memory access (DMA) to access its associated memory 1508a-f, enabling it to execute tasks independently and asynchronously from its respective microprocessor 1502a-f. In this configuration, a MAP can directly feed results to another MAP for pipeline processing and parallel execution of algorithms.

[0219] The computer architectures and systems described above are merely examples, and a wide variety of other computer, mobile phone, and personal digital assistant (PDCA) architectures and systems, including those using general-purpose processors, coprocessors, FPGAs and other programmable logic devices, systems on a chip (SOC), application-specific integrated circuits (ASICs), and any other combination of processing and logic elements, can be used in relation to exemplary instances. In some examples, all or part of the computer system can be implemented in software or hardware. Various data storage media can be used with examples including random-access memory, hard drives, flash memory, tape drives, disk arrays, network-attached storage (NAS), and other local or distributed data storage devices and systems.

[0220] In exemplary cases, a computer system may be implemented using software modules that run on any of the above or other computer architectures and systems. In other cases, the functionality of a system may be partially or completely implemented in firmware, programmable logic devices such as field-programmable gate arrays (FPGAs) as referenced in Figure 19, systems-on-chip (SOCs), application-specific integrated circuits (ASICs), or other processing and logic elements. For example, a set processor and optimizer may be implemented using hardware acceleration by using a hardware accelerator card such as the accelerator card 1322 shown in Figure 17. [Examples]

[0221] The following examples are given for the purpose of illustrating various embodiments of the invention and are not intended to limit the invention in any way. Together with the methods described herein, these examples are representative and typical of preferred embodiments and are not intended to limit the scope of the invention. Variations and other uses encompassed within the spirit of the invention as defined by the claims will be anticipated by those skilled in the art.

[0222] Example 1: Functionalization of the substrate surface

[0223] To support the attachment and synthesis of polynucleotide libraries, the substrate was functionalized. First, the substrate surface was wet-washed for 20 minutes with a piranha solution containing 90% H2SO4 and 10% H2O2. Then, the substrate was rinsed in multiple beakers with deionized water, held for 5 minutes with a deionized water gooseneck tap, and dried with N2. Subsequently, the substrate was immersed in NH4OH (1:100; 3 mL:300 mL) for 5 minutes, rinsed with deionized water using a hand gun, immersed in three consecutive beakers of deionized water for 1 minute each, and rinsed again with deionized water using a hand gun. Finally, the substrate was plasma-cleaned by exposure to O2. Using a SAMCO PC-300 instrument, O2 was plasma-etched in downstream mode at 250 watts for 1 minute.

[0224] The cleaned substrate surface was actively functionalized with a solution containing N-(3-triethoxysilylpropyl)-4-hydroxybutylamide) using a YES-1224P vapor deposition oven system with the following parameters: 0.5-1 tor, 60 minutes, 70°C, vaporization at 135°C. The substrate surface was resist-coated using a Brewer Science 200X spin coater. SPR(trademark)3612 photoresist was spin-coated onto the substrate at 2500 rpm for 40 seconds. The substrate was pre-bake on a Brewer hot plate at 90°C for 30 minutes. The substrate was exposed to photolithography using a Karl Suss MA6 mask aligner. The substrate was exposed for 2.2 seconds and developed with MSF26A for 1 minute. The remaining developer was rinsed off with a hand gun, and the substrate was immersed in water for 5 minutes. The substrate was baked in a 100°C oven for 30 minutes, and then lithography defects were visually confirmed using a Nikon L200. To remove residual resist, a SAMCO PC-300 system was used to perform a descam treatment, which involves O2 plasma etching at 250 watts for 1 minute.

[0225] The substrate surface was passively functionalized with a 100 μL solution of perfluorooctyltrichlorosilane mixed with 100 μL of light mineral oil. The substrate was placed in a chamber and pumped for 10 minutes, then the valve was closed relative to the pump and allowed to stand for 10 minutes. The chamber was vented. The substrate was stripped of its resist by immersion twice in 500 mL of NMP at 70°C for 5 minutes each, while sonicating at maximum power (9 in the Crest system). The substrate was then immersed in 500 mL of isopropanol at room temperature for 5 minutes, while sonicating at maximum power. The substrate was immersed in 300 mL of 200 proof ethanol and blow-dried with N2. The functionalized surface was activated to function as a support for polynucleotide synthesis.

[0226] Example 2: Synthesis of a 50-mer sequence on a polynucleotide synthesizer

[0227] A two-dimensional polynucleotide synthesizer was incorporated into a flow cell and connected to another flow cell (Applied Biosystems (ABI394 DNA Synthesizer)). The polynucleotide synthesizer was homogenized with N-(3-triethoxysilylpropyl)-4-hydroxybutyrate amide (N-(3-TRIETHOXYS1LYLPROPYL)-4-HYDROXYBUT YR AMIDE) (Gelest) to produce the following sequence: 5'AGACAATCAACCATTTGGGGTGGACAGCCTTGACCTCTAGACTTCGGCAT##TTTTTTTTTT3'(Sequence No. 1) An exemplary 50 bp polynucleotide ("50-mer polynucleotide") having the following characteristics was used to synthesize it (where # represents thymidine-succinylhexamide CED phosphoramidite, a cleavable linker that allows for the release of the polynucleotide from the surface during deprotection).

[0228] The synthesis was carried out using standard DNA synthesis chemistry (coupling, capping, oxidation, and deblocking).

[0229] The phosphoramidite / activator combination was delivered in the same manner as bulk reagents delivered through a flow cell. No drying steps were performed because the environment was "moist" with the reagents.

[0230] The flora restrictor was removed from the ABI 394 synthesizer, allowing for faster flow. Without a flora restrictor, the flow rates for amidite (0.1 M in ACN), activator (0.25 M benzoylthiotetrazole ("BTT"; 30-3070-xx from Glen Research) in ACN), and Ox (0.02 M I2 in 20% pyridine, 10% water, and 70% THF) were approximately ~100 μL / sec; for acetonitrile ("ACN") and capping reagent (a 1:1 mixture of CapA and CapB, where CapA is acetic anhydride in THF / pyridine and CapB is 16% 1-methylimidisole in THF), approximately ~200 μL / sec; and for Deblock (3% dichloroacetic acid in toluene), approximately ~300 μL / sec (compared to ~50 μL / sec for all reagents with a flora restrictor). The time required to completely flush out the oxidizer was observed, and the timing of chemical rinsing was adjusted accordingly, introducing extra ACN washing between different chemicals. After polynucleotide synthesis, the chip was deprotected in gaseous ammonia at 75 psi overnight. Five drops of water were added to the surface to recover the polynucleotides. The recovered polynucleotides were then analyzed on a BioAnalyzer small RNA chip (data not shown).

[0231] Example 3: Synthesis of a 100-mer sequence on a polynucleotide synthesizer

[0232] For the synthesis of the 50mer sequence, use the same process as described in Example 2 to create the following sequence: 5'CGGGATCCTTATCGTCATCGTCGTACAGATCCCGACCCATTTGCTGTCCACCAGTCATGCTAGCCATACCATGATGATGATGATGATGAGAACCCCGCAT##TTTTTTTTTT3' (SEQ ID NO: 2) Exemplary 100mer polynucleotides having the following characteristics were synthesized on two different silicon chips (where # represents thymidine-succinylhexamide CED phosphoramidite (CLP-2244 from ChemGenes)), one of which was homogeneously functionalized with N-(3-triethoxysilylpropyl)-4-hydroxybutylamide, and the other silicon chip was functionalized with a 5 / 95 mixture of 11-acetoxyundecyltriethoxysilane and n-decyltriethoxysilane. Polynucleotides extracted from the surface were analyzed on a BioAnalyzer instrument (data not shown).

[0233] All 10 samples from the two chips were subjected to the following temperature cycling program using forward primer (5'ATGCGGGGTTCTCATCATC3'; SEQ ID NO: 3) and reverse primer (5'CGGGATCCTTATCGTCATCG3'; SEQ ID NO: 4) in 50 μL of PCR mix (25 μL of NEB Q5 master mix, 2.5 μL of 10 μM forward primer, 2.5 μL of 10 μM reverse primer, 1 μL of polynucleotide extracted from the surface, and up to 50 μL of water). 98C, 30 seconds; Repeat the cycle of 98C, 10 seconds; 63C, 10 seconds; 72C, 10 seconds; 12. 72C, 2 minutes Further PCR amplification was performed using [this method].

[0234] The PCR products were also run on BioAnalyzer (data not shown) to demonstrate a sharp peak at the 100mer position. Next, the PCR-amplified samples were cloned and Sanger-sequenced. Table 7 summarizes the Sanger sequencing results for samples taken from spots 1-5 on tip 1 and from spots 6-10 on tip 2.

[0235] [Table 7]

[0236] Thus, the high quality and uniformity of the synthetic polynucleotides were repeated on two chips with different interfacial chemical properties. Of the 262 sequences in the 100mers, 89% (233) were error-free and complete sequences.

[0237] Finally, Table 8 summarizes the error features for sequences obtained from polynucleotide samples from spots 1-10.

[0238] [Table 8]

[0239] Example 4: Parallel assembly of 29,040 unique polynucleotides

[0240] As shown in Figure 15, a structure containing 256 clusters, each with 121 loci, was fabricated on a flat silicon plate. A magnified view of a cluster is shown at (205), which has 121 loci. 240 loci of the 256 clusters provided adhesion and support for the synthesis of polynucleotides with characteristic sequences. Polynucleotide synthesis was carried out by phosphoramidite chemistry using the general method from Example 3 above. 16 loci of the 256 clusters were control clusters. The comprehensive distribution of the 29,040 unique polynucleotides (240 × 121) synthesized is shown in Figure 20A. The polynucleotide library was synthesized with high uniformity. 90% of the sequences were present in signals within 4 times the average, enabling 100% representation. The distribution was measured for each cluster, as shown in Figure 20B. At a comprehensive level, all polynucleotides in motion were present, and 99% of the polynucleotides were present within 2 times the average, indicating the uniformity of the synthesis. These same observations were consistent at the cluster level as well.

[0241] The error rate for each polynucleotide was determined using an Illumina MiSeq gene sequencer. The distribution of error rates for the 29,040 unique polynucleotides averaged at approximately 1 error per 500 bases, with some having as low an error rate as 1 error per 800 bases. The distribution was measured for each cluster. The library of the 29,040 unique polynucleotides was synthesized in less than 20 hours. Analysis of GC percentage versus polynucleotide expression for all 29,040 unique polynucleotides showed that synthesis was uniform regardless of GC content.

[0242] Example 5. Design and synthesis of a synthetic cfDNA variant library

[0243] A synthetic variant library was designed and synthesized using the general synthesis method described in Example 3 above. The total number of expressed target variants was 458, and each polynucleotide in the library was 167 base pairs long. The variants were present in 85 different human genes and included SNVs (228), indels (215 in total; 168 deletions, 47 insertions), fusions, and SVs (15). This included 147 clinically relevant variants (including all SVs). Variants were selected from Tables 1-6. Polynucleotides targeting a single variant were tiled with a 4-base offset using the general design in Figure 23A, and 32 polynucleotides targeting each variant. The distribution of indel sizes in the library is shown in Figure 23B. The variant library was then mixed with a background cell-free DNA (cfDNA) library obtained from the plasma of healthy male donors (under 30 years old, shown in Figure 23C). Libraries were generated with variant allele frequencies (VAFs) of 0% (wild type), 0.1%, 0.25%, 0.5%, 1%, 2%, and 5%. The precise presentation and distribution of polynucleotides in the libraries were further confirmed by next-generation sequencing (all variant sites) and ddPCR (for subsets of variant sites).

[0244] Example 6. High-sensitivity detection of specific ultra-low frequency somatic mutations for MRD monitoring.

[0245] Five panels targeting minimal residual disease (MRD) were designed and developed to demonstrate detection sensitivity. The panels specifically targeted somatic variants found in breast, lung, CRC, melanoma, and renal cell carcinoma. Each of these MRD panels was designed to contain 197 targets with a selection of 3–5 variants and passenger mutations per tissue origin. The probe sequences for each panel were designed to incorporate variant alleles into the test sample set, for example, as shown in Figure 1B. To prepare the sample sets, synthetic variant sequences were mixed with fragmented cell line gDNA (NA12878) to form target specimens that approximated the profiles of cell-free and circulating tumor DNA. Five frequency levels were prepared with mean variant allele frequencies (VAF) of 0% (WT), 0.01%, 0.05%, 0.1%, and 2%. Libraries were prepared using UMI adapters as commonly described herein, and target enrichment was performed using the MRD panels.

[0246] At a sequencing depth of 80,000x, variant calling results revealed that an average of 20 SNV targets could be reliably detected in 0.01% VAF samples for each MRD panel and were clearly distinguishable from WT control samples. In addition to demonstrating the accuracy of variant calling by targeting alternative alleles, the usefulness of targeting multiple variants for detecting MRD signatures at very low levels (e.g., 0.01% VAF) is demonstrated. In summary, the panel performance showed high sensitivity for detecting very low-frequency somatic mutations.

[0247] method

[0248] To evaluate the detection sensitivity of the MRD panel, a pooled VAF series was constructed. The pooled VAF series was generated according to the schematic diagram shown in Figure 3. The pool contained synthetically designed variant sequences that mimic ctDNA, which were fragmented, end-repaired, A-tailed, and combined with background gDNA purified by a library-prepared enzymatic fragmentation (EF) kit, and densely mimicked the DNA size profile of native cfDNA. The ctDNA sequences were designed as a tiled pool of approximately 167 bp sequences that densely mimicked native ctDNA and covered 458 distinct mutations, including single nucleotide substitutions, small (2–4 bp), medium (5–9 bp), and large (10+ bp) insertions and deletions. Samples at five devised VAF levels (0% (WT), 0.01%, 0.05%, 0.1%, and 2%) were prepared, QC'd by ddPCR, and then libraries were prepared using mechanical fragmentation (MF) kits and UMI. These UMI libraries were captured in five panels (targeting both reference and alternative alleles) using a standard hybridization protocol for MRD application. The MRD-targeted enriched libraries were sequenced using Illumina Nextseq and analyzed using the UMI pipeline.

[0249] The bioinformatics workflow for sequencing results followed the schematic diagram shown in Figure 4. After base calling and FASTQ generation, reads were first downsampled to a fixed depth based on the panel's target space. Reads were then preprocessed to mark adapter sequences (Picard), isolate UMI sequences (fgbio), and obtain unaligned BAM files. Raw reads were aligned to the human reference genome (hg38 / GRCh38) using BWA and merged with the unaligned BAMs to provide UMI information. After alignment, UMIs were error-corrected, grouped based on strand and UMI sequence, and consensus reads were called using a duplex strategy (fgbio). Unless otherwise specified, reads were subsequently filtered to retain only those containing a duplex consensus family or at least one supporting read derived from each strand. After the consensus call, raw allele counts were obtained using samtools, or variant calls were obtained using Mutect2 (GATK) as needed.

[0250] Panel performance

[0251] We designed five 200-probe MRD panels using a proprietary algorithm. These five panels specifically targeted somatic variants found in breast, lung, CRC, melanoma, and renal cell carcinoma. To demonstrate panel performance, libraries were prepared using a mechanical library preparation kit with 30 ng of WT cfDNA Pan-cancer Reference Standard and a UMI Adapter System for targeted enrichment and duplex sequencing. We also developed a standard hybridization protocol for MRD applications and optimized it to further improve MRD panel performance.

[0252] Illumina Nextseq sequencing results showed that this upgraded system dramatically improved the performance of small panels, reducing the off-target rate in each panel to as low as 10–15% (Figure 7A) and demonstrating uniform coverage across all targets of interest (Figure 7B). The off-target rate did not increase significantly in the UMI consensus analysis pipeline, suggesting that the successful probe design kept the specific off-target rate low. In addition, no targets were dropped in the MRD panels (Figure 7C), demonstrating that the standard hybridization protocol for the MRD panel manufacturing process and MRD application workflow is highly robust and accurate. With 80,000x downsampling, the average target coverage of each panel after UMI deduplication was approximately 3,000x, providing sufficient coverage for variant calling analysis (Figure 7D).

[0253] Panel detection sensitivity

[0254] Across all panel types, approximate variant allele frequencies (Figure 8) were detected at a 60% expected dilution frequency, likely due to a combination of capture and alignment bias. The mean error rate for WT samples was approximately 10 ppm (0.001%), and the error range (0.01%) between WT samples and the lowest VAF samples was approximately 7-fold, demonstrating the ability to clearly distinguish between WT samples and 0.01% VAF samples. Five MRD panels targeting different somatic variants also showed very similar performance, demonstrating consistency in design strategies across different targets. Incorporating variant alleles instead of reference alleles into the probe improved detection sensitivity by 5–10% across different panel types at 0.05% and 0.1% VAF levels. At the 0.01% VAF level, different MRD panels showed conflicting positive site calling between the targeted reference allele and the alternative allele, likely due to sampling error at this low VAF level.

[0255] Regarding target site recall (Figure 9), nearly 100% of targets were detected at the 2% VAF level. While VAF affected sensitivity, variant calling results revealed the average of at least 20 / 200 SNV targets detected at 0.01% VAF, with a sequencing depth of 80,000x across all panels. At VAF levels of 0.05% to 0.1%, we observed a slight but consistent increase in recall (detection defined as at least one supporting duplex consensus read). Targeting alternative alleles in the MRD panel resulted in approximately 5–10% improvement in recall for 0.05% samples, but little improvement at 0.01%, likely due to sampling effects as discussed above.

[0256] Indel detection sensitivity

[0257] Insertion / deletion mutations (indels) are involved as drivers in many cancers and can therefore be important in clinical NGS. Indel detection rates can generally be influenced by the mapping parameters of short-read aligners, the normalization scheme used to represent indel alignments, and biases arising from the use of targeted capture sequencing. Consequently, the concordance rate of indel detection tools from short-read targeted sequencing may be low.

[0258] Figure 10 shows the results of investigating the indel detection sensitivity of MRD panels, illustrating the recall split across variant types. Variant types were categorized into SBS, single-base indels, small indels (<5 bp), medium indels (5-10 bp), and large indels (>10 bp). Recall was plotted as strip plots for each variant type at each VAF level; for example, if there are 50 medium indels, detecting 20 / 50 in a 0.05% VAF sample is equivalent to a 40% recall. Each point in Figure 10 represents an experiment, with panels folded together to represent the conditions. General trends observed included: little difference was observed for 0.01% samples; indels tended to show more dramatic improvement in the alternative panel than in the SBS panel; and effects could be difficult to estimate (e.g., large indels were penalized, and therefore, even if there was improvement in the alt panel, it could be difficult to judge).

[0259] To further demonstrate the indel detection sensitivity of the MRD panel, variant calling was performed using both K-mer-based searches and raw pile-up from duplex-consensus three alignments, and all variants detected from each experiment were aggregated for different conditions. The results showed that the combination of targeted surrogate alleles for enrichment and K-mer-based search methods dramatically improved the detection sensitivity of larger indels (2+bp). Approximately 10% of all indels could be called at the 0.01% VAF level and were clearly distinguishable from WT control samples. Differences in indel recall tended to be most evident for VAF samples between 0.05% and 0.1%. At the 0.1% VAF level, recall could be improved from approximately 25% to approximately 75% for large events (>10bp). Furthermore, since medium and large indels were penalized, it was difficult to predict improvements within this range even with alternative panels. At a 2% VAF level, 100% of indels were detected by targeting alternative alleles along with a K-mer-based search method, demonstrating the advantages of this approach (Figure 11). Focusing on the mean VAF detection rate at each VAF level, the application of K-mer-based searches and targeting of alternative alleles achieved consistent VAFs similar to the target VAF. Targeting alternative alleles instead of the reference allele showed a much more apparent improvement, particularly at the 2% VAF condition. The rate of random error was very high for SBS and single-nucleotide indels, but low for larger tire indels (Figure 12).

[0260] ROC curve analysis

[0261] ROC (Receiver Operating Characteristic) analysis was developed as a standard methodology for quantifying the ability of a signal receiver to accurately distinguish an object of interest from background noise in a system. ROC curves are generally used to graphically illustrate the relationship / trade-off between clinical sensitivity and clinical specificity for any possible cutoff in a trial or combination thereof. In addition, the area under the ROC curve can provide an idea of ​​the benefits of using the trial in question, which can offer meaningful interpretations for disease classification from healthy subjects.

[0262] ROC analysis (Figure 13) was applied to the MRD target-enriched variant calling dataset to (1) simulate the diagnostic power of the MRD panel compared to approaches profiling fewer sites, and (2) find the optimal thresholds for detectable VAF levels and target variant sites. From the ROC analysis, a larger number of targets (>50 sites) was beneficial for detection from lower VAF (<0.01%) samples. When the number of MRD targets was less than 50 sites in 0.01% VAF samples, the simulated ROC curve was close to the diagonal, suggesting difficulty in distinguishing between true positive and false positive variants. However, when the number of MRD targets was higher than 50 sites at the 0.01% VAF level, the MRD test was able to distinguish between true positive and false positive variants with near 100% sensitivity and 100% specificity. Furthermore, even with a low number of targets of 10 sites, the test still demonstrated excellent accuracy and precision at VAF levels higher than 0.05%.

[0263] In further experiments, VAF dilution experiments were conducted using background material including ultra-low dilutions (0.001% / 10 ppm), and the results are shown in Figure 14. Based on the results, the variant was detectable against WT, which served as good proof of the UMI system concept.

[0264] Based on the results shown, better detection sensitivity and specificity of the MRD test can be achieved by incorporating more target sites (>50 sites) into the MRD panel for samples with a VAF level of 0.01% or less, or by obtaining samples with a VAF level of 0.05% or more.

[0265] While preferred embodiments of the subject matter have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Numerous variations, alterations, and substitutions will be conceivable to those skilled in the art without departing from the subject matter. It should be understood that various alternatives to the embodiments of the subject matter described herein may be adopted in the practice of the subject matter.

[0266] This disclosure is further described by the following non-limiting items.

[0267] Item 1. A polynucleotide library containing multiple polynucleotides, each containing at least one variant associated with minimal residual disease (MRD).

[0268] Item 2. A library as described in Item 1, in which at least one variant is located within 20 base pairs from the center of each sequence of multiple polynucleotides.

[0269] Item 3. A library as described in Item 1, in which at least one variant sequence is located within 10% of the center of each sequence of multiple polynucleotides.

[0270] Item 4. A library as described in Item 1, in which the respective positions of at least one variant in each sequence of multiple polynucleotides include a distribution that includes the mean.

[0271] Item 5. The mean is the center of each sequence, as described in Item 4.

[0272] Item 6. A library as described in Item 4, where the mean is within 20 base pairs from the center of each sequence.

[0273] Item 7. The library described in Item 4, where the mean is within 10% of the center of each sequence.

[0274] Embodiment 8. The library described in item 4, wherein the distribution is a normal distribution.

[0275] Item 9. A library containing multiple polynucleotides, each with a length of 150 bases or less, as described in any one of items 1-8.

[0276] Item 10. A library described in any one of items 1-9, in which at least one variant sequence is derived from a genome sequence.

[0277] Item 11. A library whose genome sequence is derived from cell-free DNA (cfDNA), as described in any one of items 1-10.

[0278] Item 12. A library described in any one of items 1 to 11, wherein at least one of the above variants is present at a frequency of 0.001% to 0.1% relative to the wild-type genome sequence.

[0279] Item 13. A library described in any one of items 1 through 12, in which at least one variant contains approximately 500 variants.

[0280] Item 14. A library as described in Item 13, in which each polynucleotide contains at least one variant of one variant.

[0281] Item 15. A library described in any one of items 1 to 14, wherein at least one variant is located in at least 150 genes.

[0282] Item 16. A library containing multiple polynucleotides that are double-stranded, as described in any one of items 1-15.

[0283] Item 17. A library described in any one of items 1 through 16, in which at least one variant involves insertion, deletion, fusion, duplication, frameshift, iterative expansion, or substitution.

[0284] Item 18. A library described in any one of items 1-17, in which at least one variant includes a copy number variant (CNV), microsatellite instability, loss of heterozygosity (LOH), DNA methylation, early stop codon, trinucleotide repeat, translocation, somatic rearrangement, allelomorphism, single nucleotide variant (SNV), indel, splice variant, regulatory variant, copy number variant, or fusion.

[0285] Item 19. A library described in any one of items 1 through 18, wherein at least one variant includes a single nucleotide variant, indel, fusion, or structural variant.

[0286] Item 20. A library described in any one of items 1-19, in which at least one variant contains a modification to a tumor suppressor gene or oncogene.

[0287] Item 21. A library that includes a buffer, as described in any one of items 1 through 20.

[0288] Item 22. A library described in any one of items 1-21, further comprising a background set containing background polynucleotides, wherein the background set contains cell-free DNA (cfDNA).

[0289] Item 23. A library as described in Item 22, in which at least one variant contains one or more changes compared to the background polynucleotides of the background set.

[0290] Item 24. A kit for detecting minimal residual disease (MRD) in a sample, (a) A library listed in any one of items 1 to 23, (b) Instructions for using the kit, (c) Packaging configured to hold and explain the contents of the kit A kit that includes this.

[0291] Item 25. The kit described in Item 24, further comprising a second library described in any one of Items 1-23.

[0292] Item 26. The kit described in Item 25, wherein the second library contains at least one variant sequence with a different frequency compared to the library.

[0293] Item 27. A kit as described in Item 25 or 26, wherein the second library contains a variant sequence different from that of the library.

[0294] Item 28. A method for preparing a library as described in any one of items 1 to 23, (a) A step of providing at least one variant sequence associated with MRD, (b) A step of synthesizing multiple polynucleotides containing at least one variant. Methods that include...

[0295] Item 29. The method of Item 28, further comprising the step of providing a background set.

[0296] Item 30. The method according to item 28 or 29, further comprising the step of mixing a background set with a plurality of polynucleotides containing at least one variant.

[0297] Item 31. The method according to Item 30, wherein the step of mixing a background set with multiple polynucleotides is to mix the background set with multiple polynucleotides such that at least one variant sequence is present at a frequency of 0%, 0.01%, 0.05%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genome sequence.

[0298] Item 32. The method described in any one of items 28 to 31, wherein the synthesis step includes chemical synthesis.

[0299] Item 33. The method according to any one of items 28 to 32, wherein the synthesis step includes synthesis on a surface.

[0300] Item 34. The method according to any one of items 28 to 33, wherein the synthesis step involves the coupling of a nucleoside phosphoramidite.

[0301] Item 35. The method described in any one of items 28-34, further comprising the step of sequencing a library.

[0302] Item 36. The method described in any one of items 28-35, further comprising ddPCR measurement of the library.

[0303] Item 37. The method according to any one of items 28-36, further comprising fluorescence / UV DNA quantification and size distribution of the library.

[0304] Item 38. A method for detecting minimal residual disease (MRD) in a sample, (a) The process of providing a library as described in any one of items 1 to 23, (b) A step of bringing the library into contact with the sample, (c) A step of detecting the presence or absence of one or more variant sequences associated with MRD in the sample. Methods that include...

[0305] Item 39. The method described in Item 38, wherein the detection process includes sequencing.

[0306] Item 40. Sequencing, including next-generation sequencing, as described in Item 39.

[0307] Item 41. The method according to Item 39, wherein the sequencing includes synthesis sequencing, nanopore sequencing, or SMRT sequencing.

[0308] Item 42. The method according to any one of items 38-41, wherein the detection step includes ddPCR or specific hybridization to an array.

[0309] Item 43. The method according to any one of items 38 to 42, wherein at least one variant is present in the sample at a frequency of approximately 0.001% to 0.1%.

[0310] Item 44. The method described in any one of items 38-43, further comprising the step of obtaining a sample from an individual.

[0311] Item 45. The method described in Item 44, which indicates that the individual has been previously treated, is currently being treated, or has had a clinical diagnosis of cancer.

[0312] Item 46. The method described in any one of items 38-45, wherein the sample includes a liquid biopsy.

[0313] Item 47. The method described in any one of items 38-46, wherein the sample contains circulating tumor DNA (ctDNA).

[0314] Item 48. A method according to any one of items 38-47, wherein the sample is obtained from blood.

[0315] Item 49. The method described in any one of items 38-48, wherein the sample is substantially cell-free.

[0316] Item 50. The method according to any one of items 38-49, further comprising the step of ligating a sequencing adapter to at least several polynucleotides in a test sample, library, or both.

[0317] Item 51. The method according to any one of items 38-50, further comprising the step of amplifying at least some polynucleotides in a sample, library, or both.

[0318] The method described in any one of items 38-51, wherein the recall of item 52.1 or multiple variant sequences is at least 5% greater than that of multiple polynucleotides that do not contain one or more variant sequences.

[0319] The method described in any one of items 38-52, wherein the reproduction of item 53.1 or multiple variant sequences is 5% to 10% greater than that of multiple polynucleotides that do not contain one or more variant sequences.

Claims

1. A polynucleotide library comprising a plurality of polynucleotides, wherein each of the plurality of polynucleotides comprises a nucleic acid sequence having a center, and the nucleic acid sequence of each polynucleotide comprises at least one variant sequence associated with minimal residual disease (MRD).

2. The polynucleotide library according to claim 1, wherein the position of at least one MRD-associated variant sequence is within 20 bases from the center of each nucleic acid sequence of each polynucleotide.

3. The polynucleotide library according to claim 1, further comprising a distribution of the positions of at least one MRD-associated variant sequence in the nucleic acid sequences of all of the plurality of polynucleotides, wherein the distribution includes the mean of the central 20 bases or less of each nucleic acid sequence of each polynucleotide.

4. The polynucleotide library according to claim 1, wherein the nucleic acid sequence of each polynucleotide is 150 bases or less in length.

5. The polynucleotide library according to claim 1, wherein at least one MRD-associated variant sequence is derived from a genome sequence.

6. The polynucleotide library according to claim 5, wherein the genome sequence is derived from cell-free DNA (cfDNA).

7. The polynucleotide library according to claim 1, wherein at least one MRD-associated variant sequence is present in the plurality of polynucleotides at a frequency of 0.001% to 0.1% relative to the wild-type genome sequence.

8. The polynucleotide library according to claim 1, wherein the plurality of polynucleotides include about 500 variant sequences associated with MRD.

9. The polynucleotide library according to claim 1, wherein the nucleic acid sequence of each polynucleotide includes a variant sequence of at least one MRD-associated variant sequence.

10. The polynucleotide library according to claim 1, wherein at least one MRD-related variant is present in the nucleic acid sequences of at least 150 genes.

11. The polynucleotide library according to claim 1, wherein at least one MRD-associated variant sequence comprises a modification of a nucleic acid sequence of a tumor suppressor gene or oncogene.

12. The polynucleotide library according to claim 1, further comprising a background set of polynucleotides, wherein the at least one MRD-associated variant sequence is at least one base pair different from the nucleic acid sequence of the polynucleotides in the background set.

13. The polynucleotide library according to claim 12, wherein the background set includes cell-free DNA (cfDNA).

14. A method for preparing a polynucleotide library containing multiple polynucleotides, A step of providing at least one variant sequence associated with minimal residual disease (MRD), The steps include synthesizing multiple polynucleotides, each containing at least one MRD-related variant sequence, to produce the polynucleotide library, and Methods that include...

15. A process of providing a background set of polynucleotides, A step of mixing the background set and the plurality of polynucleotides such that at least one MRD-associated variant sequence is present at a frequency of 2% or less relative to the wild-type genome sequence. The method according to claim 14, further comprising:

16. The method according to claim 14, wherein synthesis includes chemical synthesis.

17. The method according to claim 14, wherein the synthesis includes synthesis on a surface.

18. The method according to claim 14, wherein the synthesis involves the coupling of nucleoside phosphoramidites.

19. Steps to sequence the aforementioned polynucleotide library The method according to claim 14, further comprising:

20. A method for detecting minimal residual disease (MRD) in a sample, A step of providing the polynucleotide library described in claim 1, The process involves bringing the polynucleotide library into contact with a sample, The steps include detecting the presence or absence of at least one variant associated with MRD in the sample, and Methods that include...